Concept
Inference
Serving stacks run forward passes through trained models to answer prompts, score embeddings, or classify inputs. Autoscaling watches queue depth while quantization trades accuracy for latency. Metrics on tokens per second and error rates guide capacity planning.