Skip to main content
A service tier is the processing class a provider runs a request on. The same model can be served ahead of standard traffic, on discounted spare capacity, or in the standard queue, and a provider that sells more than one class exposes it as service_tier. The AI Gateway passes the value through and reports back the tier that served the request. Use Cases
  • Run latency-sensitive work ahead of standard traffic with fast.
  • Move background work onto discounted spare capacity with flex.
  • Record the tier that served each request, since billing follows it.

Quick Start

Set service_tier in the request body. The AI Gateway accepts six values and sends the request to the provider on that tier, where the model sells it.
service_tier works on the Responses API and on Chat Completions.

Accepted values

Any other value is rejected before the provider call, on both endpoints:

Tier availability

A model sells a tier only where its pricing declares one, so the set differs per model and changes as providers revise their line-ups. A tier appears as a pricing variant on the model, in the model catalog, and on List Models. Read service_tier back rather than assuming the tier was honoured; a model that does not sell the requested tier does not fail the request, it runs on the standard tier and reports default.

Reading back the served tier

Every response carries service_tier, naming the tier that served the request:
fast is reported as priority, and auto, scale and default are reported as default. A request that sets no tier is reported as default.