> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Service tiers

> Request a premium or discounted provider processing tier with the service_tier parameter on the AI Gateway, and read back the tier that served the request.

A service tier is the processing class a provider runs a request on. The same model can be served ahead of standard traffic, on discounted spare capacity, or in the standard queue, and a provider that sells more than one class exposes it as `service_tier`. The **AI Gateway** passes the value through and reports back the tier that served the request.

**Use Cases**

* Run latency-sensitive work ahead of standard traffic with `fast`.
* Move background work onto discounted spare capacity with `flex`.
* Record the tier that served each request, since billing follows it.

***

## Quick Start

Set `service_tier` in the request body. The **AI Gateway** accepts six values and sends the request to the provider on that tier, where the model sells it.

<CodeGroup>
  ```bash cURL theme={"theme":{"light":"github-light","dark":"github-dark"}}
  curl -X POST https://my.orq.ai/v3/router/responses \
    -H "Authorization: Bearer $ORQ_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "openai/gpt-5.4-mini",
      "input": "Summarize the report.",
      "service_tier": "flex"
    }'
  ```

  ```typescript TypeScript theme={"theme":{"light":"github-light","dark":"github-dark"}}
  import OpenAI from "openai";

  const client = new OpenAI({
    apiKey: process.env.ORQ_API_KEY,
    baseURL: "https://my.orq.ai/v3/router",
  });

  const response = await client.responses.create({
    model: "openai/gpt-5.4-mini",
    input: "Summarize the report.",
    service_tier: "flex",
  });

  console.log(response.service_tier);
  ```

  ```python Python theme={"theme":{"light":"github-light","dark":"github-dark"}}
  from openai import OpenAI
  import os

  client = OpenAI(
      api_key=os.environ.get("ORQ_API_KEY"),
      base_url="https://my.orq.ai/v3/router",
  )

  response = client.responses.create(
      model="openai/gpt-5.4-mini",
      input="Summarize the report.",
      service_tier="flex",
  )

  print(response.service_tier)
  ```
</CodeGroup>

`service_tier` works on the [Responses API](/ai-gateway/features/responses-api) and on [Chat Completions](/ai-gateway/features/openai-compatible-api).

## Accepted values

| Value | Effect |
| - | - |
| `flex` | Spare provider capacity, billed below the standard rate |
| `fast` | Premium low-latency processing, billed above the standard rate |
| `priority` | Backward-compatible alias for `fast` |
| `auto` | Standard processing |
| `default` | Standard processing |
| `scale` | Standard processing |

Any other value is rejected before the provider call, on both endpoints:

```json theme={"theme":{"light":"github-light","dark":"github-dark"}}
{
  "error": {
    "message": "Invalid value for 'ServiceTier': must be one of [auto, default, flex, fast, scale, or priority].",
    "code": "invalid_request_error",
    "param": "ServiceTier"
  }
}
```

## Tier availability

A model sells a tier only where its pricing declares one, so the set differs per model and changes as providers revise their line-ups. A tier appears as a pricing variant on the model, in the [model catalog](/ai-gateway/supported-models), and on [List Models](/reference/models/list-models). Read `service_tier` back rather than assuming the tier was honoured; a model that does not sell the requested tier does not fail the request, it runs on the standard tier and reports `default`.

## Reading back the served tier

Every response carries `service_tier`, naming the tier that served the request:

```json theme={"theme":{"light":"github-light","dark":"github-dark"}}
{
  "service_tier": "priority"
}
```

`fast` is reported as `priority`, and `auto`, `scale` and `default` are reported as `default`. A request that sets no tier is reported as `default`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.