Classify is in Beta, and its request and response contract matches the TypeSafe Classify API.
- Answers natively on OpenAI GPT-6 Luna and
typesafe/jev-latest, and through structured outputs on a set of small chat models: see Supported models. - Does not apply PII plugins, guardrails, or evaluators yet: see Enforcement during the Beta.
- Requires an API key with the
classifypermission.
Overview
Classify answers typed questions about a piece of content in one call. Send the content asstate and one or more named questions to POST /decisions or POST /classify on the AI Gateway base URL, https://my.orq.ai/v3/router.
Both endpoints accept the same request and return the same response. Existing /classify clients can continue using that path. The SDK methods are orq.router.decisions.create() and orq.router.classify.create().
The response returns one structured answer per question, so nothing has to be parsed out of free-form text. A choice or score answer also carries the probabilities behind it and a confidence for the selected option or level.
A question pairs plain-language instructions with a type that fixes the answer shape, and criteria that define the rubric:
state is not limited to text: an object or array is serialized as JSON before classification, so the same call can judge a customer message, a tool payload, or a whole record.
Several questions can be asked at once, each with its own rubric, so a request costs one round trip instead of one call per question. See Question types for the criteria each type expects.
The request is traced, priced, and rate-limited like other gateway traffic. See Enforcement during the Beta for what applies today.
Use cases
- Triage and routing. Ask
noulwhether the message reports a problem andchoicefor the topic or product, then route on the labels and escalate the rows whose confidence is low. - Building labelled data. Turn a rubric into a
choicequestion to label a corpus for Evaluators or Datasets, and useprobabilitiesto pick the rows worth human review. - Scoring open feedback. Use
scoreon survey answers, support transcripts, or review text to get an ordinal reading without writing a scoring prompt. - Gating expensive work. Use a
noulquestion on the native model, whose output tokens are free, as a pre-check, and send only the requests that pass on to a larger model. - Enriching structured records. Pass a record as
stateand ask one question per field to derive, instead of building a prompt per field. - Comparing models on one rubric. Ask the same questions against
typesafe/jev-latestand a chat model to see where they disagree. Thresholds tuned on one do not transfer, so compare answers rather than raw probabilities.
Supported models
The table lists model families. Every offering of these models, including regional and reseller variants, is in Supported models.
Chat models return the same fields with their own probabilities, which express the model’s estimate and are not calibrated, so thresholds tuned on one model do not transfer to another.
For emulated classification, reasoning is set to the minimum the model allows.
OpenAI Decisions accepts text or arrays of user messages with
input_text and inline base64 input_image parts in state. Hosted image URLs, file IDs, and non-user roles are rejected. Content parts must be nested inside user messages; bare content parts and tool items are rejected rather than serialized as text. Other objects and arrays are serialized as text. OpenAI can return {"type": "refusal"} under an individual question key; handle refusals before reading scored fields.
Any other model returns 400.
Quick start
Request
Each question has a
type and instructions, which can be a string, object, or array. The shape of criteria depends on the question type.
Question types
noul
Returns the probability that the statement in instructions is true for state. The probability is calibrated on typesafe/jev-latest and model-reported on chat models.
The optional criteria object can define what counts as true and false.
JSON
choice
Selects one option from a rubric. The required criteria object maps each option name to a description. Set a description to null when the option name is enough for the model to interpret it.
JSON
score
Returns a position on an ordered scale. Set criteria to an array with at least two level descriptions, ordered from lowest to highest. The answer index refers to this array.
JSON
Response
Native providers return their ownconfidence and score. Emulated classification uses one structured-output call. OpenAI Decisions returns a probability-weighted score and preserves its own confidence. The fields below describe these provider differences.
JSON
The response includes optional
telemetry. Read x-orq-trace-id and x-orq-trace-span-id from the response headers when older SDKs or deployments omit it. The handler sets those once the request is traced, so a request rejected before that point carries neither: authentication and authorization (401, and 403 for a missing classify permission or a model that is not enabled for the workspace or shared with the project), plan and budget limits (429), and request-body validation (the 400 for a malformed body, the 422 for an invalid request). Everything rejected after tracing carries them, including the 400 for a model that does not support classify. The telemetry object itself is optional and returned by builds newer than some deployments run, so fall back to the headers when it is absent.
Emulated chat models return the same scored fields, but the numbers do not mean the same thing: their probabilities are the model’s own estimates and are not calibrated, and their confidence is simply the largest of them. Compare choice and score across models, not confidence.
See the Create Classify API reference for the full request and response schema, including the fields accepted alongside the ones above.
Fallbacks, retries and identity
Use the same options on/decisions and /classify:
retry.count is the number of retries after the first attempt, from 1 to 5. Retries happen only for status codes in on_codes; omitting retry disables gateway retries. The gateway honors Retry-After or waits one second when it is absent. After retries are exhausted, it tries the next fallback. A request accepts at most 10 fallbacks. The response reports the successful model and that model’s costs. A chat fallback pays the chat model’s input and output rates.
For inline-image requests, every fallback must support native OpenAI image classification; incompatible chains are rejected with 422 before a provider call.
An inaccessible fallback stops the remaining fallback chain. Session keys must allow every requested model under their factory policy.
Enforcement during the Beta
Enforced on every request:
Not applied during the Beta:
- PII plugins.
stateandinstructionsreach the provider exactly as sent, even when a redaction plugin is configured for traffic that would otherwise match. - Guardrails and evaluators. They do not run, so a matching guardrail rule cannot block a Classify request.
- Routing rules. A rule does not rewrite the model, add fallbacks, or serve from cache.
- Request-level
load_balancer,timeout,cache,plugins, andguardrails. They are not part of this contract. Unknown body fields are ignored rather than rejected, so sending them has no effect and returns no error.