> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Create rerank

> Rerank a list of documents based on their relevance to a query.

<Note>
  **Related guide**: Rerank and moderations guide. See the [Rerank and moderations guide](/ai-gateway/features/rerank-and-moderations) for a walkthrough with examples.
</Note>


## OpenAPI

````yaml post /v2/router/rerank
openapi: 3.1.0
info:
  title: orq.ai API
  version: '2.0'
  description: orq.ai API documentation
servers:
  - url: https://my.orq.ai
security:
  - ApiKey: []
tags:
  - name: Chunking
    description: Split text into smaller chunks for retrieval and generation workflows.
  - name: File Systems
    description: >-
      Create and manage persistent file systems that agents and MCP clients read
      from and write to.
  - name: Knowledge Bases
    description: Create and manage knowledge bases used by agents and retrieval workflows.
  - name: Memory Stores
    description: Create and manage memory stores, memories, and memory documents.
  - name: Evals
    description: Run an evaluator against a conversation and its result
  - name: Logs
    description: >-
      OpenTelemetry log query API. Search, filter, aggregate, and facet log
      records ingested via OTLP.
  - name: Reporting
    description: >-
      GenAI reporting API over canonical analytics rollups. Accepts a metric
      name, time range, grain, group-by, and filters; returns a typed time
      series and optional totals.
  - name: Traces
    description: >-
      Query and inspect ingested trace data: search trace summaries, aggregate
      metrics, and read individual traces and their spans.
  - description: List models available through the AI Router.
    name: Models
  - name: Policies
  - name: Alerts
    description: >-
      Alerts evaluate a Reporting API metric on a fixed interval and fire
      notifications through notifiers when the value breaches a threshold. Each
      breach opens a trigger that tracks the incident until the value recovers.
  - name: Annotation Queues
    description: Annotation queues collect spans for human review.
  - name: API keys
    description: >-
      API keys authenticate programmatic access to the workspace. They expose
      opaque tokens, per-domain access grants, and budget and rate-limit
      constraints.
  - name: Audit Logs
    description: Audit logs record workspace entity changes and access-relevant events.
  - name: Budgets
    description: >-
      Budgets govern spend, token usage, and request rate across six scopes:
      workspace, project, identity, API key, provider, and model. Every
      applicable budget is enforced, and the most restrictive limit applies per
      dimension.
  - name: Files
    description: File upload and retrieval operations.
  - name: Guardrail Rules
    description: >-
      Guardrail Rules conditionally enforce evaluators and plugins for AI
      Gateway traffic. Rules may be scoped to a project or the whole workspace.
  - name: Hub
    description: Hub items are reusable templates available to a workspace.
  - name: Identities
    description: >-
      Identities represent end users from your system for usage and engagement
      tracking.
  - name: Management keys
    description: >-
      Management keys are workspace-scoped credentials that authenticate
      programmatic access to workspace administration surfaces (API keys,
      budgets). Unlike project-scoped API keys, a management key always operates
      at the workspace level.
  - name: MCP Gateway
    description: >-
      Register upstream MCP servers, discover and sync their tools, and assemble
      gateways that expose a curated tool surface to MCP clients.
  - name: Model Catalog
    description: >-
      Browse the orq.ai model catalog: every model orq offers, across every
      provider, with pricing, capabilities and benchmark data. List endpoints
      only return models that are not deprecated. This API is public, requires
      no authentication, and is rate limited to 120 requests per minute per IP.
      Responses carry a 5-minute cache-control max-age.
  - name: Notifiers
    description: Notifier destinations used to send delivery and workflow notifications.
  - name: Projects
    description: Projects organize resources within a workspace
  - name: Routing Rules
    description: >-
      Routing Rules conditionally select models and enforce request plugins for
      AI Gateway traffic. Rules are evaluated by ascending priority and may be
      scoped to a project or the whole workspace.
  - name: Threads
    description: Threads group related trace invocations and their aggregate usage
  - name: Skills
    description: >-
      Skills are modular instructions you can use to codify processes and
      conventions
  - name: Smart Routers
    description: >-
      Create and manage workspace Smart Routers. A Smart Router selects a model
      from an eligible pool for each request according to a quality, balanced,
      or cost profile.
  - name: Webhooks
    description: >-
      Create and manage webhooks that deliver workspace events to external HTTPS
      endpoints.
  - name: Workspaces
    description: >-
      A workspace is the tenant. Create is called from a user session during
      onboarding; Get, List, and Update are the public management surface.
  - name: Workspace Security
    description: >-
      Workspace-level domain verification and IP allowlist controls. These
      operations are restricted to workspace administrators.
  - name: Workspace Settings
    description: >-
      Workspace-level settings managed with a workspace credential. A workspace
      is the tenant, so these settings are a singleton — there is nothing to
      create or delete, only read and update.
  - name: Responses
  - description: Run agents on a cron cadence. Minimum firing interval is 1 hour.
    name: Agent Schedules
  - name: Embeddings
  - name: Telemetry
    description: >-
      Unified query envelope for traces, metrics, and logs. One request shape,
      one filter dialect, and one response shape per source, validated by a
      per-source registry.
  - description: Beta. Run typed classification questions against a classify model.
    name: Classify
  - description: Search Gateway with managed credits or BYOK.
    name: Web Search
externalDocs:
  url: https://docs.orq.ai
  description: orq.ai Documentation
paths:
  /v2/router/rerank:
    post:
      tags:
        - Rerank
      summary: Create rerank
      description: Rerank a list of documents based on their relevance to a query.
      operationId: createRerank
      requestBody:
        required: true
        description: input
        content:
          application/json:
            schema:
              type: object
              properties:
                query:
                  type: string
                  description: The search query
                documents:
                  type: array
                  items:
                    type: string
                  description: >-
                    A list of texts that will be compared to the `query`. For
                    optimal performance we recommend against sending more than
                    1,000 documents in a single request.
                model:
                  type: string
                  description: The identifier of the model to use
                top_n:
                  type: integer
                  minimum: 1
                  description: >-
                    The number of most relevant documents or indices to return,
                    defaults to the length of the documents
                return_documents:
                  type: boolean
                  description: Whether to return the documents in the response
                filename:
                  type:
                    - string
                    - 'null'
                  description: The filename of the document to rerank
                name:
                  description: >-
                    The name to display on the trace. If not specified, the
                    default system name will be used.
                  type: string
                fallbacks:
                  type: array
                  items:
                    type: object
                    properties:
                      model:
                        type: string
                        description: Fallback model identifier
                    required:
                      - model
                  description: Array of fallback models to use if primary model fails
                retry:
                  type: object
                  properties:
                    count:
                      type: number
                      minimum: 1
                      maximum: 5
                      default: 3
                      description: Number of retry attempts (1-5)
                      example: 3
                    on_codes:
                      type: array
                      items:
                        type: number
                        minimum: 100
                        maximum: 599
                      minItems: 1
                      description: HTTP status codes that trigger retry logic
                      example:
                        - 429
                        - 500
                        - 502
                        - 503
                        - 504
                  description: Retry configuration for the request
                cache:
                  type: object
                  properties:
                    ttl:
                      type: number
                      minimum: 1
                      maximum: 259200
                      default: 1800
                      description: >-
                        Time to live for cached responses in seconds. Maximum
                        259200 seconds (3 days).
                      example: 3600
                    type:
                      type: string
                      enum:
                        - exact_match
                  required:
                    - type
                  description: Cache configuration for the request.
                load_balancer:
                  oneOf:
                    - type: object
                      properties:
                        type:
                          type: string
                          enum:
                            - weight_based
                        models:
                          type: array
                          items:
                            type: object
                            properties:
                              model:
                                type: string
                                description: Model identifier for load balancing
                                example: openai/gpt-5.6-sol
                              weight:
                                type: number
                                minimum: 0.001
                                maximum: 1
                                default: 0.5
                                description: >-
                                  Weight assigned to this model for load
                                  balancing
                                example: 0.7
                            required:
                              - model
                      required:
                        - type
                        - models
                  description: Load balancer configuration for the request.
                timeout:
                  type: object
                  properties:
                    call_timeout:
                      type: number
                      minimum: 1
                      description: Timeout value in milliseconds
                      example: 30000
                  required:
                    - call_timeout
                  description: >-
                    Timeout configuration to apply to the request. If the
                    request exceeds the timeout, it will be retried or fallback
                    to the next model if configured.
                plugins:
                  type: array
                  items:
                    $ref: '#/components/schemas/PIIRedactionPlugin'
                  description: >-
                    Request-scoped transforms applied to the text exchanged with
                    the model. Currently supports `pii_redaction`, which
                    replaces PII with placeholders before the provider sees it
                    and restores the original values in the response.
                orq:
                  type: object
                  properties:
                    name:
                      description: >-
                        The name to display on the trace. If not specified, the
                        default system name will be used.
                      type: string
                    fallbacks:
                      type: array
                      items:
                        type: object
                        properties:
                          model:
                            type: string
                            description: Fallback model identifier
                            example: openai/gpt-5.4-mini
                        required:
                          - model
                      description: Array of fallback models to use if primary model fails
                    cache:
                      type: object
                      properties:
                        ttl:
                          type: number
                          minimum: 1
                          maximum: 259200
                          default: 1800
                          description: >-
                            Time to live for cached responses in seconds.
                            Maximum 259200 seconds (3 days).
                          example: 3600
                        type:
                          type: string
                          enum:
                            - exact_match
                      required:
                        - type
                      description: Cache configuration for the request.
                    retry:
                      type: object
                      properties:
                        count:
                          type: number
                          minimum: 1
                          maximum: 5
                          default: 3
                          description: Number of retry attempts (1-5)
                          example: 3
                        on_codes:
                          type: array
                          items:
                            type: number
                            minimum: 100
                            maximum: 599
                          minItems: 1
                          description: HTTP status codes that trigger retry logic
                          example:
                            - 429
                            - 500
                            - 502
                            - 503
                            - 504
                      description: Retry configuration for the request
                    identity:
                      $ref: '#/components/schemas/PublicIdentity'
                    contact:
                      $ref: '#/components/schemas/PublicContact'
                    load_balancer:
                      oneOf:
                        - type: object
                          properties:
                            type:
                              type: string
                              enum:
                                - weight_based
                            models:
                              type: array
                              items:
                                type: object
                                properties:
                                  model:
                                    type: string
                                    description: Model identifier for load balancing
                                    example: openai/gpt-5.6-sol
                                  weight:
                                    type: number
                                    minimum: 0.001
                                    maximum: 1
                                    default: 0.5
                                    description: >-
                                      Weight assigned to this model for load
                                      balancing
                                    example: 0.7
                                required:
                                  - model
                          required:
                            - type
                            - models
                      description: Load balancer configuration for the request.
                      example:
                        type: weight_based
                        models:
                          - model: openai/gpt-4o
                            weight: 0.7
                          - model: anthropic/claude-3-5-sonnet
                            weight: 0.3
                    timeout:
                      type: object
                      properties:
                        call_timeout:
                          type: number
                          minimum: 1
                          description: Timeout value in milliseconds
                          example: 30000
                      required:
                        - call_timeout
                      description: >-
                        Timeout configuration to apply to the request. If the
                        request exceeds the timeout, it will be retried or
                        fallback to the next model if configured.
              required:
                - query
                - documents
                - model
      responses:
        '200':
          description: Returns the reranked documents.
          content:
            application/json:
              schema:
                type: object
                properties:
                  id:
                    type: string
                    description: A unique identifier for the rerank.
                  object:
                    enum:
                      - list
                    type: string
                  results:
                    type: array
                    items:
                      type: object
                      properties:
                        object:
                          type: string
                          enum:
                            - rerank
                          description: The object type, which is always `rerank`.
                        index:
                          type: number
                          description: >-
                            Corresponds to the index in the original list of
                            documents to which the ranked document belongs.
                        relevance_score:
                          type: number
                          description: >-
                            Relevance scores are normalized to be in the range
                            [0, 1]. Scores close to 1 indicate a high relevance
                            to the query, and scores closer to 0 indicate low
                            relevance.
                        document:
                          type: object
                          properties:
                            text:
                              type: string
                              description: The text of the document to rerank
                          required:
                            - text
                          description: >-
                            If return_documents is set as false this will return
                            none, if true it will return the documents passed in
                      required:
                        - object
                        - index
                        - relevance_score
                    description: An ordered list of ranked documents
                  usage:
                    type: object
                    properties:
                      total_tokens:
                        type: number
                        description: The total number of tokens used in the rerank
                    required:
                      - total_tokens
                required:
                  - object
                  - results
components:
  schemas:
    PIIRedactionPlugin:
      type: object
      properties:
        id:
          type: string
          enum:
            - pii_redaction
          description: PII redaction plugin.
        language:
          type: string
          description: >-
            Detector language. Accepts "auto" to detect the language per
            request; GET /v2/pii/capabilities lists the concrete languages and
            does not include "auto". Omitting the field falls back to en.
        regions:
          type: array
          items:
            type: string
          description: >-
            Region codes selecting whole regions of coverage (e.g. "nl", "gb").
            Every entity type those regions cover is redacted, alongside the
            base catalog. ["all"] is exclusive, and leaving both this and
            entities empty also runs every region, so selecting nothing is the
            widest request rather than the narrowest. Combines with entities:
            the two selections are unioned, so entities adds types on top of the
            region coverage.
        entities:
          type: array
          items:
            type: string
          description: >-
            The entity types to redact. A named type is redacted even when it
            belongs to a region, so a region's types can be selected
            individually without naming the region. On its own this is a strict
            allowlist; alongside regions it adds to the region coverage. See GET
            /v2/pii/capabilities for valid types.
        entity_thresholds:
          type: object
          additionalProperties:
            type: number
            minimum: 0
            maximum: 1
          description: >-
            Per-entity confidence cutoff overrides in [0,1]. Tunes confidence
            only and never changes which types are redacted, so every key must
            also appear in entities.
        threshold:
          type: number
          minimum: 0
          maximum: 1
          description: >-
            Baseline confidence cutoff applied to every entity type without a
            per-entity override.
        on_failure:
          type: string
          enum:
            - block
            - passthrough
          description: Behaviour when detection fails.
      required:
        - id
      additionalProperties: false
      title: PII redaction plugin
    PublicIdentity:
      type: object
      properties:
        id:
          type: string
          description: Unique identifier for the contact
          example: contact_01ARZ3NDEKTSV4RRFFQ69G5FAV
        display_name:
          type: string
          description: Display name of the contact
          example: Jane Doe
        email:
          type: string
          format: email
          description: Email address of the contact
          example: jane.doe@example.com
        metadata:
          type: array
          items:
            type: object
            additionalProperties: {}
          description: >-
            A hash of key/value pairs containing any other data about the
            contact
          example:
            - department: Engineering
              role: Senior Developer
        logo_url:
          type: string
          description: URL to the contact's avatar or logo
          example: https://example.com/avatars/jane-doe.jpg
        tags:
          type: array
          items:
            type: string
          description: A list of tags associated with the contact
          example:
            - hr
            - engineering
      required:
        - id
      description: >-
        Information about the identity making the request. If the identity does
        not exist, it will be created automatically.
    PublicContact:
      type: object
      properties:
        id:
          type: string
          description: Unique identifier for the contact
          example: contact_01ARZ3NDEKTSV4RRFFQ69G5FAV
        display_name:
          type: string
          description: Display name of the contact
          example: Jane Doe
        email:
          type: string
          format: email
          description: Email address of the contact
          example: jane.doe@example.com
        metadata:
          type: array
          items:
            type: object
            additionalProperties: {}
          description: >-
            A hash of key/value pairs containing any other data about the
            contact
          example:
            - department: Engineering
              role: Senior Developer
        logo_url:
          type: string
          description: URL to the contact's avatar or logo
          example: https://example.com/avatars/jane-doe.jpg
        tags:
          type: array
          items:
            type: string
          description: A list of tags associated with the contact
          example:
            - hr
            - engineering
      required:
        - id
      description: >-
        @deprecated Use identity instead. Information about the contact making
        the request.
      deprecated: true
  securitySchemes:
    ApiKey:
      type: http
      scheme: bearer
      bearerFormat: JWT

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.