> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Create embeddings

> Get a vector representation of a given input that can be easily consumed by machine learning models and algorithms.

<Note>
  **Related guide**: Embeddings guide. See the [Embeddings guide](/ai-gateway/features/embeddings) for a walkthrough with examples.
</Note>


## OpenAPI

````yaml post /v2/router/embeddings
openapi: 3.1.0
info:
  title: orq.ai API
  version: '2.0'
  description: orq.ai API documentation
servers:
  - url: https://my.orq.ai
security:
  - ApiKey: []
tags:
  - name: Chunking
    description: Split text into smaller chunks for retrieval and generation workflows.
  - name: File Systems
    description: >-
      Create and manage persistent file systems that agents and MCP clients read
      from and write to.
  - name: Knowledge Bases
    description: Create and manage knowledge bases used by agents and retrieval workflows.
  - name: Memory Stores
    description: Create and manage memory stores, memories, and memory documents.
  - name: Evals
    description: Run an evaluator against a conversation and its result
  - name: Logs
    description: >-
      OpenTelemetry log query API. Search, filter, aggregate, and facet log
      records ingested via OTLP.
  - name: Reporting
    description: >-
      GenAI reporting API over canonical analytics rollups. Accepts a metric
      name, time range, grain, group-by, and filters; returns a typed time
      series and optional totals.
  - name: Traces
    description: >-
      Query and inspect ingested trace data: search trace summaries, aggregate
      metrics, and read individual traces and their spans.
  - description: List models available through the AI Router.
    name: Models
  - name: Policies
  - name: Alerts
    description: >-
      Alerts evaluate a Reporting API metric on a fixed interval and fire
      notifications through notifiers when the value breaches a threshold. Each
      breach opens a trigger that tracks the incident until the value recovers.
  - name: Annotation Queues
    description: Annotation queues collect spans for human review.
  - name: API keys
    description: >-
      API keys authenticate programmatic access to the workspace. They expose
      opaque tokens, per-domain access grants, and budget and rate-limit
      constraints.
  - name: Audit Logs
    description: Audit logs record workspace entity changes and access-relevant events.
  - name: Budgets
    description: >-
      Budgets govern spend, token usage, and request rate across six scopes:
      workspace, project, identity, API key, provider, and model. Every
      applicable budget is enforced, and the most restrictive limit applies per
      dimension.
  - name: Files
    description: File upload and retrieval operations.
  - name: Guardrail Rules
    description: >-
      Guardrail Rules conditionally enforce evaluators and plugins for AI
      Gateway traffic. Rules may be scoped to a project or the whole workspace.
  - name: Hub
    description: Hub items are reusable templates available to a workspace.
  - name: Identities
    description: >-
      Identities represent end users from your system for usage and engagement
      tracking.
  - name: Management keys
    description: >-
      Management keys are workspace-scoped credentials that authenticate
      programmatic access to workspace administration surfaces (API keys,
      budgets). Unlike project-scoped API keys, a management key always operates
      at the workspace level.
  - name: MCP Gateway
    description: >-
      Register upstream MCP servers, discover and sync their tools, and assemble
      gateways that expose a curated tool surface to MCP clients.
  - name: Model Catalog
    description: >-
      Browse the orq.ai model catalog: every model orq offers, across every
      provider, with pricing, capabilities and benchmark data. List endpoints
      only return models that are not deprecated. This API is public, requires
      no authentication, and is rate limited to 120 requests per minute per IP.
      Responses carry a 5-minute cache-control max-age.
  - name: Notifiers
    description: Notifier destinations used to send delivery and workflow notifications.
  - name: Projects
    description: Projects organize resources within a workspace
  - name: Routing Rules
    description: >-
      Routing Rules conditionally select models and enforce request plugins for
      AI Gateway traffic. Rules are evaluated by ascending priority and may be
      scoped to a project or the whole workspace.
  - name: Threads
    description: Threads group related trace invocations and their aggregate usage
  - name: Skills
    description: >-
      Skills are modular instructions you can use to codify processes and
      conventions
  - name: Smart Routers
    description: >-
      Create and manage workspace Smart Routers. A Smart Router selects a model
      from an eligible pool for each request according to a quality, balanced,
      or cost profile.
  - name: Webhooks
    description: >-
      Create and manage webhooks that deliver workspace events to external HTTPS
      endpoints.
  - name: Workspaces
    description: >-
      A workspace is the tenant. Create is called from a user session during
      onboarding; Get, List, and Update are the public management surface.
  - name: Workspace Security
    description: >-
      Workspace-level domain verification and IP allowlist controls. These
      operations are restricted to workspace administrators.
  - name: Workspace Settings
    description: >-
      Workspace-level settings managed with a workspace credential. A workspace
      is the tenant, so these settings are a singleton — there is nothing to
      create or delete, only read and update.
  - name: Responses
  - description: Run agents on a cron cadence. Minimum firing interval is 1 hour.
    name: Agent Schedules
  - name: Embeddings
  - name: Telemetry
    description: >-
      Unified query envelope for traces, metrics, and logs. One request shape,
      one filter dialect, and one response shape per source, validated by a
      per-source registry.
  - description: Beta. Run typed classification questions against a classify model.
    name: Classify
  - description: Search Gateway with managed credits or BYOK.
    name: Web Search
externalDocs:
  url: https://docs.orq.ai
  description: orq.ai Documentation
paths:
  /v2/router/embeddings:
    post:
      tags:
        - Embeddings
      summary: Create embeddings
      description: >-
        Get a vector representation of a given input that can be easily consumed
        by machine learning models and algorithms.
      operationId: createEmbedding
      requestBody:
        content:
          application/json:
            examples:
              array_of_strings:
                summary: Array of strings input
                value:
                  input:
                    - The food was delicious
                    - And the waiter was friendly
                  model: openai/text-embedding-3-small
              single_string:
                summary: Single string input
                value:
                  input: The food was delicious and the waiter...
                  model: openai/text-embedding-3-small
            schema:
              additionalProperties: false
              properties:
                cache:
                  $ref: '#/components/schemas/EmbeddingCacheConfig'
                  description: Cache configuration for the request.
                dimensions:
                  description: >-
                    The number of dimensions the resulting output embeddings
                    should have.
                  format: int64
                  type: integer
                encoding_format:
                  description: >-
                    The format to return the embeddings in. Can be either float
                    or base64.
                  enum:
                    - float
                    - base64
                  type: string
                fallbacks:
                  description: Array of fallback models to use if primary model fails.
                  items:
                    $ref: '#/components/schemas/FallbackConfig'
                  type:
                    - array
                    - 'null'
                input:
                  anyOf:
                    - description: A single text string to embed. Must be non-empty.
                      minLength: 1
                      title: String
                      type: string
                    - description: >-
                        An array of strings or token arrays to embed. Must be
                        non-empty.
                      items:
                        anyOf:
                          - type: string
                            minLength: 1
                          - items:
                              type: integer
                            minItems: 1
                            type: array
                      minItems: 1
                      title: Array
                      type: array
                  description: Input text to embed, encoded as a string or array of tokens.
                load_balancer:
                  $ref: '#/components/schemas/EmbeddingLoadBalancerConfig'
                  description: Load balancer configuration for the request.
                model:
                  description: ID of the model to use.
                  type: string
                name:
                  description: >-
                    The name to display on the trace. If not specified, the
                    default system name will be used.
                  type: string
                orq:
                  $ref: '#/components/schemas/EmbeddingOrqParams'
                  description: >-
                    Orq platform extension parameters. Top-level equivalents
                    take priority when both are set.
                retry:
                  $ref: '#/components/schemas/EmbeddingRetryConfig'
                  description: Retry configuration for the request.
                timeout:
                  $ref: '#/components/schemas/EmbeddingTimeoutConfig'
                  description: Timeout configuration to apply to the request.
                user:
                  description: A unique identifier representing your end-user.
                  type: string
              required:
                - model
                - input
              type: object
        required: true
      responses:
        '200':
          content:
            application/json:
              schema:
                additionalProperties: false
                properties:
                  data:
                    description: List of embedding objects.
                    items:
                      $ref: '#/components/schemas/PublicEmbeddingData'
                    minItems: 0
                    type:
                      - array
                      - 'null'
                  model:
                    description: ID of the model used.
                    type: string
                  object:
                    description: Always "list".
                    enum:
                      - list
                    type: string
                  usage:
                    $ref: '#/components/schemas/PublicEmbeddingUsage'
                    description: The usage information for the request.
                required:
                  - object
                  - data
                  - model
                  - usage
                type: object
          description: Returns the embedding vector.
components:
  schemas:
    EmbeddingCacheConfig:
      additionalProperties: false
      properties:
        ttl:
          description: >-
            Time to live for cached responses in seconds. Maximum 259200 seconds
            (3 days).
          format: int64
          type: integer
        type:
          description: Cache type.
          enum:
            - exact_match
          type: string
      required:
        - type
      type: object
    FallbackConfig:
      additionalProperties: false
      properties:
        model:
          type: string
      required:
        - model
      type: object
    EmbeddingLoadBalancerConfig:
      additionalProperties: false
      properties:
        models:
          description: Array of models with weights for load balancing requests.
          items:
            $ref: '#/components/schemas/EmbeddingLoadBalancerModelConfig'
          type:
            - array
            - 'null'
        type:
          description: Load balancer type.
          enum:
            - weight_based
          type: string
      required:
        - type
        - models
      type: object
    EmbeddingOrqParams:
      additionalProperties: false
      properties:
        cache:
          $ref: '#/components/schemas/EmbeddingCacheConfig'
          description: 'Deprecated: use top-level cache instead.'
        contact:
          $ref: '#/components/schemas/EmbeddingContactParams'
          description: >-
            Deprecated: use identity instead. Information about the contact
            making the request.
        fallbacks:
          description: 'Deprecated: use top-level fallbacks instead.'
          items:
            $ref: '#/components/schemas/FallbackConfig'
          type:
            - array
            - 'null'
        identity:
          $ref: '#/components/schemas/ResponseIdentity'
          description: Information about the identity making the request.
        load_balancer:
          $ref: '#/components/schemas/EmbeddingLoadBalancerConfig'
          description: 'Deprecated: use top-level load_balancer instead.'
        name:
          description: 'Deprecated: use top-level name instead.'
          type: string
        retry:
          $ref: '#/components/schemas/EmbeddingRetryConfig'
          description: 'Deprecated: use top-level retry instead.'
        timeout:
          $ref: '#/components/schemas/EmbeddingTimeoutConfig'
          description: 'Deprecated: use top-level timeout instead.'
      type: object
    EmbeddingRetryConfig:
      additionalProperties: false
      properties:
        count:
          description: Number of retry attempts (1-5).
          format: int64
          type: integer
        on_codes:
          description: HTTP status codes that trigger retry logic.
          items:
            format: int64
            type: integer
          type:
            - array
            - 'null'
      required:
        - count
        - on_codes
      type: object
    EmbeddingTimeoutConfig:
      additionalProperties: false
      properties:
        call_timeout:
          description: Timeout value in milliseconds.
          format: int64
          type: integer
      required:
        - call_timeout
      type: object
    PublicEmbeddingData:
      additionalProperties: false
      properties:
        embedding:
          description: >-
            The embedding vector, which is a list of floats. The length of
            vector depends on the model. Can also be a base64-encoded string
            when encoding_format is base64.
        index:
          description: The index of the embedding in the list of embeddings.
          format: int64
          type: integer
        object:
          description: The object type, which is always "embedding".
          enum:
            - embedding
          type: string
      required:
        - object
        - index
        - embedding
      type: object
    PublicEmbeddingUsage:
      additionalProperties: false
      properties:
        prompt_tokens:
          description: The number of tokens used by the prompt.
          format: int64
          type: integer
        total_tokens:
          description: The total number of tokens used by the request.
          format: int64
          type: integer
      required:
        - prompt_tokens
        - total_tokens
      type: object
    EmbeddingLoadBalancerModelConfig:
      additionalProperties: false
      properties:
        model:
          description: Model identifier for load balancing.
          type: string
        weight:
          description: Weight assigned to this model for load balancing.
          format: double
          type: number
      required:
        - model
        - weight
      type: object
    EmbeddingContactParams:
      additionalProperties: false
      properties:
        display_name:
          type: string
        email:
          type: string
        id:
          type: string
        metadata:
          items:
            type: object
            additionalProperties: {}
          type:
            - array
            - 'null'
        tags:
          items:
            type: string
          type:
            - array
            - 'null'
      required:
        - id
      type: object
    ResponseIdentity:
      additionalProperties: false
      properties:
        display_name:
          type: string
        email:
          type: string
        id:
          type: string
        metadata:
          items:
            type: object
            additionalProperties: {}
          type:
            - array
            - 'null'
        tags:
          items:
            type: string
          type:
            - array
            - 'null'
      required:
        - id
      type: object
  securitySchemes:
    ApiKey:
      type: http
      scheme: bearer
      bearerFormat: JWT

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.