> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Create ocr

> Extracts text content while maintaining document structure and hierarchy

<Note>
  **Related guide**: OCR guide. See the [OCR guide](/ai-gateway/features/ocr) for a walkthrough with examples.
</Note>


## OpenAPI

````yaml post /v2/router/ocr
openapi: 3.1.0
info:
  title: orq.ai API
  version: '2.0'
  description: orq.ai API documentation
servers:
  - url: https://my.orq.ai
security:
  - ApiKey: []
tags:
  - name: Chunking
    description: Split text into smaller chunks for retrieval and generation workflows.
  - name: File Systems
    description: >-
      Create and manage persistent file systems that agents and MCP clients read
      from and write to.
  - name: Knowledge Bases
    description: Create and manage knowledge bases used by agents and retrieval workflows.
  - name: Memory Stores
    description: Create and manage memory stores, memories, and memory documents.
  - name: Evals
    description: Run an evaluator against a conversation and its result
  - name: Logs
    description: >-
      OpenTelemetry log query API. Search, filter, aggregate, and facet log
      records ingested via OTLP.
  - name: Reporting
    description: >-
      GenAI reporting API over canonical analytics rollups. Accepts a metric
      name, time range, grain, group-by, and filters; returns a typed time
      series and optional totals.
  - name: Traces
    description: >-
      Query and inspect ingested trace data: search trace summaries, aggregate
      metrics, and read individual traces and their spans.
  - description: List models available through the AI Router.
    name: Models
  - name: Policies
  - name: Alerts
    description: >-
      Alerts evaluate a Reporting API metric on a fixed interval and fire
      notifications through notifiers when the value breaches a threshold. Each
      breach opens a trigger that tracks the incident until the value recovers.
  - name: Annotation Queues
    description: Annotation queues collect spans for human review.
  - name: API keys
    description: >-
      API keys authenticate programmatic access to the workspace. They expose
      opaque tokens, per-domain access grants, and budget and rate-limit
      constraints.
  - name: Audit Logs
    description: Audit logs record workspace entity changes and access-relevant events.
  - name: Budgets
    description: >-
      Budgets govern spend, token usage, and request rate across six scopes:
      workspace, project, identity, API key, provider, and model. Every
      applicable budget is enforced, and the most restrictive limit applies per
      dimension.
  - name: Files
    description: File upload and retrieval operations.
  - name: Guardrail Rules
    description: >-
      Guardrail Rules conditionally enforce evaluators and plugins for AI
      Gateway traffic. Rules may be scoped to a project or the whole workspace.
  - name: Hub
    description: Hub items are reusable templates available to a workspace.
  - name: Identities
    description: >-
      Identities represent end users from your system for usage and engagement
      tracking.
  - name: Management keys
    description: >-
      Management keys are workspace-scoped credentials that authenticate
      programmatic access to workspace administration surfaces (API keys,
      budgets). Unlike project-scoped API keys, a management key always operates
      at the workspace level.
  - name: MCP Gateway
    description: >-
      Register upstream MCP servers, discover and sync their tools, and assemble
      gateways that expose a curated tool surface to MCP clients.
  - name: Model Catalog
    description: >-
      Browse the orq.ai model catalog: every model orq offers, across every
      provider, with pricing, capabilities and benchmark data. List endpoints
      only return models that are not deprecated. This API is public, requires
      no authentication, and is rate limited to 120 requests per minute per IP.
      Responses carry a 5-minute cache-control max-age.
  - name: Notifiers
    description: Notifier destinations used to send delivery and workflow notifications.
  - name: Projects
    description: Projects organize resources within a workspace
  - name: Routing Rules
    description: >-
      Routing Rules conditionally select models and enforce request plugins for
      AI Gateway traffic. Rules are evaluated by ascending priority and may be
      scoped to a project or the whole workspace.
  - name: Threads
    description: Threads group related trace invocations and their aggregate usage
  - name: Skills
    description: >-
      Skills are modular instructions you can use to codify processes and
      conventions
  - name: Smart Routers
    description: >-
      Create and manage workspace Smart Routers. A Smart Router selects a model
      from an eligible pool for each request according to a quality, balanced,
      or cost profile.
  - name: Webhooks
    description: >-
      Create and manage webhooks that deliver workspace events to external HTTPS
      endpoints.
  - name: Workspaces
    description: >-
      A workspace is the tenant. Create is called from a user session during
      onboarding; Get, List, and Update are the public management surface.
  - name: Workspace Security
    description: >-
      Workspace-level domain verification and IP allowlist controls. These
      operations are restricted to workspace administrators.
  - name: Workspace Settings
    description: >-
      Workspace-level settings managed with a workspace credential. A workspace
      is the tenant, so these settings are a singleton — there is nothing to
      create or delete, only read and update.
  - name: Responses
  - description: Run agents on a cron cadence. Minimum firing interval is 1 hour.
    name: Agent Schedules
  - name: Embeddings
  - name: Telemetry
    description: >-
      Unified query envelope for traces, metrics, and logs. One request shape,
      one filter dialect, and one response shape per source, validated by a
      per-source registry.
  - description: Beta. Run typed classification questions against a classify model.
    name: Classify
  - description: Search Gateway with managed credits or BYOK.
    name: Web Search
externalDocs:
  url: https://docs.orq.ai
  description: orq.ai Documentation
paths:
  /v2/router/ocr:
    post:
      tags:
        - Router
      description: Extracts text content while maintaining document structure and hierarchy
      requestBody:
        required: true
        description: input
        content:
          application/json:
            schema:
              type: object
              properties:
                model:
                  type: string
                  description: ID of the model to use for OCR.
                document:
                  anyOf:
                    - type: object
                      properties:
                        type:
                          type: string
                          enum:
                            - document_url
                        document_url:
                          type: string
                          format: uri
                          description: URL of the document to process
                        document_name:
                          type: string
                          description: The name of the document
                      required:
                        - type
                        - document_url
                    - type: object
                      properties:
                        type:
                          type: string
                          enum:
                            - image_url
                        image_url:
                          anyOf:
                            - type: string
                              description: Base64 encoded image
                            - type: object
                              properties:
                                url:
                                  type: string
                                  format: uri
                                detail:
                                  type: string
                              required:
                                - url
                              description: URL of the image to process
                      required:
                        - type
                        - image_url
                  description: >-
                    Document to run OCR on. Can be a DocumentURLChunk or
                    ImageURLChunk.
                pages:
                  type:
                    - array
                    - 'null'
                  items:
                    type: integer
                  description: >-
                    Specific pages to process. Can be a single number, range, or
                    list. Starts from 0. Null for all pages.
                ocr_settings:
                  type: object
                  properties:
                    include_image_base64:
                      type:
                        - boolean
                        - 'null'
                      description: >-
                        Whether to include image Base64 in the response. Null
                        for default.
                    max_images_to_include:
                      type: integer
                      description: Maximum number of images to extract. Null for no limit.
                    image_min_size:
                      type: integer
                      description: >-
                        Minimum height and width of image to extract. Null for
                        no minimum.
                  description: Optional settings for the OCR run
              required:
                - model
                - document
      responses:
        '200':
          description: Represents an OCR response from the API.
          content:
            application/json:
              schema:
                type: object
                properties:
                  model:
                    type: string
                    description: ID of the model used for OCR.
                  pages:
                    type: array
                    items:
                      type: object
                      properties:
                        index:
                          type: number
                          description: The page index in a pdf document starting from 0
                        markdown:
                          type: string
                          description: The markdown string response of the page
                        images:
                          type: array
                          items:
                            type: object
                            properties:
                              id:
                                type: string
                                description: The id of the image
                              image_base64:
                                type:
                                  - string
                                  - 'null'
                                description: The base64 encoded image
                            required:
                              - id
                        dimensions:
                          type:
                            - object
                            - 'null'
                          properties:
                            dpi:
                              type: integer
                              description: Dots per inch of the page-image
                            height:
                              type: integer
                              description: Height of the image in pixels
                            width:
                              type: integer
                              description: Width of the image in pixels
                          required:
                            - dpi
                            - height
                            - width
                          description: The dimensions of the PDF Page's screenshot image
                      required:
                        - index
                        - markdown
                        - images
                  usage:
                    oneOf:
                      - type: object
                        properties:
                          type:
                            type: string
                            enum:
                              - pages
                          pages_processed:
                            type: integer
                            description: The number of pages processed
                          input_cost:
                            type: number
                            description: >-
                              Cost (USD) attributed to input processing. Present
                              when billing was computed for this response.
                          output_cost:
                            type: number
                            description: >-
                              Cost (USD) attributed to output processing.
                              Present when billing was computed for this
                              response.
                          total_cost:
                            type: number
                            description: >-
                              Total cost (USD) of the response. Present when
                              billing was computed for this response.
                        required:
                          - type
                          - pages_processed
                        description: >-
                          The usage information for the OCR run counted as pages
                          processed
                      - type: object
                        properties:
                          type:
                            type: string
                            enum:
                              - tokens
                          tokens_processed:
                            type: integer
                            description: The number of tokens processed
                          input_cost:
                            type: number
                            description: >-
                              Cost (USD) attributed to input processing. Present
                              when billing was computed for this response.
                          output_cost:
                            type: number
                            description: >-
                              Cost (USD) attributed to output processing.
                              Present when billing was computed for this
                              response.
                          total_cost:
                            type: number
                            description: >-
                              Total cost (USD) of the response. Present when
                              billing was computed for this response.
                        required:
                          - type
                          - tokens_processed
                        description: >-
                          The usage information for the OCR run counted as
                          tokens processed
                required:
                  - model
                  - pages
                  - usage
components:
  securitySchemes:
    ApiKey:
      type: http
      scheme: bearer
      bearerFormat: JWT

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.