> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Create agent

> Create a new agent with the specified model, instructions, tools, and knowledge bases. Supports fallback models and configurable execution settings.

<Note>
  **Related guide**: Build agents guide. See the [Build agents guide](/ai-studio/ai-engineering/build-agents) for a walkthrough with examples.
</Note>


## OpenAPI

````yaml post /v2/agents
openapi: 3.1.0
info:
  title: orq.ai API
  version: '2.0'
  description: orq.ai API documentation
servers:
  - url: https://my.orq.ai
security:
  - ApiKey: []
tags:
  - name: Chunking
    description: Split text into smaller chunks for retrieval and generation workflows.
  - name: File Systems
    description: >-
      Create and manage persistent file systems that agents and MCP clients read
      from and write to.
  - name: Knowledge Bases
    description: Create and manage knowledge bases used by agents and retrieval workflows.
  - name: Memory Stores
    description: Create and manage memory stores, memories, and memory documents.
  - name: Evals
    description: Run an evaluator against a conversation and its result
  - name: Logs
    description: >-
      OpenTelemetry log query API. Search, filter, aggregate, and facet log
      records ingested via OTLP.
  - name: Reporting
    description: >-
      GenAI reporting API over canonical analytics rollups. Accepts a metric
      name, time range, grain, group-by, and filters; returns a typed time
      series and optional totals.
  - name: Traces
    description: >-
      Query and inspect ingested trace data: search trace summaries, aggregate
      metrics, and read individual traces and their spans.
  - description: List models available through the AI Router.
    name: Models
  - name: Policies
  - name: Alerts
    description: >-
      Alerts evaluate a Reporting API metric on a fixed interval and fire
      notifications through notifiers when the value breaches a threshold. Each
      breach opens a trigger that tracks the incident until the value recovers.
  - name: Annotation Queues
    description: Annotation queues collect spans for human review.
  - name: API keys
    description: >-
      API keys authenticate programmatic access to the workspace. They expose
      opaque tokens, per-domain access grants, and budget and rate-limit
      constraints.
  - name: Audit Logs
    description: Audit logs record workspace entity changes and access-relevant events.
  - name: Budgets
    description: >-
      Budgets govern spend, token usage, and request rate across six scopes:
      workspace, project, identity, API key, provider, and model. Every
      applicable budget is enforced, and the most restrictive limit applies per
      dimension.
  - name: Files
    description: File upload and retrieval operations.
  - name: Guardrail Rules
    description: >-
      Guardrail Rules conditionally enforce evaluators and plugins for AI
      Gateway traffic. Rules may be scoped to a project or the whole workspace.
  - name: Hub
    description: Hub items are reusable templates available to a workspace.
  - name: Identities
    description: >-
      Identities represent end users from your system for usage and engagement
      tracking.
  - name: Management keys
    description: >-
      Management keys are workspace-scoped credentials that authenticate
      programmatic access to workspace administration surfaces (API keys,
      budgets). Unlike project-scoped API keys, a management key always operates
      at the workspace level.
  - name: MCP Gateway
    description: >-
      Register upstream MCP servers, discover and sync their tools, and assemble
      gateways that expose a curated tool surface to MCP clients.
  - name: Model Catalog
    description: >-
      Browse the orq.ai model catalog: every model orq offers, across every
      provider, with pricing, capabilities and benchmark data. List endpoints
      only return models that are not deprecated. This API is public, requires
      no authentication, and is rate limited to 120 requests per minute per IP.
      Responses carry a 5-minute cache-control max-age.
  - name: Notifiers
    description: Notifier destinations used to send delivery and workflow notifications.
  - name: Projects
    description: Projects organize resources within a workspace
  - name: Routing Rules
    description: >-
      Routing Rules conditionally select models and enforce request plugins for
      AI Gateway traffic. Rules are evaluated by ascending priority and may be
      scoped to a project or the whole workspace.
  - name: Threads
    description: Threads group related trace invocations and their aggregate usage
  - name: Skills
    description: >-
      Skills are modular instructions you can use to codify processes and
      conventions
  - name: Smart Routers
    description: >-
      Create and manage workspace Smart Routers. A Smart Router selects a model
      from an eligible pool for each request according to a quality, balanced,
      or cost profile.
  - name: Webhooks
    description: >-
      Create and manage webhooks that deliver workspace events to external HTTPS
      endpoints.
  - name: Workspaces
    description: >-
      A workspace is the tenant. Create is called from a user session during
      onboarding; Get, List, and Update are the public management surface.
  - name: Workspace Security
    description: >-
      Workspace-level domain verification and IP allowlist controls. These
      operations are restricted to workspace administrators.
  - name: Workspace Settings
    description: >-
      Workspace-level settings managed with a workspace credential. A workspace
      is the tenant, so these settings are a singleton — there is nothing to
      create or delete, only read and update.
  - name: Responses
  - description: Run agents on a cron cadence. Minimum firing interval is 1 hour.
    name: Agent Schedules
  - name: Embeddings
  - name: Telemetry
    description: >-
      Unified query envelope for traces, metrics, and logs. One request shape,
      one filter dialect, and one response shape per source, validated by a
      per-source registry.
  - description: Beta. Run typed classification questions against a classify model.
    name: Classify
  - description: Search Gateway with managed credits or BYOK.
    name: Web Search
externalDocs:
  url: https://docs.orq.ai
  description: orq.ai Documentation
paths:
  /v2/agents:
    post:
      tags:
        - Agents
      summary: Create agent
      description: >-
        Create a new agent with the specified model, instructions, tools, and
        knowledge bases. Supports fallback models and configurable execution
        settings.
      operationId: CreateAgentRequest
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                key:
                  type: string
                  minLength: 1
                  maxLength: 255
                  pattern: ^[A-Za-z][A-Za-z0-9]*([._-][A-Za-z0-9]+)*$
                  description: Unique identifier for the agent within the workspace
                display_name:
                  type: string
                  maxLength: 255
                  description: agent display name within the workspace
                role:
                  type: string
                  minLength: 1
                  description: The role or function of the agent
                description:
                  type: string
                  minLength: 1
                  description: A brief description of what the agent does
                instructions:
                  type: string
                  description: Detailed instructions that guide the agent's behavior
                system_prompt:
                  type:
                    - string
                    - 'null'
                  minLength: 1
                  description: >-
                    A custom system prompt template for the agent. If omitted,
                    the default template is used.
                path:
                  type: string
                  description: >-
                    The path where the agent will be stored in the project
                    structure. The first element identifies the project,
                    followed by nested folders (auto-created as needed).


                    With project-based API keys, the first element is treated as
                    a folder name, as the project is predetermined by the API
                    key.
                  example: Default Project
                model:
                  anyOf:
                    - type: string
                      description: >-


                        A model ID string (e.g., `openai/gpt-5.6-sol` or
                        `anthropic/claude-sonnet-5`). The agent can be run with
                        a wide range of models with different capabilities,
                        performance characteristics, and price points. Only
                        models that support tool calling (function_calling) can
                        be used to run agents. See [supported
                        models](/ai-gateway/supported-models) documentation for
                        the complete list of available models.
                    - type: object
                      properties:
                        id:
                          type: string
                          description: >-
                            A model ID string (e.g., `openai/gpt-5.6-sol` or
                            `anthropic/claude-sonnet-5`). Only models that
                            support tool calling can be used with agents.
                        parameters:
                          type: object
                          properties:
                            name:
                              description: >-
                                The name to display on the trace. If not
                                specified, the default system name will be used.
                              type: string
                            frequency_penalty:
                              type:
                                - number
                                - 'null'
                              description: >-
                                Number between -2.0 and 2.0. Positive values
                                penalize new tokens based on their existing
                                frequency in the text so far, decreasing the
                                model's likelihood to repeat the same line
                                verbatim.
                            max_tokens:
                              type:
                                - integer
                                - 'null'
                              description: >-
                                `[Deprecated]`. The maximum number of tokens
                                that can be generated in the chat completion.
                                This value can be used to control costs for text
                                generated via API. 

                                 This value is now `deprecated` in favor of `max_completion_tokens`, and is not compatible with o1 series models.
                            max_completion_tokens:
                              type:
                                - integer
                                - 'null'
                              exclusiveMinimum: 0
                              description: >-
                                An upper bound for the number of tokens that can
                                be generated for a completion, including visible
                                output tokens and reasoning tokens
                            presence_penalty:
                              type:
                                - number
                                - 'null'
                              description: >-
                                Number between -2.0 and 2.0. Positive values
                                penalize new tokens based on whether they appear
                                in the text so far, increasing the model's
                                likelihood to talk about new topics.
                            response_format:
                              oneOf:
                                - type: object
                                  properties:
                                    type:
                                      type: string
                                      enum:
                                        - text
                                  required:
                                    - type
                                  title: Text
                                  description: >-


                                    Default response format. Used to generate
                                    text responses
                                - type: object
                                  properties:
                                    type:
                                      type: string
                                      enum:
                                        - json_object
                                  required:
                                    - type
                                  title: JSON object
                                  description: >-


                                    JSON object response format. An older method
                                    of generating JSON responses. Using
                                    `json_schema` is recommended for models that
                                    support it. Note that the model will not
                                    generate JSON without a system or user
                                    message instructing it to do so.
                                - type: object
                                  properties:
                                    type:
                                      type: string
                                      enum:
                                        - json_schema
                                    json_schema:
                                      type: object
                                      properties:
                                        description:
                                          description: >-
                                            A description of what the response
                                            format is for, used by the model to
                                            determine how to respond in the format.
                                          type: string
                                        name:
                                          type: string
                                          description: >-
                                            The name of the response format. Must be
                                            a-z, A-Z, 0-9, or contain underscores
                                            and dashes, with a maximum length of 64.
                                        schema:
                                          description: >-
                                            The schema for the response format,
                                            described as a JSON Schema object.
                                        strict:
                                          type: boolean
                                          default: false
                                          description: >-
                                            Whether to enable strict schema
                                            adherence when generating the output. If
                                            set to true, the model will always
                                            follow the exact schema defined in the
                                            schema field. Only a subset of JSON
                                            Schema is supported when strict is true.
                                      required:
                                        - name
                                  required:
                                    - type
                                    - json_schema
                                  title: JSON schema
                                  description: >-


                                    JSON Schema response format. Used to
                                    generate structured JSON responses
                              description: >-
                                An object specifying the format that the model
                                must output
                            reasoning_effort:
                              type: string
                              enum:
                                - none
                                - minimal
                                - low
                                - medium
                                - high
                                - xhigh
                                - max
                              description: >-
                                Constrains effort on reasoning for [reasoning
                                models](https://platform.openai.com/docs/guides/reasoning).
                                Currently supported values are `none`,
                                `minimal`, `low`, `medium`, `high`, `xhigh`, and
                                `max`. Reducing reasoning effort can result in
                                faster responses and fewer tokens used on
                                reasoning in a response.


                                - `gpt-5.1` defaults to `none`, which does not
                                perform reasoning. The supported reasoning
                                values for `gpt-5.1` are `none`, `low`,
                                `medium`, and `high`. Tool calls are supported
                                for all reasoning values in gpt-5.1.

                                - All models before `gpt-5.1` default to
                                `medium` reasoning effort, and do not support
                                `none`.

                                - The `gpt-5-pro` model defaults to (and only
                                supports) `high` reasoning effort.

                                - `xhigh` is currently only supported for
                                `gpt-5.1-codex-max`.


                                Any of "none", "minimal", "low", "medium",
                                "high", "xhigh", "max".
                            verbosity:
                              type: string
                              description: >-
                                Adjusts response verbosity. Lower levels yield
                                shorter answers.
                            seed:
                              type:
                                - number
                                - 'null'
                              description: >-
                                If specified, our system will make a best effort
                                to sample deterministically, such that repeated
                                requests with the same seed and parameters
                                should return the same result.
                            stop:
                              anyOf:
                                - type: string
                                - type: array
                                  items:
                                    type: string
                                  maxItems: 4
                                - type: 'null'
                              description: >-
                                Up to 4 sequences where the API will stop
                                generating further tokens.
                            thinking:
                              oneOf:
                                - $ref: >-
                                    #/components/schemas/ThinkingConfigDisabledSchema
                                - $ref: >-
                                    #/components/schemas/ThinkingConfigEnabledSchema
                                - $ref: >-
                                    #/components/schemas/ThinkingConfigAdaptiveSchema
                              discriminator:
                                propertyName: type
                                mapping:
                                  disabled: >-
                                    #/components/schemas/ThinkingConfigDisabledSchema
                                  enabled: >-
                                    #/components/schemas/ThinkingConfigEnabledSchema
                                  adaptive: >-
                                    #/components/schemas/ThinkingConfigAdaptiveSchema
                            temperature:
                              type:
                                - number
                                - 'null'
                              minimum: 0
                              maximum: 2
                              description: >-
                                What sampling temperature to use, between 0 and
                                2. Higher values like 0.8 will make the output
                                more random, while lower values like 0.2 will
                                make it more focused and deterministic.
                            top_p:
                              type:
                                - number
                                - 'null'
                              minimum: 0
                              maximum: 1
                              description: >-
                                An alternative to sampling with temperature,
                                called nucleus sampling, where the model
                                considers the results of the tokens with top_p
                                probability mass. 
                            top_k:
                              type:
                                - number
                                - 'null'
                              description: >-
                                Limits the model to consider only the top k most
                                likely tokens at each step.
                            tool_choice:
                              anyOf:
                                - type: string
                                  enum:
                                    - none
                                    - auto
                                    - required
                                - type: object
                                  properties:
                                    type:
                                      type: string
                                      enum:
                                        - function
                                      description: >-
                                        The type of the tool. Currently, only
                                        function is supported.
                                    function:
                                      type: object
                                      properties:
                                        name:
                                          type: string
                                          description: The name of the function to call.
                                      required:
                                        - name
                                  required:
                                    - function
                              description: >-
                                Controls which (if any) tool is called by the
                                model.
                            parallel_tool_calls:
                              type: boolean
                              description: >-
                                Whether to enable parallel function calling
                                during tool use.
                            modalities:
                              type:
                                - array
                                - 'null'
                              items:
                                type: string
                                enum:
                                  - text
                                  - audio
                              description: >-
                                Output types that you would like the model to
                                generate. Most models are capable of generating
                                text, which is the default: ["text"]. The
                                gpt-4o-audio-preview model can also be used to
                                generate audio. To request that this model
                                generate both text and audio responses, you can
                                use: ["text", "audio"].
                            guardrails:
                              type: array
                              items:
                                type: object
                                properties:
                                  id:
                                    anyOf:
                                      - type: string
                                        enum:
                                          - orq_pii_detection
                                          - orq_secret_detection
                                          - orq_sexual_moderation
                                          - orq_harmful_moderation
                                        description: The key of the guardrail.
                                      - type: string
                                        description: >-
                                          Unique key or identifier of the
                                          evaluator
                                  execute_on:
                                    type: string
                                    enum:
                                      - input
                                      - output
                                    description: >-
                                      Determines whether the guardrail runs on
                                      the input (user message) or output (model
                                      response).
                                required:
                                  - id
                                  - execute_on
                              description: A list of guardrails to apply to the request.
                            plugins:
                              type: array
                              items:
                                anyOf:
                                  - $ref: '#/components/schemas/PIIRedactionPlugin'
                                  - $ref: '#/components/schemas/ResponseHealingPlugin'
                                  - $ref: '#/components/schemas/TraceScrubbingPlugin'
                              description: >-
                                Request-scoped transforms applied to the text
                                exchanged with the model. Supports
                                `pii_redaction`, which replaces PII with
                                placeholders before the provider sees it and
                                restores the original values in the response,
                                and `response_healing`, which repairs malformed
                                JSON in non-streaming output.
                            fallbacks:
                              type: array
                              items:
                                type: object
                                properties:
                                  model:
                                    type: string
                                    description: Fallback model identifier
                                    example: openai/gpt-5.4-mini
                                required:
                                  - model
                              description: >-
                                Array of fallback models to use if primary model
                                fails
                            cache:
                              type: object
                              properties:
                                ttl:
                                  type: number
                                  minimum: 1
                                  maximum: 259200
                                  default: 1800
                                  description: >-
                                    Time to live for cached responses in
                                    seconds. Maximum 259200 seconds (3 days).
                                  example: 3600
                                type:
                                  type: string
                                  enum:
                                    - exact_match
                              required:
                                - type
                              description: Cache configuration for the request.
                            load_balancer:
                              oneOf:
                                - type: object
                                  properties:
                                    type:
                                      type: string
                                      enum:
                                        - weight_based
                                    models:
                                      type: array
                                      items:
                                        type: object
                                        properties:
                                          model:
                                            type: string
                                            description: Model identifier for load balancing
                                            example: openai/gpt-5.6-sol
                                          weight:
                                            type: number
                                            minimum: 0.001
                                            maximum: 1
                                            default: 0.5
                                            description: >-
                                              Weight assigned to this model for load
                                              balancing
                                            example: 0.7
                                        required:
                                          - model
                                  required:
                                    - type
                                    - models
                              description: Load balancer configuration for the request.
                              example:
                                type: weight_based
                                models:
                                  - model: openai/gpt-4o
                                    weight: 0.7
                                  - model: anthropic/claude-3-5-sonnet
                                    weight: 0.3
                            timeout:
                              type: object
                              properties:
                                call_timeout:
                                  type: number
                                  minimum: 1
                                  description: Timeout value in milliseconds
                                  example: 30000
                              required:
                                - call_timeout
                              description: >-
                                Timeout configuration to apply to the request.
                                If the request exceeds the timeout, it will be
                                retried or fallback to the next model if
                                configured.
                            cache_control:
                              type: object
                              properties:
                                type:
                                  type: string
                                  enum:
                                    - ephemeral
                                  description: >-
                                    Create a cache control breakpoint at this
                                    content block. Accepts only the value
                                    "ephemeral".
                                ttl:
                                  type: string
                                  enum:
                                    - 5m
                                    - 1h
                                  default: 5m
                                  description: >-
                                    The time-to-live for the cache control
                                    breakpoint. This may be one of the following
                                    values:


                                    - `5m`: 5 minutes

                                    - `1h`: 1 hour


                                    Defaults to `5m`. Only supported by
                                    `Anthropic` Claude models.
                              required:
                                - type
                              description: >-
                                Provider-level prompt caching configuration
                                applied to the request. Creates a cache control
                                breakpoint covering the request content. Only
                                supported by `Anthropic` Claude models.
                            prompt_cache_key:
                              type: string
                              description: >-
                                Used by OpenAI to cache responses for similar
                                requests to optimize your cache hit rates.
                                Replaces the legacy `user` field for prompt
                                caching.
                          description: >-
                            Model behavior parameters that control how the model
                            generates responses. Common parameters:
                            `temperature` (0-2, randomness; the selected model
                            may impose a lower maximum), `max_completion_tokens`
                            (max output length), `top_p` (sampling diversity).
                            Advanced: `frequency_penalty`, `presence_penalty`,
                            `response_format` (JSON/structured),
                            `reasoning_effort`, `seed` (reproducibility).
                            Support varies by model - consult AI Gateway
                            documentation.
                        retry:
                          type: object
                          properties:
                            count:
                              type: number
                              minimum: 1
                              maximum: 5
                              default: 3
                              description: Number of retry attempts (1-5)
                              example: 3
                            on_codes:
                              type: array
                              items:
                                type: number
                                minimum: 100
                                maximum: 599
                              minItems: 1
                              description: HTTP status codes that trigger retry logic
                              example:
                                - 429
                                - 500
                                - 502
                                - 503
                                - 504
                          description: >-
                            Retry configuration for model requests. Retries are
                            triggered for specific HTTP status codes (e.g., 500,
                            429, 502, 503, 504). Supports configurable retry
                            count (1-5) and custom status codes.
                      required:
                        - id
                      description: |-


                        Model configuration with parameters and retry settings.
                  title: Model Configuration
                  description: >-
                    Model configuration for agent execution. Can be a simple
                    model ID string or a configuration object with optional
                    behavior parameters and retry settings.
                fallback_models:
                  type: array
                  items:
                    anyOf:
                      - type: string
                        description: >-
                          A fallback model ID string (e.g.,
                          `openai/gpt-4o-mini`). Will be used if the primary
                          model request fails. Must support tool calling.
                      - type: object
                        properties:
                          id:
                            type: string
                            description: >-
                              A fallback model ID string. Must support tool
                              calling.
                          parameters:
                            type: object
                            properties:
                              name:
                                description: >-
                                  The name to display on the trace. If not
                                  specified, the default system name will be
                                  used.
                                type: string
                              frequency_penalty:
                                type:
                                  - number
                                  - 'null'
                                description: >-
                                  Number between -2.0 and 2.0. Positive values
                                  penalize new tokens based on their existing
                                  frequency in the text so far, decreasing the
                                  model's likelihood to repeat the same line
                                  verbatim.
                              max_tokens:
                                type:
                                  - integer
                                  - 'null'
                                description: >-
                                  `[Deprecated]`. The maximum number of tokens
                                  that can be generated in the chat completion.
                                  This value can be used to control costs for
                                  text generated via API. 

                                   This value is now `deprecated` in favor of `max_completion_tokens`, and is not compatible with o1 series models.
                              max_completion_tokens:
                                type:
                                  - integer
                                  - 'null'
                                exclusiveMinimum: 0
                                description: >-
                                  An upper bound for the number of tokens that
                                  can be generated for a completion, including
                                  visible output tokens and reasoning tokens
                              presence_penalty:
                                type:
                                  - number
                                  - 'null'
                                description: >-
                                  Number between -2.0 and 2.0. Positive values
                                  penalize new tokens based on whether they
                                  appear in the text so far, increasing the
                                  model's likelihood to talk about new topics.
                              response_format:
                                oneOf:
                                  - type: object
                                    properties:
                                      type:
                                        type: string
                                        enum:
                                          - text
                                    required:
                                      - type
                                    title: Text
                                    description: >-


                                      Default response format. Used to generate
                                      text responses
                                  - type: object
                                    properties:
                                      type:
                                        type: string
                                        enum:
                                          - json_object
                                    required:
                                      - type
                                    title: JSON object
                                    description: >-


                                      JSON object response format. An older
                                      method of generating JSON responses. Using
                                      `json_schema` is recommended for models
                                      that support it. Note that the model will
                                      not generate JSON without a system or user
                                      message instructing it to do so.
                                  - type: object
                                    properties:
                                      type:
                                        type: string
                                        enum:
                                          - json_schema
                                      json_schema:
                                        type: object
                                        properties:
                                          description:
                                            description: >-
                                              A description of what the response
                                              format is for, used by the model to
                                              determine how to respond in the format.
                                            type: string
                                          name:
                                            type: string
                                            description: >-
                                              The name of the response format. Must be
                                              a-z, A-Z, 0-9, or contain underscores
                                              and dashes, with a maximum length of 64.
                                          schema:
                                            description: >-
                                              The schema for the response format,
                                              described as a JSON Schema object.
                                          strict:
                                            type: boolean
                                            default: false
                                            description: >-
                                              Whether to enable strict schema
                                              adherence when generating the output. If
                                              set to true, the model will always
                                              follow the exact schema defined in the
                                              schema field. Only a subset of JSON
                                              Schema is supported when strict is true.
                                        required:
                                          - name
                                    required:
                                      - type
                                      - json_schema
                                    title: JSON schema
                                    description: >-


                                      JSON Schema response format. Used to
                                      generate structured JSON responses
                                description: >-
                                  An object specifying the format that the model
                                  must output
                              reasoning_effort:
                                type: string
                                enum:
                                  - none
                                  - minimal
                                  - low
                                  - medium
                                  - high
                                  - xhigh
                                  - max
                                description: >-
                                  Constrains effort on reasoning for [reasoning
                                  models](https://platform.openai.com/docs/guides/reasoning).
                                  Currently supported values are `none`,
                                  `minimal`, `low`, `medium`, `high`, `xhigh`,
                                  and `max`. Reducing reasoning effort can
                                  result in faster responses and fewer tokens
                                  used on reasoning in a response.


                                  - `gpt-5.1` defaults to `none`, which does not
                                  perform reasoning. The supported reasoning
                                  values for `gpt-5.1` are `none`, `low`,
                                  `medium`, and `high`. Tool calls are supported
                                  for all reasoning values in gpt-5.1.

                                  - All models before `gpt-5.1` default to
                                  `medium` reasoning effort, and do not support
                                  `none`.

                                  - The `gpt-5-pro` model defaults to (and only
                                  supports) `high` reasoning effort.

                                  - `xhigh` is currently only supported for
                                  `gpt-5.1-codex-max`.


                                  Any of "none", "minimal", "low", "medium",
                                  "high", "xhigh", "max".
                              verbosity:
                                type: string
                                description: >-
                                  Adjusts response verbosity. Lower levels yield
                                  shorter answers.
                              seed:
                                type:
                                  - number
                                  - 'null'
                                description: >-
                                  If specified, our system will make a best
                                  effort to sample deterministically, such that
                                  repeated requests with the same seed and
                                  parameters should return the same result.
                              stop:
                                anyOf:
                                  - type: string
                                  - type: array
                                    items:
                                      type: string
                                    maxItems: 4
                                  - type: 'null'
                                description: >-
                                  Up to 4 sequences where the API will stop
                                  generating further tokens.
                              thinking:
                                oneOf:
                                  - $ref: >-
                                      #/components/schemas/ThinkingConfigDisabledSchema
                                  - $ref: >-
                                      #/components/schemas/ThinkingConfigEnabledSchema
                                  - $ref: >-
                                      #/components/schemas/ThinkingConfigAdaptiveSchema
                                discriminator:
                                  propertyName: type
                                  mapping:
                                    disabled: >-
                                      #/components/schemas/ThinkingConfigDisabledSchema
                                    enabled: >-
                                      #/components/schemas/ThinkingConfigEnabledSchema
                                    adaptive: >-
                                      #/components/schemas/ThinkingConfigAdaptiveSchema
                              temperature:
                                type:
                                  - number
                                  - 'null'
                                minimum: 0
                                maximum: 2
                                description: >-
                                  What sampling temperature to use, between 0
                                  and 2. Higher values like 0.8 will make the
                                  output more random, while lower values like
                                  0.2 will make it more focused and
                                  deterministic.
                              top_p:
                                type:
                                  - number
                                  - 'null'
                                minimum: 0
                                maximum: 1
                                description: >-
                                  An alternative to sampling with temperature,
                                  called nucleus sampling, where the model
                                  considers the results of the tokens with top_p
                                  probability mass. 
                              top_k:
                                type:
                                  - number
                                  - 'null'
                                description: >-
                                  Limits the model to consider only the top k
                                  most likely tokens at each step.
                              tool_choice:
                                anyOf:
                                  - type: string
                                    enum:
                                      - none
                                      - auto
                                      - required
                                  - type: object
                                    properties:
                                      type:
                                        type: string
                                        enum:
                                          - function
                                        description: >-
                                          The type of the tool. Currently, only
                                          function is supported.
                                      function:
                                        type: object
                                        properties:
                                          name:
                                            type: string
                                            description: The name of the function to call.
                                        required:
                                          - name
                                    required:
                                      - function
                                description: >-
                                  Controls which (if any) tool is called by the
                                  model.
                              parallel_tool_calls:
                                type: boolean
                                description: >-
                                  Whether to enable parallel function calling
                                  during tool use.
                              modalities:
                                type:
                                  - array
                                  - 'null'
                                items:
                                  type: string
                                  enum:
                                    - text
                                    - audio
                                description: >-
                                  Output types that you would like the model to
                                  generate. Most models are capable of
                                  generating text, which is the default:
                                  ["text"]. The gpt-4o-audio-preview model can
                                  also be used to generate audio. To request
                                  that this model generate both text and audio
                                  responses, you can use: ["text", "audio"].
                              guardrails:
                                type: array
                                items:
                                  type: object
                                  properties:
                                    id:
                                      anyOf:
                                        - type: string
                                          enum:
                                            - orq_pii_detection
                                            - orq_secret_detection
                                            - orq_sexual_moderation
                                            - orq_harmful_moderation
                                          description: The key of the guardrail.
                                        - type: string
                                          description: >-
                                            Unique key or identifier of the
                                            evaluator
                                    execute_on:
                                      type: string
                                      enum:
                                        - input
                                        - output
                                      description: >-
                                        Determines whether the guardrail runs on
                                        the input (user message) or output
                                        (model response).
                                  required:
                                    - id
                                    - execute_on
                                description: A list of guardrails to apply to the request.
                              plugins:
                                type: array
                                items:
                                  anyOf:
                                    - $ref: '#/components/schemas/PIIRedactionPlugin'
                                    - $ref: >-
                                        #/components/schemas/ResponseHealingPlugin
                                    - $ref: >-
                                        #/components/schemas/TraceScrubbingPlugin
                                description: >-
                                  Request-scoped transforms applied to the text
                                  exchanged with the model. Supports
                                  `pii_redaction`, which replaces PII with
                                  placeholders before the provider sees it and
                                  restores the original values in the response,
                                  and `response_healing`, which repairs
                                  malformed JSON in non-streaming output.
                              fallbacks:
                                type: array
                                items:
                                  type: object
                                  properties:
                                    model:
                                      type: string
                                      description: Fallback model identifier
                                      example: openai/gpt-5.4-mini
                                  required:
                                    - model
                                description: >-
                                  Array of fallback models to use if primary
                                  model fails
                              cache:
                                type: object
                                properties:
                                  ttl:
                                    type: number
                                    minimum: 1
                                    maximum: 259200
                                    default: 1800
                                    description: >-
                                      Time to live for cached responses in
                                      seconds. Maximum 259200 seconds (3 days).
                                    example: 3600
                                  type:
                                    type: string
                                    enum:
                                      - exact_match
                                required:
                                  - type
                                description: Cache configuration for the request.
                              load_balancer:
                                oneOf:
                                  - type: object
                                    properties:
                                      type:
                                        type: string
                                        enum:
                                          - weight_based
                                      models:
                                        type: array
                                        items:
                                          type: object
                                          properties:
                                            model:
                                              type: string
                                              description: Model identifier for load balancing
                                              example: openai/gpt-5.6-sol
                                            weight:
                                              type: number
                                              minimum: 0.001
                                              maximum: 1
                                              default: 0.5
                                              description: >-
                                                Weight assigned to this model for load
                                                balancing
                                              example: 0.7
                                          required:
                                            - model
                                    required:
                                      - type
                                      - models
                                description: Load balancer configuration for the request.
                                example:
                                  type: weight_based
                                  models:
                                    - model: openai/gpt-4o
                                      weight: 0.7
                                    - model: anthropic/claude-3-5-sonnet
                                      weight: 0.3
                              timeout:
                                type: object
                                properties:
                                  call_timeout:
                                    type: number
                                    minimum: 1
                                    description: Timeout value in milliseconds
                                    example: 30000
                                required:
                                  - call_timeout
                                description: >-
                                  Timeout configuration to apply to the request.
                                  If the request exceeds the timeout, it will be
                                  retried or fallback to the next model if
                                  configured.
                              cache_control:
                                type: object
                                properties:
                                  type:
                                    type: string
                                    enum:
                                      - ephemeral
                                    description: >-
                                      Create a cache control breakpoint at this
                                      content block. Accepts only the value
                                      "ephemeral".
                                  ttl:
                                    type: string
                                    enum:
                                      - 5m
                                      - 1h
                                    default: 5m
                                    description: >-
                                      The time-to-live for the cache control
                                      breakpoint. This may be one of the
                                      following values:


                                      - `5m`: 5 minutes

                                      - `1h`: 1 hour


                                      Defaults to `5m`. Only supported by
                                      `Anthropic` Claude models.
                                required:
                                  - type
                                description: >-
                                  Provider-level prompt caching configuration
                                  applied to the request. Creates a cache
                                  control breakpoint covering the request
                                  content. Only supported by `Anthropic` Claude
                                  models.
                              prompt_cache_key:
                                type: string
                                description: >-
                                  Used by OpenAI to cache responses for similar
                                  requests to optimize your cache hit rates.
                                  Replaces the legacy `user` field for prompt
                                  caching.
                            description: >-
                              Optional model parameters specific to this
                              fallback model. Overrides primary model parameters
                              if this fallback is used.
                          retry:
                            type: object
                            properties:
                              count:
                                type: number
                                minimum: 1
                                maximum: 5
                                default: 3
                                description: Number of retry attempts (1-5)
                                example: 3
                              on_codes:
                                type: array
                                items:
                                  type: number
                                  minimum: 100
                                  maximum: 599
                                minItems: 1
                                description: HTTP status codes that trigger retry logic
                                example:
                                  - 429
                                  - 500
                                  - 502
                                  - 503
                                  - 504
                            description: >-
                              Retry configuration for this fallback model.
                              Allows customizing retry count (1-5) and HTTP
                              status codes that trigger retries.
                        required:
                          - id
                        description: >-
                          Fallback model configuration with optional parameters
                          and retry settings.
                    title: Fallback Model Configuration
                    description: >-
                      Fallback model for automatic failover when primary model
                      request fails. Supports optional parameter overrides. Can
                      be a simple model ID string or a configuration object with
                      model-specific parameters. Fallbacks are tried in order.
                  description: >-
                    Optional array of fallback models used when the primary
                    model fails. Fallbacks are attempted in order. All models
                    must support tool calling.
                settings:
                  type: object
                  properties:
                    max_iterations:
                      type: integer
                      exclusiveMinimum: 0
                      maximum: 100
                      minimum: 1
                      default: 100
                      description: >-
                        Maximum iterations(llm calls) before the agent will stop
                        executing.
                    max_execution_time:
                      type: integer
                      minimum: 2
                      exclusiveMinimum: 0
                      maximum: 600
                      default: 600
                      description: >-
                        Maximum time (in seconds) for the agent thinking
                        process. This does not include the time for tool calls
                        and sub agent calls. It will be loosely enforced, the in
                        progress LLM calls will not be terminated and the last
                        assistant message will be returned.
                    max_cost:
                      type: number
                      minimum: 0
                      default: 0
                      description: >-
                        Maximum cost in USD for the agent execution. When the
                        accumulated cost exceeds this limit, the agent will stop
                        executing. Set to 0 for unlimited. Only supported in v3
                        responses
                    tool_approval_required:
                      type: string
                      enum:
                        - all
                        - respect_tool
                        - none
                      default: respect_tool
                      description: >-
                        If all, the agent will require approval for all tools.
                        If respect_tool, the agent will require approval for
                        tools that have the requires_approval flag set to true.
                        If none, the agent will not require approval for any
                        tools.
                    chat_exposed:
                      type: boolean
                      description: >-
                        When enabled, this agent is exposed as a selectable
                        target in AI Chat for users to consume.
                    tools:
                      type: array
                      items:
                        $ref: '#/components/schemas/AgentToolInputCRUD'
                      default: []
                      description: >-
                        Tools available to the agent. Built-in tools only need a
                        type, while custom tools (http, code, function) must
                        reference pre-created tools by key or id.
                    evaluators:
                      type: array
                      items:
                        type: object
                        properties:
                          id:
                            type: string
                            description: Unique key or identifier of the evaluator
                          sample_rate:
                            type: number
                            minimum: 1
                            maximum: 100
                            default: 50
                            description: >-
                              The percentage of executions to evaluate with this
                              evaluator (1-100). For example, a value of 50
                              means the evaluator will run on approximately half
                              of the executions.
                          execute_on:
                            type: string
                            enum:
                              - input
                              - output
                            description: >-
                              Determines whether the evaluator runs on the agent
                              input (user message) or output (agent response).
                          options:
                            type: object
                            additionalProperties: {}
                            description: >-
                              Evaluator-specific configuration, passed through
                              to the evaluator at run time. For
                              orq_pii_detection this carries regions, entities,
                              entity_thresholds, language and threshold, and is
                              validated against PIIDetectionGuardrailOptions:
                              regions and entities are two mutually exclusive
                              coverage modes, and every entity_thresholds key
                              must also appear in entities. on_failure is
                              rejected: an evaluator acting as a guardrail
                              always fails closed.
                        required:
                          - id
                          - execute_on
                      title: Agent evaluator configuration
                      description: Configuration for an evaluator applied to the agent
                    guardrails:
                      type: array
                      items:
                        type: object
                        properties:
                          id:
                            type: string
                            description: Unique key or identifier of the evaluator
                          sample_rate:
                            type: number
                            minimum: 1
                            maximum: 100
                            default: 50
                            description: >-
                              The percentage of executions to evaluate with this
                              evaluator (1-100). For example, a value of 50
                              means the evaluator will run on approximately half
                              of the executions.
                          execute_on:
                            type: string
                            enum:
                              - input
                              - output
                            description: >-
                              Determines whether the evaluator runs on the agent
                              input (user message) or output (agent response).
                          options:
                            type: object
                            additionalProperties: {}
                            description: >-
                              Evaluator-specific configuration, passed through
                              to the evaluator at run time. For
                              orq_pii_detection this carries regions, entities,
                              entity_thresholds, language and threshold, and is
                              validated against PIIDetectionGuardrailOptions:
                              regions and entities are two mutually exclusive
                              coverage modes, and every entity_thresholds key
                              must also appear in entities. on_failure is
                              rejected: an evaluator acting as a guardrail
                              always fails closed.
                        required:
                          - id
                          - execute_on
                      title: Agent guardrail configuration
                      description: Configuration for a guardrail applied to the agent
                  description: Configuration settings for the agent's behavior
                memory_stores:
                  type: array
                  items:
                    type: string
                  default: []
                  description: >-
                    Optional array of memory store identifiers for the agent to
                    access. Accepts both memory store IDs and keys.
                knowledge_bases:
                  type: array
                  items:
                    type: object
                    properties:
                      knowledge_id:
                        type: string
                        description: Unique identifier of the knowledge base to search
                        example: customer-knowledge-base
                    required:
                      - knowledge_id
                  default: []
                  description: >-
                    Optional array of knowledge base configurations for the
                    agent to access
                team_of_agents:
                  type: array
                  items:
                    type: object
                    properties:
                      key:
                        type: string
                        description: The unique key of the agent within the workspace
                      role:
                        type: string
                        description: >-
                          The role of the agent in this context. This is used to
                          give extra information to the leader to help it decide
                          which agent to hand off to.
                    required:
                      - key
                  default: []
                  title: Team of agents
                  description: >-
                    The agents that are accessible to this orchestrator. The
                    main agent can hand off to these agents to perform tasks.
                skills:
                  type:
                    - array
                    - 'null'
                  items:
                    type: string
                  description: >-
                    List of skills that the agent can utilize. This field allows
                    you to specify which skills the agent has access to,
                    enabling more complex and dynamic behavior.
                variables:
                  type: object
                  additionalProperties: {}
                source:
                  type: string
                  enum:
                    - internal
                    - external
                    - experiment
                engine:
                  type: string
                  enum:
                    - text
                    - jinja
                    - mustache
                  default: text
              required:
                - key
                - role
                - description
                - instructions
                - path
                - model
                - settings
      responses:
        '201':
          description: >-
            Agent successfully created and ready for use. Returns the complete
            agent manifest including the generated ID, configuration, and all
            settings.
          content:
            application/json:
              schema:
                type: object
                properties:
                  _id:
                    type: string
                  key:
                    type: string
                    pattern: ^[A-Za-z][A-Za-z0-9]*([._-][A-Za-z0-9]+)*$
                    description: Unique identifier for the agent within the workspace
                  display_name:
                    type: string
                  project_id:
                    type: string
                  created_by_id:
                    type:
                      - string
                      - 'null'
                  updated_by_id:
                    type:
                      - string
                      - 'null'
                  created:
                    type: string
                  updated:
                    type: string
                  status:
                    type: string
                    enum:
                      - live
                      - draft
                      - pending
                      - published
                    description: >-
                      The status of the agent. `Live` is the latest version of
                      the agent. `Draft` is a version that is not yet published.
                      `Pending` is a version that is pending approval.
                      `Published` is a version that was live and has been
                      replaced by a new version.
                  version:
                    type: string
                    description: Current semantic version of the agent manifest.
                  path:
                    type: string
                    description: >-
                      Entity storage path.


                      With workspace-level API keys, use the format
                      `project/folder/subfolder/...`. The first element must be
                      the display name of an existing project, followed by
                      nested folders (auto-created as needed). Example: `Default
                      Project/agents`.


                      With project-level API keys, the project is predetermined
                      by the API key, so the path is relative to that project.
                      Example: `agents`. For backward compatibility, a leading
                      project name is ignored when it matches the scoped
                      project.
                    example: Default Project
                  memory_stores:
                    type: array
                    items:
                      type: string
                    default: []
                    description: >-
                      Array of memory store identifiers. Accepts both memory
                      store IDs and keys.
                  team_of_agents:
                    type: array
                    items:
                      type: object
                      properties:
                        key:
                          type: string
                          description: The unique key of the agent within the workspace
                        role:
                          type: string
                          description: >-
                            The role of the agent in this context. This is used
                            to give extra information to the leader to help it
                            decide which agent to hand off to.
                      required:
                        - key
                    default: []
                    description: >-
                      The agents that are accessible to this orchestrator. The
                      main agent can hand off to these agents to perform tasks.
                  skills:
                    type: array
                    items:
                      type: string
                    default: []
                    description: >-
                      List of skills that the agent can utilize. This field
                      allows you to specify which skills the agent has access
                      to, enabling more complex and dynamic behavior.
                  metrics:
                    type: object
                    properties:
                      total_cost:
                        type: number
                        minimum: 0
                        default: 0
                    default:
                      total_cost: 0
                  variables:
                    type: object
                    additionalProperties: {}
                    description: Extracted variables from agent instructions
                  knowledge_bases:
                    type: array
                    items:
                      type: object
                      properties:
                        knowledge_id:
                          type: string
                          description: Unique identifier of the knowledge base to search
                          example: customer-knowledge-base
                      required:
                        - knowledge_id
                    description: Agent knowledge bases reference
                  source:
                    type: string
                    enum:
                      - internal
                      - external
                      - experiment
                  engine:
                    type: string
                    enum:
                      - text
                      - jinja
                      - mustache
                    default: text
                  type:
                    type: string
                    enum:
                      - internal
                      - a2a
                    default: internal
                    description: >-
                      Agent type: internal (orq.ai-managed) or a2a (external
                      A2A-compliant)
                  role:
                    type: string
                    minLength: 1
                  description:
                    type: string
                  system_prompt:
                    type:
                      - string
                      - 'null'
                    minLength: 1
                  instructions:
                    type: string
                  settings:
                    type: object
                    properties:
                      max_iterations:
                        type: integer
                        exclusiveMinimum: 0
                        maximum: 100
                        minimum: 1
                        default: 100
                        description: >-
                          Maximum iterations(llm calls) before the agent will
                          stop executing.
                      max_execution_time:
                        type: integer
                        minimum: 2
                        exclusiveMinimum: 0
                        maximum: 600
                        default: 600
                        description: >-
                          Maximum time (in seconds) for the agent thinking
                          process. This does not include the time for tool calls
                          and sub agent calls. It will be loosely enforced, the
                          in progress LLM calls will not be terminated and the
                          last assistant message will be returned.
                      max_cost:
                        type: number
                        minimum: 0
                        default: 0
                        description: >-
                          Maximum cost in USD for the agent execution. When the
                          accumulated cost exceeds this limit, the agent will
                          stop executing. Set to 0 for unlimited. Only supported
                          in v3 responses
                      tool_approval_required:
                        type: string
                        enum:
                          - all
                          - respect_tool
                          - none
                        default: respect_tool
                        description: >-
                          If all, the agent will require approval for all tools.
                          If respect_tool, the agent will require approval for
                          tools that have the requires_approval flag set to
                          true. If none, the agent will not require approval for
                          any tools.
                      chat_exposed:
                        type: boolean
                        description: >-
                          When enabled, this agent is exposed as a selectable
                          target in AI Chat for users to consume.
                      tools:
                        type: array
                        items:
                          type: object
                          properties:
                            id:
                              type: string
                              format: ulid
                              pattern: ^[0-9A-HJKMNP-TV-Z]{26}$
                              readOnly: true
                              description: The id of the resource
                            key:
                              type: string
                              description: Optional tool key for custom tools
                            action_type:
                              type: string
                            display_name:
                              type: string
                            description:
                              type: string
                              description: Optional tool description
                            configuration:
                              type: object
                              additionalProperties: {}
                              description: >-
                                Static tool configuration set at design time.
                                Merged over LLM-provided arguments at execution
                                time.
                            requires_approval:
                              type: boolean
                              default: false
                            tool_id:
                              type: string
                              description: >-
                                Nested tool ID for MCP tools (identifies
                                specific tool within MCP server)
                            conditions:
                              type: array
                              items:
                                type: object
                                properties:
                                  condition:
                                    type: string
                                    description: The argument of the tool call to evaluate
                                  operator:
                                    type: string
                                    description: The operator to use
                                  value:
                                    type: string
                                    description: The value to compare against
                                required:
                                  - condition
                                  - operator
                                  - value
                              default: []
                            timeout:
                              type: number
                              minimum: 1
                              maximum: 600
                              description: >-
                                Tool execution timeout in seconds for this agent
                                (max: 10 minutes). Overrides the timeout
                                configured on the tool definition.
                          required:
                            - id
                            - action_type
                        default: []
                      evaluators:
                        type: array
                        items:
                          type: object
                          properties:
                            id:
                              type: string
                              description: Unique key or identifier of the evaluator
                            sample_rate:
                              type: number
                              minimum: 1
                              maximum: 100
                              default: 50
                              description: >-
                                The percentage of executions to evaluate with
                                this evaluator (1-100). For example, a value of
                                50 means the evaluator will run on approximately
                                half of the executions.
                            execute_on:
                              type: string
                              enum:
                                - input
                                - output
                              description: >-
                                Determines whether the evaluator runs on the
                                agent input (user message) or output (agent
                                response).
                            options:
                              type: object
                              additionalProperties: {}
                              description: >-
                                Evaluator-specific configuration, passed through
                                to the evaluator at run time. For
                                orq_pii_detection this carries regions,
                                entities, entity_thresholds, language and
                                threshold, and is validated against
                                PIIDetectionGuardrailOptions: regions and
                                entities are two mutually exclusive coverage
                                modes, and every entity_thresholds key must also
                                appear in entities. on_failure is rejected: an
                                evaluator acting as a guardrail always fails
                                closed.
                          required:
                            - id
                            - execute_on
                        title: Agent evaluator configuration
                        description: Configuration for an evaluator applied to the agent
                      guardrails:
                        type: array
                        items:
                          type: object
                          properties:
                            id:
                              type: string
                              description: Unique key or identifier of the evaluator
                            sample_rate:
                              type: number
                              minimum: 1
                              maximum: 100
                              default: 50
                              description: >-
                                The percentage of executions to evaluate with
                                this evaluator (1-100). For example, a value of
                                50 means the evaluator will run on approximately
                                half of the executions.
                            execute_on:
                              type: string
                              enum:
                                - input
                                - output
                              description: >-
                                Determines whether the evaluator runs on the
                                agent input (user message) or output (agent
                                response).
                            options:
                              type: object
                              additionalProperties: {}
                              description: >-
                                Evaluator-specific configuration, passed through
                                to the evaluator at run time. For
                                orq_pii_detection this carries regions,
                                entities, entity_thresholds, language and
                                threshold, and is validated against
                                PIIDetectionGuardrailOptions: regions and
                                entities are two mutually exclusive coverage
                                modes, and every entity_thresholds key must also
                                appear in entities. on_failure is rejected: an
                                evaluator acting as a guardrail always fails
                                closed.
                          required:
                            - id
                            - execute_on
                        title: Agent guardrail configuration
                        description: >-
                          Configuration for a guardrail applied to the agent.
                          sample_rate has no effect here: a guardrail is a gate
                          rather than a measurement, so it runs on every
                          request.
                    default:
                      max_execution_time: 600
                      max_iterations: 100
                      max_cost: 0
                      tool_approval_required: respect_tool
                      tools: []
                  model:
                    type: object
                    properties:
                      id:
                        type: string
                        description: >-
                          ID of the primary model, in provider/model-id format
                          (for example `openai/gpt-5.6-sol`)
                      integration_id:
                        type:
                          - string
                          - 'null'
                        description: >-
                          Optional integration ID for custom model
                          configurations
                      parameters:
                        type:
                          - object
                          - 'null'
                        properties:
                          name:
                            description: >-
                              The name to display on the trace. If not
                              specified, the default system name will be used.
                            type: string
                          frequency_penalty:
                            type:
                              - number
                              - 'null'
                            description: >-
                              Number between -2.0 and 2.0. Positive values
                              penalize new tokens based on their existing
                              frequency in the text so far, decreasing the
                              model's likelihood to repeat the same line
                              verbatim.
                          max_tokens:
                            type:
                              - integer
                              - 'null'
                            description: >-
                              `[Deprecated]`. The maximum number of tokens that
                              can be generated in the chat completion. This
                              value can be used to control costs for text
                              generated via API. 

                               This value is now `deprecated` in favor of `max_completion_tokens`, and is not compatible with o1 series models.
                          max_completion_tokens:
                            type:
                              - integer
                              - 'null'
                            exclusiveMinimum: 0
                            description: >-
                              An upper bound for the number of tokens that can
                              be generated for a completion, including visible
                              output tokens and reasoning tokens
                          presence_penalty:
                            type:
                              - number
                              - 'null'
                            description: >-
                              Number between -2.0 and 2.0. Positive values
                              penalize new tokens based on whether they appear
                              in the text so far, increasing the model's
                              likelihood to talk about new topics.
                          response_format:
                            oneOf:
                              - type: object
                                properties:
                                  type:
                                    type: string
                                    enum:
                                      - text
                                required:
                                  - type
                                title: Text
                                description: >-


                                  Default response format. Used to generate text
                                  responses
                              - type: object
                                properties:
                                  type:
                                    type: string
                                    enum:
                                      - json_object
                                required:
                                  - type
                                title: JSON object
                                description: >-


                                  JSON object response format. An older method
                                  of generating JSON responses. Using
                                  `json_schema` is recommended for models that
                                  support it. Note that the model will not
                                  generate JSON without a system or user message
                                  instructing it to do so.
                              - type: object
                                properties:
                                  type:
                                    type: string
                                    enum:
                                      - json_schema
                                  json_schema:
                                    type: object
                                    properties:
                                      description:
                                        description: >-
                                          A description of what the response
                                          format is for, used by the model to
                                          determine how to respond in the format.
                                        type: string
                                      name:
                                        type: string
                                        description: >-
                                          The name of the response format. Must be
                                          a-z, A-Z, 0-9, or contain underscores
                                          and dashes, with a maximum length of 64.
                                      schema:
                                        description: >-
                                          The schema for the response format,
                                          described as a JSON Schema object.
                                      strict:
                                        type: boolean
                                        default: false
                                        description: >-
                                          Whether to enable strict schema
                                          adherence when generating the output. If
                                          set to true, the model will always
                                          follow the exact schema defined in the
                                          schema field. Only a subset of JSON
                                          Schema is supported when strict is true.
                                    required:
                                      - name
                                required:
                                  - type
                                  - json_schema
                                title: JSON schema
                                description: >-


                                  JSON Schema response format. Used to generate
                                  structured JSON responses
                            description: >-
                              An object specifying the format that the model
                              must output
                          reasoning_effort:
                            type: string
                            enum:
                              - none
                              - minimal
                              - low
                              - medium
                              - high
                              - xhigh
                              - max
                            description: >-
                              Constrains effort on reasoning for [reasoning
                              models](https://platform.openai.com/docs/guides/reasoning).
                              Currently supported values are `none`, `minimal`,
                              `low`, `medium`, `high`, `xhigh`, and `max`.
                              Reducing reasoning effort can result in faster
                              responses and fewer tokens used on reasoning in a
                              response.


                              - `gpt-5.1` defaults to `none`, which does not
                              perform reasoning. The supported reasoning values
                              for `gpt-5.1` are `none`, `low`, `medium`, and
                              `high`. Tool calls are supported for all reasoning
                              values in gpt-5.1.

                              - All models before `gpt-5.1` default to `medium`
                              reasoning effort, and do not support `none`.

                              - The `gpt-5-pro` model defaults to (and only
                              supports) `high` reasoning effort.

                              - `xhigh` is currently only supported for
                              `gpt-5.1-codex-max`.


                              Any of "none", "minimal", "low", "medium", "high",
                              "xhigh", "max".
                          verbosity:
                            type: string
                            description: >-
                              Adjusts response verbosity. Lower levels yield
                              shorter answers.
                          seed:
                            type:
                              - number
                              - 'null'
                            description: >-
                              If specified, our system will make a best effort
                              to sample deterministically, such that repeated
                              requests with the same seed and parameters should
                              return the same result.
                          stop:
                            anyOf:
                              - type: string
                              - type: array
                                items:
                                  type: string
                                maxItems: 4
                              - type: 'null'
                            description: >-
                              Up to 4 sequences where the API will stop
                              generating further tokens.
                          thinking:
                            oneOf:
                              - $ref: >-
                                  #/components/schemas/ThinkingConfigDisabledSchema
                              - $ref: >-
                                  #/components/schemas/ThinkingConfigEnabledSchema
                              - $ref: >-
                                  #/components/schemas/ThinkingConfigAdaptiveSchema
                            discriminator:
                              propertyName: type
                              mapping:
                                disabled: >-
                                  #/components/schemas/ThinkingConfigDisabledSchema
                                enabled: >-
                                  #/components/schemas/ThinkingConfigEnabledSchema
                                adaptive: >-
                                  #/components/schemas/ThinkingConfigAdaptiveSchema
                          temperature:
                            type:
                              - number
                              - 'null'
                            minimum: 0
                            maximum: 2
                            description: >-
                              What sampling temperature to use, between 0 and 2.
                              Higher values like 0.8 will make the output more
                              random, while lower values like 0.2 will make it
                              more focused and deterministic.
                          top_p:
                            type:
                              - number
                              - 'null'
                            minimum: 0
                            maximum: 1
                            description: >-
                              An alternative to sampling with temperature,
                              called nucleus sampling, where the model considers
                              the results of the tokens with top_p probability
                              mass. 
                          top_k:
                            type:
                              - number
                              - 'null'
                            description: >-
                              Limits the model to consider only the top k most
                              likely tokens at each step.
                          tool_choice:
                            anyOf:
                              - type: string
                                enum:
                                  - none
                                  - auto
                                  - required
                              - type: object
                                properties:
                                  type:
                                    type: string
                                    enum:
                                      - function
                                    description: >-
                                      The type of the tool. Currently, only
                                      function is supported.
                                  function:
                                    type: object
                                    properties:
                                      name:
                                        type: string
                                        description: The name of the function to call.
                                    required:
                                      - name
                                required:
                                  - function
                            description: >-
                              Controls which (if any) tool is called by the
                              model.
                          parallel_tool_calls:
                            type: boolean
                            description: >-
                              Whether to enable parallel function calling during
                              tool use.
                          modalities:
                            type:
                              - array
                              - 'null'
                            items:
                              type: string
                              enum:
                                - text
                                - audio
                            description: >-
                              Output types that you would like the model to
                              generate. Most models are capable of generating
                              text, which is the default: ["text"]. The
                              gpt-4o-audio-preview model can also be used to
                              generate audio. To request that this model
                              generate both text and audio responses, you can
                              use: ["text", "audio"].
                          guardrails:
                            type: array
                            items:
                              type: object
                              properties:
                                id:
                                  anyOf:
                                    - type: string
                                      enum:
                                        - orq_pii_detection
                                        - orq_secret_detection
                                        - orq_sexual_moderation
                                        - orq_harmful_moderation
                                      description: The key of the guardrail.
                                    - type: string
                                      description: >-
                                        Unique key or identifier of the
                                        evaluator
                                execute_on:
                                  type: string
                                  enum:
                                    - input
                                    - output
                                  description: >-
                                    Determines whether the guardrail runs on the
                                    input (user message) or output (model
                                    response).
                              required:
                                - id
                                - execute_on
                            description: A list of guardrails to apply to the request.
                          plugins:
                            type: array
                            items:
                              anyOf:
                                - $ref: '#/components/schemas/PIIRedactionPlugin'
                                - $ref: '#/components/schemas/ResponseHealingPlugin'
                                - $ref: '#/components/schemas/TraceScrubbingPlugin'
                            description: >-
                              Request-scoped transforms applied to the text
                              exchanged with the model. Supports
                              `pii_redaction`, which replaces PII with
                              placeholders before the provider sees it and
                              restores the original values in the response, and
                              `response_healing`, which repairs malformed JSON
                              in non-streaming output.
                          fallbacks:
                            type: array
                            items:
                              type: object
                              properties:
                                model:
                                  type: string
                                  description: Fallback model identifier
                                  example: openai/gpt-5.4-mini
                              required:
                                - model
                            description: >-
                              Array of fallback models to use if primary model
                              fails
                          cache:
                            type: object
                            properties:
                              ttl:
                                type: number
                                minimum: 1
                                maximum: 259200
                                default: 1800
                                description: >-
                                  Time to live for cached responses in seconds.
                                  Maximum 259200 seconds (3 days).
                                example: 3600
                              type:
                                type: string
                                enum:
                                  - exact_match
                            required:
                              - type
                            description: Cache configuration for the request.
                          load_balancer:
                            oneOf:
                              - type: object
                                properties:
                                  type:
                                    type: string
                                    enum:
                                      - weight_based
                                  models:
                                    type: array
                                    items:
                                      type: object
                                      properties:
                                        model:
                                          type: string
                                          description: Model identifier for load balancing
                                          example: openai/gpt-5.6-sol
                                        weight:
                                          type: number
                                          minimum: 0.001
                                          maximum: 1
                                          default: 0.5
                                          description: >-
                                            Weight assigned to this model for load
                                            balancing
                                          example: 0.7
                                      required:
                                        - model
                                required:
                                  - type
                                  - models
                            description: Load balancer configuration for the request.
                            example:
                              type: weight_based
                              models:
                                - model: openai/gpt-4o
                                  weight: 0.7
                                - model: anthropic/claude-3-5-sonnet
                                  weight: 0.3
                          timeout:
                            type: object
                            properties:
                              call_timeout:
                                type: number
                                minimum: 1
                                description: Timeout value in milliseconds
                                example: 30000
                            required:
                              - call_timeout
                            description: >-
                              Timeout configuration to apply to the request. If
                              the request exceeds the timeout, it will be
                              retried or fallback to the next model if
                              configured.
                          cache_control:
                            type: object
                            properties:
                              type:
                                type: string
                                enum:
                                  - ephemeral
                                description: >-
                                  Create a cache control breakpoint at this
                                  content block. Accepts only the value
                                  "ephemeral".
                              ttl:
                                type: string
                                enum:
                                  - 5m
                                  - 1h
                                default: 5m
                                description: >-
                                  The time-to-live for the cache control
                                  breakpoint. This may be one of the following
                                  values:


                                  - `5m`: 5 minutes

                                  - `1h`: 1 hour


                                  Defaults to `5m`. Only supported by
                                  `Anthropic` Claude models.
                            required:
                              - type
                            description: >-
                              Provider-level prompt caching configuration
                              applied to the request. Creates a cache control
                              breakpoint covering the request content. Only
                              supported by `Anthropic` Claude models.
                          prompt_cache_key:
                            type: string
                            description: >-
                              Used by OpenAI to cache responses for similar
                              requests to optimize your cache hit rates.
                              Replaces the legacy `user` field for prompt
                              caching.
                        description: >-
                          Model behavior parameters (snake_case) stored as part
                          of the agent configuration. These become the default
                          parameters used when the agent is executed. Commonly
                          used: temperature (0-2, controls randomness; the
                          selected model may impose a lower maximum),
                          max_completion_tokens (response length), top_p
                          (nucleus sampling). Advanced: frequency_penalty,
                          presence_penalty, response_format (JSON/structured
                          output), reasoning_effort (for o1/thinking models),
                          seed (reproducibility), stop sequences. Model-specific
                          support varies. Runtime parameters in agent execution
                          requests can override these defaults.
                      retry:
                        type: object
                        properties:
                          count:
                            type: number
                            minimum: 1
                            maximum: 5
                            default: 3
                            description: Number of retry attempts (1-5)
                            example: 3
                          on_codes:
                            type: array
                            items:
                              type: number
                              minimum: 100
                              maximum: 599
                            minItems: 1
                            description: HTTP status codes that trigger retry logic
                            example:
                              - 429
                              - 500
                              - 502
                              - 503
                              - 504
                        description: >-
                          Retry configuration for model requests. Allows
                          customizing retry count (1-5) and HTTP status codes
                          that trigger retries. Default codes: [429]. Common
                          codes: 500 (internal error), 429 (rate limit),
                          502/503/504 (gateway errors).
                      fallback_models:
                        type:
                          - array
                          - 'null'
                        items:
                          anyOf:
                            - type: string
                              description: >-
                                A fallback model ID string (e.g.,
                                `openai/gpt-4o-mini`). Will be used if the
                                primary model request fails. Must support tool
                                calling.
                            - type: object
                              properties:
                                id:
                                  type: string
                                  description: >-
                                    A fallback model ID string. Must support
                                    tool calling.
                                parameters:
                                  type: object
                                  properties:
                                    name:
                                      description: >-
                                        The name to display on the trace. If not
                                        specified, the default system name will
                                        be used.
                                      type: string
                                    frequency_penalty:
                                      type:
                                        - number
                                        - 'null'
                                      description: >-
                                        Number between -2.0 and 2.0. Positive
                                        values penalize new tokens based on
                                        their existing frequency in the text so
                                        far, decreasing the model's likelihood
                                        to repeat the same line verbatim.
                                    max_tokens:
                                      type:
                                        - integer
                                        - 'null'
                                      description: >-
                                        `[Deprecated]`. The maximum number of
                                        tokens that can be generated in the chat
                                        completion. This value can be used to
                                        control costs for text generated via
                                        API. 

                                         This value is now `deprecated` in favor of `max_completion_tokens`, and is not compatible with o1 series models.
                                    max_completion_tokens:
                                      type:
                                        - integer
                                        - 'null'
                                      exclusiveMinimum: 0
                                      description: >-
                                        An upper bound for the number of tokens
                                        that can be generated for a completion,
                                        including visible output tokens and
                                        reasoning tokens
                                    presence_penalty:
                                      type:
                                        - number
                                        - 'null'
                                      description: >-
                                        Number between -2.0 and 2.0. Positive
                                        values penalize new tokens based on
                                        whether they appear in the text so far,
                                        increasing the model's likelihood to
                                        talk about new topics.
                                    response_format:
                                      oneOf:
                                        - type: object
                                          properties:
                                            type:
                                              type: string
                                              enum:
                                                - text
                                          required:
                                            - type
                                          title: Text
                                          description: >-


                                            Default response format. Used to
                                            generate text responses
                                        - type: object
                                          properties:
                                            type:
                                              type: string
                                              enum:
                                                - json_object
                                          required:
                                            - type
                                          title: JSON object
                                          description: >-


                                            JSON object response format. An older
                                            method of generating JSON responses.
                                            Using `json_schema` is recommended for
                                            models that support it. Note that the
                                            model will not generate JSON without a
                                            system or user message instructing it to
                                            do so.
                                        - type: object
                                          properties:
                                            type:
                                              type: string
                                              enum:
                                                - json_schema
                                            json_schema:
                                              type: object
                                              properties:
                                                description:
                                                  description: >-
                                                    A description of what the response
                                                    format is for, used by the model to
                                                    determine how to respond in the format.
                                                  type: string
                                                name:
                                                  type: string
                                                  description: >-
                                                    The name of the response format. Must be
                                                    a-z, A-Z, 0-9, or contain underscores
                                                    and dashes, with a maximum length of 64.
                                                schema:
                                                  description: >-
                                                    The schema for the response format,
                                                    described as a JSON Schema object.
                                                strict:
                                                  type: boolean
                                                  default: false
                                                  description: >-
                                                    Whether to enable strict schema
                                                    adherence when generating the output. If
                                                    set to true, the model will always
                                                    follow the exact schema defined in the
                                                    schema field. Only a subset of JSON
                                                    Schema is supported when strict is true.
                                              required:
                                                - name
                                          required:
                                            - type
                                            - json_schema
                                          title: JSON schema
                                          description: >-


                                            JSON Schema response format. Used to
                                            generate structured JSON responses
                                      description: >-
                                        An object specifying the format that the
                                        model must output
                                    reasoning_effort:
                                      type: string
                                      enum:
                                        - none
                                        - minimal
                                        - low
                                        - medium
                                        - high
                                        - xhigh
                                        - max
                                      description: >-
                                        Constrains effort on reasoning for
                                        [reasoning
                                        models](https://platform.openai.com/docs/guides/reasoning).
                                        Currently supported values are `none`,
                                        `minimal`, `low`, `medium`, `high`,
                                        `xhigh`, and `max`. Reducing reasoning
                                        effort can result in faster responses
                                        and fewer tokens used on reasoning in a
                                        response.


                                        - `gpt-5.1` defaults to `none`, which
                                        does not perform reasoning. The
                                        supported reasoning values for `gpt-5.1`
                                        are `none`, `low`, `medium`, and `high`.
                                        Tool calls are supported for all
                                        reasoning values in gpt-5.1.

                                        - All models before `gpt-5.1` default to
                                        `medium` reasoning effort, and do not
                                        support `none`.

                                        - The `gpt-5-pro` model defaults to (and
                                        only supports) `high` reasoning effort.

                                        - `xhigh` is currently only supported
                                        for `gpt-5.1-codex-max`.


                                        Any of "none", "minimal", "low",
                                        "medium", "high", "xhigh", "max".
                                    verbosity:
                                      type: string
                                      description: >-
                                        Adjusts response verbosity. Lower levels
                                        yield shorter answers.
                                    seed:
                                      type:
                                        - number
                                        - 'null'
                                      description: >-
                                        If specified, our system will make a
                                        best effort to sample deterministically,
                                        such that repeated requests with the
                                        same seed and parameters should return
                                        the same result.
                                    stop:
                                      anyOf:
                                        - type: string
                                        - type: array
                                          items:
                                            type: string
                                          maxItems: 4
                                        - type: 'null'
                                      description: >-
                                        Up to 4 sequences where the API will
                                        stop generating further tokens.
                                    thinking:
                                      oneOf:
                                        - $ref: >-
                                            #/components/schemas/ThinkingConfigDisabledSchema
                                        - $ref: >-
                                            #/components/schemas/ThinkingConfigEnabledSchema
                                        - $ref: >-
                                            #/components/schemas/ThinkingConfigAdaptiveSchema
                                      discriminator:
                                        propertyName: type
                                        mapping:
                                          disabled: >-
                                            #/components/schemas/ThinkingConfigDisabledSchema
                                          enabled: >-
                                            #/components/schemas/ThinkingConfigEnabledSchema
                                          adaptive: >-
                                            #/components/schemas/ThinkingConfigAdaptiveSchema
                                    temperature:
                                      type:
                                        - number
                                        - 'null'
                                      minimum: 0
                                      maximum: 2
                                      description: >-
                                        What sampling temperature to use,
                                        between 0 and 2. Higher values like 0.8
                                        will make the output more random, while
                                        lower values like 0.2 will make it more
                                        focused and deterministic.
                                    top_p:
                                      type:
                                        - number
                                        - 'null'
                                      minimum: 0
                                      maximum: 1
                                      description: >-
                                        An alternative to sampling with
                                        temperature, called nucleus sampling,
                                        where the model considers the results of
                                        the tokens with top_p probability mass. 
                                    top_k:
                                      type:
                                        - number
                                        - 'null'
                                      description: >-
                                        Limits the model to consider only the
                                        top k most likely tokens at each step.
                                    tool_choice:
                                      anyOf:
                                        - type: string
                                          enum:
                                            - none
                                            - auto
                                            - required
                                        - type: object
                                          properties:
                                            type:
                                              type: string
                                              enum:
                                                - function
                                              description: >-
                                                The type of the tool. Currently, only
                                                function is supported.
                                            function:
                                              type: object
                                              properties:
                                                name:
                                                  type: string
                                                  description: The name of the function to call.
                                              required:
                                                - name
                                          required:
                                            - function
                                      description: >-
                                        Controls which (if any) tool is called
                                        by the model.
                                    parallel_tool_calls:
                                      type: boolean
                                      description: >-
                                        Whether to enable parallel function
                                        calling during tool use.
                                    modalities:
                                      type:
                                        - array
                                        - 'null'
                                      items:
                                        type: string
                                        enum:
                                          - text
                                          - audio
                                      description: >-
                                        Output types that you would like the
                                        model to generate. Most models are
                                        capable of generating text, which is the
                                        default: ["text"]. The
                                        gpt-4o-audio-preview model can also be
                                        used to generate audio. To request that
                                        this model generate both text and audio
                                        responses, you can use: ["text",
                                        "audio"].
                                    guardrails:
                                      type: array
                                      items:
                                        type: object
                                        properties:
                                          id:
                                            anyOf:
                                              - type: string
                                                enum:
                                                  - orq_pii_detection
                                                  - orq_secret_detection
                                                  - orq_sexual_moderation
                                                  - orq_harmful_moderation
                                                description: The key of the guardrail.
                                              - type: string
                                                description: >-
                                                  Unique key or identifier of the
                                                  evaluator
                                          execute_on:
                                            type: string
                                            enum:
                                              - input
                                              - output
                                            description: >-
                                              Determines whether the guardrail runs on
                                              the input (user message) or output
                                              (model response).
                                        required:
                                          - id
                                          - execute_on
                                      description: >-
                                        A list of guardrails to apply to the
                                        request.
                                    plugins:
                                      type: array
                                      items:
                                        anyOf:
                                          - $ref: '#/components/schemas/PIIRedactionPlugin'
                                          - $ref: >-
                                              #/components/schemas/ResponseHealingPlugin
                                          - $ref: >-
                                              #/components/schemas/TraceScrubbingPlugin
                                      description: >-
                                        Request-scoped transforms applied to the
                                        text exchanged with the model. Supports
                                        `pii_redaction`, which replaces PII with
                                        placeholders before the provider sees it
                                        and restores the original values in the
                                        response, and `response_healing`, which
                                        repairs malformed JSON in non-streaming
                                        output.
                                    fallbacks:
                                      type: array
                                      items:
                                        type: object
                                        properties:
                                          model:
                                            type: string
                                            description: Fallback model identifier
                                            example: openai/gpt-5.4-mini
                                        required:
                                          - model
                                      description: >-
                                        Array of fallback models to use if
                                        primary model fails
                                    cache:
                                      type: object
                                      properties:
                                        ttl:
                                          type: number
                                          minimum: 1
                                          maximum: 259200
                                          default: 1800
                                          description: >-
                                            Time to live for cached responses in
                                            seconds. Maximum 259200 seconds (3
                                            days).
                                          example: 3600
                                        type:
                                          type: string
                                          enum:
                                            - exact_match
                                      required:
                                        - type
                                      description: Cache configuration for the request.
                                    load_balancer:
                                      oneOf:
                                        - type: object
                                          properties:
                                            type:
                                              type: string
                                              enum:
                                                - weight_based
                                            models:
                                              type: array
                                              items:
                                                type: object
                                                properties:
                                                  model:
                                                    type: string
                                                    description: Model identifier for load balancing
                                                    example: openai/gpt-5.6-sol
                                                  weight:
                                                    type: number
                                                    minimum: 0.001
                                                    maximum: 1
                                                    default: 0.5
                                                    description: >-
                                                      Weight assigned to this model for load
                                                      balancing
                                                    example: 0.7
                                                required:
                                                  - model
                                          required:
                                            - type
                                            - models
                                      description: >-
                                        Load balancer configuration for the
                                        request.
                                      example:
                                        type: weight_based
                                        models:
                                          - model: openai/gpt-4o
                                            weight: 0.7
                                          - model: anthropic/claude-3-5-sonnet
                                            weight: 0.3
                                    timeout:
                                      type: object
                                      properties:
                                        call_timeout:
                                          type: number
                                          minimum: 1
                                          description: Timeout value in milliseconds
                                          example: 30000
                                      required:
                                        - call_timeout
                                      description: >-
                                        Timeout configuration to apply to the
                                        request. If the request exceeds the
                                        timeout, it will be retried or fallback
                                        to the next model if configured.
                                    cache_control:
                                      type: object
                                      properties:
                                        type:
                                          type: string
                                          enum:
                                            - ephemeral
                                          description: >-
                                            Create a cache control breakpoint at
                                            this content block. Accepts only the
                                            value "ephemeral".
                                        ttl:
                                          type: string
                                          enum:
                                            - 5m
                                            - 1h
                                          default: 5m
                                          description: >-
                                            The time-to-live for the cache control
                                            breakpoint. This may be one of the
                                            following values:


                                            - `5m`: 5 minutes

                                            - `1h`: 1 hour


                                            Defaults to `5m`. Only supported by
                                            `Anthropic` Claude models.
                                      required:
                                        - type
                                      description: >-
                                        Provider-level prompt caching
                                        configuration applied to the request.
                                        Creates a cache control breakpoint
                                        covering the request content. Only
                                        supported by `Anthropic` Claude models.
                                    prompt_cache_key:
                                      type: string
                                      description: >-
                                        Used by OpenAI to cache responses for
                                        similar requests to optimize your cache
                                        hit rates. Replaces the legacy `user`
                                        field for prompt caching.
                                  description: >-
                                    Optional model parameters specific to this
                                    fallback model. Overrides primary model
                                    parameters if this fallback is used.
                                retry:
                                  type: object
                                  properties:
                                    count:
                                      type: number
                                      minimum: 1
                                      maximum: 5
                                      default: 3
                                      description: Number of retry attempts (1-5)
                                      example: 3
                                    on_codes:
                                      type: array
                                      items:
                                        type: number
                                        minimum: 100
                                        maximum: 599
                                      minItems: 1
                                      description: >-
                                        HTTP status codes that trigger retry
                                        logic
                                      example:
                                        - 429
                                        - 500
                                        - 502
                                        - 503
                                        - 504
                                  description: >-
                                    Retry configuration for this fallback model.
                                    Allows customizing retry count (1-5) and
                                    HTTP status codes that trigger retries.
                              required:
                                - id
                              description: >-
                                Fallback model configuration with optional
                                parameters and retry settings.
                          title: Fallback Model Configuration
                          description: >-
                            Fallback model for automatic failover when primary
                            model request fails. Supports optional parameter
                            overrides. Can be a simple model ID string or a
                            configuration object with model-specific parameters.
                            Fallbacks are tried in order.
                        description: >-
                          Optional array of fallback models (string IDs or
                          config objects) that will be used automatically in
                          order if the primary model fails
                    required:
                      - id
                required:
                  - _id
                  - key
                  - project_id
                  - status
                  - path
                  - role
                  - description
                  - instructions
                  - model
        '401':
          description: Unauthorized.
        '409':
          description: An agent with the same key already exists in this workspace.
components:
  schemas:
    ThinkingConfigDisabledSchema:
      type: object
      properties:
        type:
          type: string
          enum:
            - disabled
          description: Disables the thinking mode capability
      required:
        - type
      title: Thinking config disabled
      description: Disables the thinking mode capability
    ThinkingConfigEnabledSchema:
      type: object
      properties:
        type:
          type: string
          enum:
            - enabled
          description: Enables or disables the thinking mode capability
        budget_tokens:
          type: number
          description: >-
            Determines how many tokens the model can use for its internal
            reasoning process. Larger budgets can enable more thorough analysis
            for complex problems, improving response quality. Must be ≥1024 and
            less than `max_tokens`.
        thinking_level:
          type: string
          enum:
            - minimal
            - low
            - medium
            - high
          description: >-
            The level of reasoning the model should use. This setting is
            supported only by `gemini-3` models. If budget_tokens is specified
            and `thinking_level` is available, `budget_tokens` will be ignored.
      required:
        - type
        - budget_tokens
      title: Thinking config enabled
      description: Enables the thinking mode capability
    ThinkingConfigAdaptiveSchema:
      type: object
      properties:
        type:
          type: string
          enum:
            - adaptive
          description: >-
            Lets the model dynamically determine when and how much to use
            extended thinking based on the complexity of each request. Supported
            on Claude Opus 4.6 and Sonnet 4.6.
      required:
        - type
      title: Thinking config adaptive
      description: >-
        Enables adaptive thinking mode where the model dynamically determines
        thinking depth
    PIIRedactionPlugin:
      type: object
      properties:
        id:
          type: string
          enum:
            - pii_redaction
          description: PII redaction plugin.
        language:
          type: string
          description: >-
            Detector language. Accepts "auto" to detect the language per
            request; GET /v2/pii/capabilities lists the concrete languages and
            does not include "auto". Omitting the field falls back to en.
        regions:
          type: array
          items:
            type: string
          description: >-
            Region codes selecting whole regions of coverage (e.g. "nl", "gb").
            Every entity type those regions cover is redacted, alongside the
            base catalog. ["all"] is exclusive, and leaving both this and
            entities empty also runs every region, so selecting nothing is the
            widest request rather than the narrowest. Combines with entities:
            the two selections are unioned, so entities adds types on top of the
            region coverage.
        entities:
          type: array
          items:
            type: string
          description: >-
            The entity types to redact. A named type is redacted even when it
            belongs to a region, so a region's types can be selected
            individually without naming the region. On its own this is a strict
            allowlist; alongside regions it adds to the region coverage. See GET
            /v2/pii/capabilities for valid types.
        entity_thresholds:
          type: object
          additionalProperties:
            type: number
            minimum: 0
            maximum: 1
          description: >-
            Per-entity confidence cutoff overrides in [0,1]. Tunes confidence
            only and never changes which types are redacted, so every key must
            also appear in entities.
        threshold:
          type: number
          minimum: 0
          maximum: 1
          description: >-
            Baseline confidence cutoff applied to every entity type without a
            per-entity override.
        on_failure:
          type: string
          enum:
            - block
            - passthrough
          description: Behaviour when detection fails.
      required:
        - id
      additionalProperties: false
      title: PII redaction plugin
    ResponseHealingPlugin:
      type: object
      properties:
        id:
          type: string
          enum:
            - response_healing
          description: Plugin discriminator. Must be `response_healing`.
      required:
        - id
      additionalProperties: false
      title: Response healing plugin
    TraceScrubbingPlugin:
      type: object
      properties:
        id:
          type: string
          enum:
            - trace_scrubbing
          description: Plugin discriminator. Must be `trace_scrubbing`.
        mask:
          type: array
          items:
            type: string
            enum:
              - all
              - system
              - input
              - output
              - metadata
              - variables
          minItems: 1
          description: >-
            Trace surfaces to scrub. `all` includes system, input, output,
            metadata, and variables.
      required:
        - id
        - mask
      additionalProperties: false
      title: Trace scrubbing plugin
    AgentToolInputCRUD:
      anyOf:
        - $ref: '#/components/schemas/GoogleSearchToolInput'
        - $ref: '#/components/schemas/WebScraperToolInput'
        - $ref: '#/components/schemas/CallSubAgentToolInput'
        - $ref: '#/components/schemas/RetrieveAgentsToolInput'
        - $ref: '#/components/schemas/QueryMemoryStoreToolInput'
        - $ref: '#/components/schemas/WriteMemoryStoreToolInput'
        - $ref: '#/components/schemas/RetrieveMemoryStoresToolInput'
        - $ref: '#/components/schemas/DeleteMemoryDocumentToolInput'
        - $ref: '#/components/schemas/RetrieveKnowledgeBasesToolInput'
        - $ref: '#/components/schemas/QueryKnowledgeBaseToolInput'
        - $ref: '#/components/schemas/CurrentDateToolInput'
        - $ref: '#/components/schemas/AdvisorToolInput'
        - $ref: '#/components/schemas/SidekickToolInput'
        - $ref: '#/components/schemas/CodeInterpreterToolInput'
        - $ref: '#/components/schemas/FileSystemToolInput'
        - $ref: '#/components/schemas/HttpToolInput'
        - $ref: '#/components/schemas/CodeToolInput'
        - $ref: '#/components/schemas/FunctionToolInput'
        - $ref: '#/components/schemas/JsonSchemaToolInput'
        - $ref: '#/components/schemas/McpToolInput'
        - $ref: '#/components/schemas/ProviderToolInput'
      description: >-
        Tool configuration for agent create/update operations. Built-in tools
        only require a type, while custom tools (HTTP, Code, Function, JSON
        Schema, MCP) must reference pre-created tools by key or id.
        Provider-prefixed tools (e.g., openai:web_search) are passed through to
        the provider.
      title: Agent Tool Input (CRUD)
      discriminator:
        propertyName: type
        mapping:
          google_search: '#/components/schemas/GoogleSearchToolInput'
          web_scraper: '#/components/schemas/WebScraperToolInput'
          call_sub_agent: '#/components/schemas/CallSubAgentToolInput'
          retrieve_agents: '#/components/schemas/RetrieveAgentsToolInput'
          query_memory_store: '#/components/schemas/QueryMemoryStoreToolInput'
          write_memory_store: '#/components/schemas/WriteMemoryStoreToolInput'
          retrieve_memory_stores: '#/components/schemas/RetrieveMemoryStoresToolInput'
          delete_memory_document: '#/components/schemas/DeleteMemoryDocumentToolInput'
          retrieve_knowledge_bases: '#/components/schemas/RetrieveKnowledgeBasesToolInput'
          query_knowledge_base: '#/components/schemas/QueryKnowledgeBaseToolInput'
          current_date: '#/components/schemas/CurrentDateToolInput'
          advisor: '#/components/schemas/AdvisorToolInput'
          sidekick: '#/components/schemas/SidekickToolInput'
          code_interpreter: '#/components/schemas/CodeInterpreterToolInput'
          file_system: '#/components/schemas/FileSystemToolInput'
          http: '#/components/schemas/HttpToolInput'
          code: '#/components/schemas/CodeToolInput'
          function: '#/components/schemas/FunctionToolInput'
          json_schema: '#/components/schemas/JsonSchemaToolInput'
          mcp: '#/components/schemas/McpToolInput'
    GoogleSearchToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - google_search
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Google search tool
      description: Performs Google searches to retrieve web content
    WebScraperToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - web_scraper
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Web scraper tool
      description: Scrapes and extracts content from web pages
    CallSubAgentToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - call_sub_agent
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Call sub agent tool
      description: Delegates tasks to specialized sub-agents
    RetrieveAgentsToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - retrieve_agents
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Retrieve agents tool
      description: Retrieves available agents in the system
    QueryMemoryStoreToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - query_memory_store
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Query memory store tool
      description: Queries agent memory stores for context
    WriteMemoryStoreToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - write_memory_store
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Write memory store tool
      description: Writes information to agent memory stores
    RetrieveMemoryStoresToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - retrieve_memory_stores
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Retrieve memory stores tool
      description: Lists available memory stores
    DeleteMemoryDocumentToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - delete_memory_document
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Delete memory document tool
      description: Deletes documents from memory stores
    RetrieveKnowledgeBasesToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - retrieve_knowledge_bases
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Retrieve knowledge bases tool
      description: Lists available knowledge bases
    QueryKnowledgeBaseToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - query_knowledge_base
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Query knowledge base tool
      description: Queries knowledge bases for information
    CurrentDateToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - current_date
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Current date tool
      description: Returns the current date and time
    AdvisorToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - advisor
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Advisor tool
      description: Consult a secondary model for advice on the current task
    SidekickToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - sidekick
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Sidekick tool
      description: Delegate a subtask to a secondary model for execution
    CodeInterpreterToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - code_interpreter
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Code interpreter tool
      description: >-
        Executes model-written Python code. Uses provider-native code execution
        when the model supports it, otherwise a secure orq-managed sandbox.
    FileSystemToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - file_system
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          properties:
            file_system_id:
              type: string
              minLength: 1
              description: The id of the file system to attach.
            file_system_key:
              type: string
              minLength: 1
              description: The key of the file system to attach.
            access_mode:
              type: string
              enum:
                - read_only
                - read_write
              default: read_only
              description: >-
                Whether the agent may only read this file system, or also write
                to it.
      required:
        - type
        - configuration
      title: File system tool
      description: >-
        Attaches a persistent file system to the agent, named by file_system_id
        or file_system_key in the configuration. The agent may only read it
        unless the configured access_mode is read_write.
    HttpToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - http
          default: http
          description: HTTP tool type
        key:
          type: string
          description: The key of the pre-created HTTP tool
        id:
          type: string
          description: The ID of the pre-created HTTP tool
        requires_approval:
          type: boolean
          default: false
          description: Whether this tool requires approval before execution
        timeout:
          type: number
          minimum: 1
          maximum: 600
          description: >-
            Tool execution timeout in seconds for this agent (max: 10 minutes).
            Overrides the timeout configured on the tool definition.
      title: HTTP Tool
      description: >-
        Executes HTTP requests to interact with external APIs and web services.
        Must reference a pre-created HTTP tool by key or id.
    CodeToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - code
          default: code
          description: Code execution tool type
        key:
          type: string
          description: The key of the pre-created code tool
        id:
          type: string
          description: The ID of the pre-created code tool
        requires_approval:
          type: boolean
          default: false
          description: Whether this tool requires approval before execution
        timeout:
          type: number
          minimum: 1
          maximum: 120
          description: >-
            Tool execution timeout in seconds for this agent (max: 2 minutes,
            the code sandbox cap). Overrides the timeout configured on the tool
            definition.
      title: Code Execution Tool
      description: >-
        Executes code snippets in a sandboxed environment. Must reference a
        pre-created code tool by key or id.
    FunctionToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - function
          default: function
          description: Function tool type
        key:
          type: string
          description: The key of the pre-created function tool
        id:
          type: string
          description: The ID of the pre-created function tool
        requires_approval:
          type: boolean
          default: false
          description: Whether this tool requires approval before execution
      title: Function Tool
      description: >-
        Calls custom function tools defined in the agent configuration. Must
        reference a pre-created function tool by key or id.
    JsonSchemaToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - json_schema
          default: json_schema
          description: JSON Schema tool type
        key:
          type: string
          description: The key of the pre-created JSON Schema tool
        id:
          type: string
          description: The ID of the pre-created JSON Schema tool
        requires_approval:
          type: boolean
          default: false
          description: Whether this tool requires approval before execution
      title: JSON Schema Tool
      description: >-
        Enforces structured output format using JSON Schema. Must reference a
        pre-created JSON Schema tool by key or id.
    McpToolInput:
      type: object
      properties:
        type:
          type: string
          enum:
            - mcp
          default: mcp
          description: MCP tool type
        key:
          type: string
          description: The key of the parent MCP tool
        id:
          type: string
          description: The ID of the parent MCP tool
        tool_id:
          type: string
          description: The ID of the specific nested tool within the MCP server
        requires_approval:
          type: boolean
          default: false
          description: Whether this tool requires approval before execution
      required:
        - tool_id
      title: MCP Tool
      description: >-
        Executes tools from Model Context Protocol (MCP) servers. Specify the
        parent MCP tool using "key" or "id", and the specific nested tool using
        "tool_id".
      example:
        type: mcp
        id: 01KA84ND5J0SWQMA2Q8HY5WZZZ
        tool_id: 01KXYZ123456789
        requires_approval: false
    ProviderToolInput:
      type: object
      properties:
        type:
          type: string
          pattern: ^[a-z][a-z0-9_-]*:.+$
          description: Provider-prefixed tool type
        key:
          type: string
          description: The key of a pre-created tool
        id:
          type: string
          description: The ID of a pre-created tool
        timeout:
          type: number
          minimum: 1
          maximum: 600
          description: >-
            Tool execution timeout in seconds for this agent (max: 10 minutes).
            Overrides the timeout configured on the tool definition.
        requires_approval:
          type: boolean
          description: Whether this tool requires approval before execution
        configuration:
          type: object
          additionalProperties: {}
          description: >-
            Static tool configuration set at design time. Merged over
            LLM-provided arguments at execution time.
      required:
        - type
      title: Provider Built-in Tool
      description: >-
        Provider-specific built-in tools that are passed through to the
        provider. Must be prefixed with the provider name (e.g.,
        openai:web_search, anthropic:web_search_20250305, google:google_search).
      example:
        type: openai:web_search
  securitySchemes:
    ApiKey:
      type: http
      scheme: bearer
      bearerFormat: JWT

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.