> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Invoke a Custom Evaluator

> Runs an evaluator that already exists in the workspace. Accepts either a conversation or the structured input and output fields; when both are present the conversation wins.

<Note>
  **Related guide**: Evaluators guide. See the [Evaluators guide](/ai-studio/optimize/evaluators) for a walkthrough with examples.
</Note>


## OpenAPI

````yaml post /v3/evaluators/{id}/invoke
openapi: 3.1.0
info:
  title: orq.ai API
  version: '2.0'
  description: orq.ai API documentation
servers:
  - url: https://my.orq.ai
security:
  - ApiKey: []
tags:
  - name: Chunking
    description: Split text into smaller chunks for retrieval and generation workflows.
  - name: File Systems
    description: >-
      Create and manage persistent file systems that agents and MCP clients read
      from and write to.
  - name: Knowledge Bases
    description: Create and manage knowledge bases used by agents and retrieval workflows.
  - name: Memory Stores
    description: Create and manage memory stores, memories, and memory documents.
  - name: Evals
    description: Run an evaluator against a conversation and its result
  - name: Logs
    description: >-
      OpenTelemetry log query API. Search, filter, aggregate, and facet log
      records ingested via OTLP.
  - name: Reporting
    description: >-
      GenAI reporting API over canonical analytics rollups. Accepts a metric
      name, time range, grain, group-by, and filters; returns a typed time
      series and optional totals.
  - name: Traces
    description: >-
      Query and inspect ingested trace data: search trace summaries, aggregate
      metrics, and read individual traces and their spans.
  - description: List models available through the AI Router.
    name: Models
  - name: Policies
  - name: Alerts
    description: >-
      Alerts evaluate a Reporting API metric on a fixed interval and fire
      notifications through notifiers when the value breaches a threshold. Each
      breach opens a trigger that tracks the incident until the value recovers.
  - name: Annotation Queues
    description: Annotation queues collect spans for human review.
  - name: API keys
    description: >-
      API keys authenticate programmatic access to the workspace. They expose
      opaque tokens, per-domain access grants, and budget and rate-limit
      constraints.
  - name: Audit Logs
    description: Audit logs record workspace entity changes and access-relevant events.
  - name: Budgets
    description: >-
      Budgets govern spend, token usage, and request rate across six scopes:
      workspace, project, identity, API key, provider, and model. Every
      applicable budget is enforced, and the most restrictive limit applies per
      dimension.
  - name: Files
    description: File upload and retrieval operations.
  - name: Guardrail Rules
    description: >-
      Guardrail Rules conditionally enforce evaluators and plugins for AI
      Gateway traffic. Rules may be scoped to a project or the whole workspace.
  - name: Hub
    description: Hub items are reusable templates available to a workspace.
  - name: Identities
    description: >-
      Identities represent end users from your system for usage and engagement
      tracking.
  - name: Management keys
    description: >-
      Management keys are workspace-scoped credentials that authenticate
      programmatic access to workspace administration surfaces (API keys,
      budgets). Unlike project-scoped API keys, a management key always operates
      at the workspace level.
  - name: MCP Gateway
    description: >-
      Register upstream MCP servers, discover and sync their tools, and assemble
      gateways that expose a curated tool surface to MCP clients.
  - name: Model Catalog
    description: >-
      Browse the orq.ai model catalog: every model orq offers, across every
      provider, with pricing, capabilities and benchmark data. List endpoints
      only return models that are not deprecated. This API is public, requires
      no authentication, and is rate limited to 120 requests per minute per IP.
      Responses carry a 5-minute cache-control max-age.
  - name: Notifiers
    description: Notifier destinations used to send delivery and workflow notifications.
  - name: Projects
    description: Projects organize resources within a workspace
  - name: Routing Rules
    description: >-
      Routing Rules conditionally select models and enforce request plugins for
      AI Gateway traffic. Rules are evaluated by ascending priority and may be
      scoped to a project or the whole workspace.
  - name: Threads
    description: Threads group related trace invocations and their aggregate usage
  - name: Skills
    description: >-
      Skills are modular instructions you can use to codify processes and
      conventions
  - name: Smart Routers
    description: >-
      Create and manage workspace Smart Routers. A Smart Router selects a model
      from an eligible pool for each request according to a quality, balanced,
      or cost profile.
  - name: Webhooks
    description: >-
      Create and manage webhooks that deliver workspace events to external HTTPS
      endpoints.
  - name: Workspaces
    description: >-
      A workspace is the tenant. Create is called from a user session during
      onboarding; Get, List, and Update are the public management surface.
  - name: Workspace Security
    description: >-
      Workspace-level domain verification and IP allowlist controls. These
      operations are restricted to workspace administrators.
  - name: Workspace Settings
    description: >-
      Workspace-level settings managed with a workspace credential. A workspace
      is the tenant, so these settings are a singleton — there is nothing to
      create or delete, only read and update.
  - name: Responses
  - description: Run agents on a cron cadence. Minimum firing interval is 1 hour.
    name: Agent Schedules
  - name: Embeddings
  - name: Telemetry
    description: >-
      Unified query envelope for traces, metrics, and logs. One request shape,
      one filter dialect, and one response shape per source, validated by a
      per-source registry.
  - description: Beta. Run typed classification questions against a classify model.
    name: Classify
  - description: Search Gateway with managed credits or BYOK.
    name: Web Search
externalDocs:
  url: https://docs.orq.ai
  description: orq.ai Documentation
paths:
  /v3/evaluators/{id}/invoke:
    post:
      tags:
        - Evals
      summary: Invoke a Custom Evaluator
      description: >-
        Runs an evaluator that already exists in the workspace. Accepts either a
        conversation or the structured input and output fields; when both are
        present the conversation wins.
      operationId: InvokeEval
      parameters:
        - name: id
          in: path
          description: Accepts a bare id, `id@version`, or `id@environment`.
          required: true
          schema:
            type: string
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/InvokeEvaluatorRequest'
        required: true
      responses:
        '200':
          description: OK
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/EvaluationResult'
      x-code-samples:
        - lang: curl
          label: Core - Run an evaluator
          source: >
            curl
            'https://my.orq.ai/v3/evaluators/01KT1FCSA8N3YD1K8YBPVTAV9E/invoke'
            \
              --header "Authorization: Bearer $ORQ_API_KEY" \
              --header 'Content-Type: application/json' \
              --data-raw '{
                "context": {
                  "input": {
                    "user_query": "What is the capital of France?",
                    "expected_output": "Paris",
                    "retrievals": ["The capital of France is Paris."]
                  },
                  "output": {
                    "response": "The capital of France is Paris."
                  },
                  "variables": {
                    "tone": "formal"
                  }
                }
              }'
        - lang: python
          label: Python - Run an evaluator
          source: >
            import os

            from orq_ai_sdk import Orq


            client = Orq(api_key=os.environ["ORQ_API_KEY"])


            evaluation = client.evals.invoke(
                id="01KT1FCSA8N3YD1K8YBPVTAV9E",
                context={
                    "input": {
                        "user_query": "What is the capital of France?",
                        "expected_output": "Paris",
                        "retrievals": ["The capital of France is Paris."],
                    },
                    "output": {"response": "The capital of France is Paris."},
                    "variables": {"tone": "formal"},
                },
            )


            # `passed` is the guardrail's decision when the evaluator has one,
            and the

            # grader's own judgement otherwise. `value` carries the verdict
            itself,

            # whose type follows the evaluator: a boolean, a score, a label.

            print(evaluation.passed, evaluation.value)
        - lang: typescript
          label: Node.js - Run an evaluator
          source: |
            import { Orq } from '@orq-ai/node';

            const orq = new Orq({ apiKey: process.env.ORQ_API_KEY });

            const evaluation = await orq.evals.invoke({
              id: '01KT1FCSA8N3YD1K8YBPVTAV9E',
              invokeEvaluatorRequest: {
                context: {
                  input: {
                    user_query: 'What is the capital of France?',
                    expected_output: 'Paris',
                    retrievals: ['The capital of France is Paris.'],
                  },
                  output: { response: 'The capital of France is Paris.' },
                  variables: { tone: 'formal' },
                },
              },
            });

            console.log(evaluation.passed, evaluation.value);
        - lang: curl
          label: Core - Grade a conversation instead of a single turn
          source: >
            curl
            'https://my.orq.ai/v3/evaluators/01KT1FCSA8N3YD1K8YBPVTAV9E/invoke'
            \
              --header "Authorization: Bearer $ORQ_API_KEY" \
              --header 'Content-Type: application/json' \
              --data-raw '{
                "context": {
                  "messages": [
                    {"role": "user", "content": "What is the capital of France?"},
                    {"role": "assistant", "content": "The capital of France is Paris."}
                  ],
                  "input": {
                    "expected_output": "Paris"
                  }
                }
              }'
components:
  schemas:
    InvokeEvaluatorRequest:
      type: object
      properties:
        context:
          $ref: '#/components/schemas/EvaluationContext'
        model:
          type: string
          description: |-
            Model to grade with, as a catalog id such as "openai/gpt-4o".

             Only meaningful for a hub template of type llm_eval or ragas, which has no
             model of its own. A stored evaluator uses the model on its own definition
             and ignores this.
        query:
          type: string
          description: Latest user message. Folds into `context.input.user_query`.
        output:
          type: string
          description: |-
            The generated response from the model. Folds into
             `context.output.response`.
        reference:
          type: string
          description: |-
            The reference used to compare the output. Folds into
             `context.input.expected_output`.
        retrievals:
          type: array
          items:
            type: string
          description: Knowledge base retrievals. Folds into `context.input.retrievals`.
        messages:
          type: array
          items:
            type: object
            additionalProperties: true
          description: |-
            The conversation that produced the output. Folds into
             `context.messages`.
        variables:
          type: object
          additionalProperties:
            $ref: '#/components/schemas/JsonValue'
          description: |-
            Template variables for evaluator prompt substitution. Folds into
             `context.variables`.
      description: >-
        Accepts two shapes. `context` names its fields after the template
        variables
         they feed and is the one to use; the flat fields below are folded into
         `context` when it is absent. Setting `context` wins.
    EvaluationResult:
      type: object
      properties:
        type:
          type: string
          description: |-
            Discriminator for the verdict shape: "string", "number", "boolean",
             "string_array", "rouge_n", "bert_score", "llm_evaluator", "http_eval".
        value:
          allOf:
            - $ref: '#/components/schemas/JsonValue'
          description: |-
            The verdict. Dynamic by design — a boolean pass, a numeric score, a
             categorical label, a list of labels and the nested rouge_n object all
             arrive here.
        trace_id:
          type: string
          description: |-
            Trace reference of the evaluator's own span. Optional so an absent
             reference is omitted rather than emitted as an empty string.
        span_id:
          type: string
        evaluator_id:
          type: string
        status:
          type: string
          description: >-
            How the run ended, as distinct from `passed`: "passed",
            "condition_failed",
             "failed" or "timed_out". A string, not an enum, because the engine owns the
             vocabulary.
        passed:
          type: boolean
          description: >-
            The guardrail's decision when the evaluator has one, the grader's
            own
             judgement otherwise. Always present, so read `guardrail_config` to detect
             a guardrail, not this.
        explanation:
          type: string
        categories:
          type: array
          items:
            type: string
          description: >-
            Set by the classifying graders (moderation, PII, secret detection),
            which
             report which categories tripped rather than a single verdict.
        confidence:
          type: number
          description: >-
            Optional so an absent score is omitted rather than reported as 0.0,
            which
             a consumer would read as maximum uncertainty.
          format: double
      description: >-
        The verdict. Its shape is fixed so existing consumers read the same
        JSON.
    EvaluationContext:
      type: object
      properties:
        messages:
          type: array
          items:
            type: object
            additionalProperties: true
        input:
          $ref: '#/components/schemas/StructuredInput'
        output:
          $ref: '#/components/schemas/StructuredOutput'
        variables:
          type: object
          additionalProperties:
            $ref: '#/components/schemas/JsonValue'
      description: |-
        The data to grade. When `messages` is present it is the conversation and
         `input.user_query` is ignored; `output.response` is appended only when the
         conversation carries no assistant turn.
    JsonValue:
      description: >-
        Represents a dynamically typed value which can be either null, a number,
        a string, a boolean, a recursive struct value, or a list of values.
    StructuredInput:
      type: object
      properties:
        system_instructions:
          type: string
        user_query:
          type: string
        retrievals:
          type: array
          items:
            type: string
        expected_output:
          type: string
      description: >-
        StructuredInput names its fields after the template variables they feed,
        so
         input.user_query in a prompt is user_query here.
    StructuredOutput:
      type: object
      properties:
        response:
          type: string
        tools_called:
          type: array
          items:
            $ref: '#/components/schemas/StructuredToolCall'
    StructuredToolCall:
      type: object
      properties:
        name:
          type: string
        arguments:
          type: string
        output:
          type: string
      description: >-
        StructuredToolCall mirrors one tool invocation and its result. Name is
        the
         only field the grader requires; a call without one is dropped.
  securitySchemes:
    ApiKey:
      type: http
      scheme: bearer
      bearerFormat: JWT

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.