| build-agent | Design, create, and configure an Orq.ai Agent with tools, instructions, Knowledge Bases, and Memory | SKILL.md |
| build-evaluator | Create validated LLM-as-a-Judge Evaluators following evaluation best practices | SKILL.md |
| recommend-evaluators | Recommend the Evaluators an agent is missing, from its traces or, with no traffic, from its instructions and config; skips the ones already attached and creates each only after approval | SKILL.md |
| evaluator-alignment | Align an existing LLM judge (boolean, categorical, or numeric) to human judgment: measure how often it changes its mind, group the least reliable cases, rewrite the judge prompt, and recreate the Evaluator after approval | SKILL.md |
| analyze-traces | Read production Traces, identify what is failing, build failure taxonomies, and categorize issues | SKILL.md |
| improve-agent | Improve an underperforming Orq.ai Agent or Deployment: rewrite its instructions against a prompting framework, or move a configuration knob, grounded in the error-analysis file analyze-traces writes | SKILL.md |
| run-experiment | Create and run Orq.ai Experiments: compare configurations with specialized agent, conversation, and RAG evaluation | SKILL.md |
| generate-synthetic-dataset | Generate and curate evaluation Datasets: structured generation, quick from description, expansion, and dataset maintenance | SKILL.md |
| invoke-deployment | Invoke Orq.ai Deployments, Agents, and models via the Python SDK or HTTP API, with correct variable substitution, streaming, and identity tracking | SKILL.md |
| setup-observability | Instrument LLM applications with Orq.ai tracing. Covers AI Gateway (zero-code traces) and OpenTelemetry/OpenInference. Guides from framework detection through baseline verification to trace enrichment | SKILL.md |
| compare-agents | Run cross-framework agent comparisons: compare any combination of Orq.ai, LangGraph, CrewAI, OpenAI Agents SDK, or Vercel AI SDK agents head-to-head on the same dataset using evaluatorq | SKILL.md |
| red-team | Run adversarial attacks against deployed agents or static datasets with the evaluatorq red team CLI. Covers OWASP-ASI (agentic: goal hijacking, tool misuse) and OWASP-LLM (model-level: prompt injection, system prompt leakage) | SKILL.md |
| evaluatorq | Write and run evaluatorq evaluation scripts (Python or TypeScript) for a single agent or deployment. Supports custom scorers, dataset-driven runs, and LLM-as-a-Judge Evaluators | SKILL.md |
| simulate-agent | Run multi-turn simulations with evaluatorq primitives (simulate(), generate_and_simulate(), wrap_simulation_agent()): drive an agent under test with a simulated user and score each turn with a built-in judge | SKILL.md |
| manage-skills | List, inspect, create, update, retire, and delete Orq.ai Skills (platform entities). Handles naming rules, template integration ({{skill.key}}), reference scanning, and safe deletion | SKILL.md |
| create-skill | Build or update an agent skill from an API, CLI, or MCP surface: probe the surface, verify each claim, write a tested contract, and register it | SKILL.md |
| shared | Reference bundle the other skills read: the verified orq CLI trace query contract and the run-key preflight. Not invoked on its own | SKILL.md |
| orq-cli | Drive the orq command-line interface: install check, authentication, workspace selection, orq doctor troubleshooting, and read/write commands with JSON output | SKILL.md |