Skip to content
Agent observability

Ship agent changes without guessing

Orq AI traces every agent run, scores it against your eval suites, and holds the release when a regression shows up.

run_a3f7e2b1 PASS
Step Dur ms Tokens Eval
plan 142 312 / 48 pass
tool-call-search 871 88 / 214 pass
tool-call-retrieve 603 64 / 512 skip
synthesize 1240 704 / 381 pass
validate 218 128 / 32 pass
What Orq does in one integration call
Trace any agent run Score every output Block regressions in CI Alert before users feel it
How it works

From deploy to confidence in three steps

Orq AI wraps around your existing agent stack. No rewrites, no new frameworks.

01

Instrument your agent

Add one SDK call. Orq captures every span, tool call, and LLM turn automatically, with latency, token counts, and raw inputs and outputs.

02

Define your eval suite

Write metric thresholds once: coherence, grounding, latency, safety. Orq scores each run against your suite and tracks drift over time.

03

Gate your releases

Connect Orq to your CI pipeline. A commit that drops a metric below threshold blocks the merge, not your on-call at 2 am.

See the full walkthrough
Core capabilities

Everything your team needs to ship agent changes safely

Built for the realities of production LLM agents, not research prototypes.

Distributed tracing

See the full execution graph of every agent run. Every span linked, every tool call timed, every prompt recorded.

Automated eval scoring

LLM-as-judge and deterministic checks run on every trace. Score coherence, grounding, safety and format compliance automatically.

Release gating

Tie eval thresholds to your Git branch. Orq blocks merges when any metric regresses, before it reaches your users.

Regression alerts

Statistical anomaly detection spots metric drift between versions. Get a Slack or webhook alert before users feel the difference.

Version comparison

Side-by-side diff of any two agent versions. Compare latency percentiles, eval distributions, and token budgets at a glance.

Native integrations

Works with LangChain, LlamaIndex, OpenAI Agents SDK, Anthropic SDK, and any Python or TypeScript agent you write yourself.

12ms

Median SDK overhead added per traced run

40+

Pre-built eval metrics ready on day one

1 call

To instrument an agent and start tracing

0 rewrites

Orq wraps your existing stack, not the other way around

Built for production

The problem Orq solves

Agent quality is invisible until something breaks. Orq makes it visible before the deploy.

The problem: silent regression

A prompt change ships. Grounding drops 8 points on one document format. No alert fires. A customer notices three weeks later. Orq exists to make that three-week gap zero.

What existing tools miss

Distributed tracing systems were not built for the LLM call shape. Eval frameworks run offline in notebooks. Nothing connects tracing to evals to CI in one place. That gap is what we close.

What Orq gives you instead

A trace of every run, a score on every output, and a gate in CI that blocks the merge when a metric regresses. All connected, all running automatically after one integration call.

Stop guessing what changed in your agent

Get a complete trace, eval score, and release gate running in under an hour. Free plan includes 5,000 traces per month, no card required.