Everything AgentSpeed watches, in one place.
Run + span traces, cost and latency, failures, journey canaries, alerts, and public status pages: the full picture of how your agents behave in production.
Metadata-first. We never store your prompts or model outputs.
| Status | Started | Latency | Model | Tokens | Cost | Error |
|---|---|---|---|---|---|---|
| running | just now | — | claude-opus-4-8 | — | — | |
| succeeded | 2m ago | 3.1s | claude-opus-4-8 | 1,540 | $0.021 | |
| succeeded | 4m ago | 2.7s | claude-sonnet-4-6 | 1,120 | $0.009 | |
| failed | 6m ago | 9.0s | gpt-4o | 2,310 | $0.038 | RateLimitError |
| succeeded | 8m ago | 3.4s | claude-opus-4-8 | 1,705 | $0.024 | |
| timeout | 11m ago | 30.0s | claude-opus-4-8 | 980 | $0.014 | TimeoutError |
| succeeded | 13m ago | 2.9s | claude-sonnet-4-6 | 1,260 | $0.010 |
Illustrative data. Click around the real demo →
Runs, spans, cost & latency
Every run as a timeline of LLM calls, tool calls, and retrievals, with status, duration, and tokens per step.
p50 / p95 latency per agent over rolling windows, computed in-database, not sampled estimates.
Input/output tokens and dollar cost rolled up by agent, model, and run. Catch the expensive regressions.
Success / failure / timeout rates with the error type on every failed run, so you see what broke and why.
Attach a quality score to each run and we aggregate it per agent and per model. Spot the silent slide where success stays high but answers got worse, all without us seeing a prompt or output.
Test the flows your users actually take
Write the flow as plain steps. “Open /pricing, start a trial, land on the dashboard.”
On your schedule, we probe the live site and verify each step is still supported by what it serves.
Each run is saved as a full trace: every step, how long it took, pass or fail.
When a step fails, your alert links straight to the step that broke.
Hit a CAPTCHA or bot wall? We tell you it blocked us. We never try to sneak past it.
Know before your users do
Set thresholds per agent or per project. A 30-minute cooldown keeps a flapping metric from turning into an alert storm.
Every alert links straight to the run (or the failing journey step) that tripped it, so you start debugging on the trace, not in a dashboard hunt.
Prove uptime to your customers
Publish a hosted status page with 30-day uptime bars per agent: a public, always-on signal of reliability you can hand to customers and stakeholders.