agentspeed.
Research

What we measure, and what we have not measured yet

A score is only worth what its method is worth. Everything below is either a published measurement you can check, or a method we committed to in public before collecting any data. This page keeps those two apart, because a list that mixed them would read as more evidence than we hold.


Published — 2 with results you can read
How often the score is wrongPublished

The measured error rate of every scored check against a frozen, hand-labelled adversarial corpus.

127 observations, one false positive, no false negatives — and no per-check rate is publishable yet, because no check has enough labelled cases to carry one.

The scoring rubricPublished

The versioned specification behind every score: checks, weights, grade bands, and the rules for changing them.

Published and frozen per version; a published rubric is never edited in place.


No results yet — 3

Method and machinery, not findings. Where an entry says registered, the method was written down in public before any data was collected — which is what makes the eventual result worth anything, since a study designed after seeing the numbers can always be made to say something. Those pre-registrations include the conditions under which we would report a null result, or decline to report at all.

The Agent Readiness IndexBuilt · not yet published

A frozen, citable benchmark of how ready the public web is for AI agents, computed from the scored corpus.

The generator is built and no edition has been published yet.

The Agent Citation StudyRegistered · no data yet

Does higher agent-readiness correspond to being cited more often by AI assistants? Pre-registered with seven gates that can refuse to report.

Registered before any observation was collected. Wave 1 has not run, so there is nothing to report — and the gates are written so that a null or unusable result is reported as one.

Pre-registration: docs/research/AGENT-CITATION-STUDY.md

The Vibe Coding IndexRegistered · no data yet

How agent-ready are sites built on AI app builders? A population study of sites on each builder’s default subdomain.

Registered, with the sampling frame captured from Certificate Transparency. No site has been scanned yet, so there are no findings — distributions only when there are, and never individual site names.

Pre-registration: docs/research/VIBE-CODING-INDEX.md


Nothing here claims that agent-readiness causes anything. Whether a readier site is cited more often by AI assistants is an open question, it is the thing the Agent Citation Study was registered to answer, and it has not run. 5 artifacts are listed; the status on each is the one fact this page most needs to get right.

Research and evidence · AgentSpeed