scanned 2026-08-23 17:00 UTC · cached · refresh in 0h · rubric 2026.11.0
○Unverified public scan
Built with Next.js
Good
AgentSpeed tests how easily automated agents can discover, understand, and act on weather.com. The score weights five categories of machine-readability; the ticks on the arc are the 37 individual checks behind it.
Send this to whoever edits the site. The link shows the score and grade as a preview card, and it stays current: it re-reads the latest scan rather than freezing today’s number.
Top fixes
Checks tagged “emerging” are 2026 agent-protocol standards most of the web hasn’t adopted yet. Adopting early is an edge, not a defect.
warnrobots.txt allows AI agents+1 ptDiscoverability
2 access agent(s) blocked: PerplexityBot, Amazonbot. These keep your site out of AI answers. 11 training-data crawler(s) blocked (GPTBot, ClaudeBot, anthropic-ai, Google-Extended, Applebot-Extended, Bytespider, CCBot, cohere-ai, Diffbot, Meta-ExternalAgent, FacebookBot), a licensing choice, not scored.
Show me exactly what to do
What this means
robots.txt is a small text file at yoursite.com/robots.txt where you tell visiting robots what they may read. Many sites carry an old rule that blocks ALL robots, which now also blocks the AI assistants your customers ask for recommendations.
Why it matters
ChatGPT, Claude, and Perplexity each send a named crawler (GPTBot, ClaudeBot, PerplexityBot). If your robots.txt turns them away, they cannot read your pages. So when a customer asks an assistant about your product, the assistant answers from what everyone ELSE says about you, or recommends a competitor it could read.
Do this
Next.js: Allow GPTBot, ClaudeBot, and PerplexityBot in robots.txt (or remove the blanket Disallow).
Don’t lose this report
Get it in your inbox now, plus a heads-up whenever weather.com’s agent-readiness score changes. Free, unsubscribe anytime.
Prefer we just watch it for you? Monitor rescans weekly and emails what changed, $29/mo. Start monitoring →
Fix it
Generated, ready-to-ship files for the gaps above. Copy them, or download and drop them into your repo.
Allow AI agents in robots.txt
Merge these groups into your robots.txt at the site root.
This scan graded one page. Your site is more than one page.
Weakest on this scan: discoverability, 73/100 — the signals agents use to find and navigate this site are thin or missing.
This graded your homepage. Your pricing, docs, and checkout live on other pages: the ones assistants actually quote when they recommend you. A 86/100 here means agents are already missing things on the page you polish most. The $29 Deep Scan grades your homepage plus up to 10 more pages we discover, against all 37 checks, then emails a PDF with every failing check and its exact fix, ranked by what costs you the most visibility. One-time, delivered in minutes, no subscription.
Right now, only you know this score. Your buyers ask AI before they buy, and certification is how you show them, and anyone comparing you to a competitor, that assistants can actually read, quote, and use your site.
A public verification URL you can drop into a sales deck, a procurement questionnaire, or a partner review. Anyone can check it, dated and signed.
An embeddable badge for your site, the visible version of a property competitors can’t claim without earning the score.
Weekly re-checks for a full year. A deploy that breaks agent access gets caught before it costs you recommendations, and before the certificate would lapse.
Sites under 70 can’t buy this at any price. That’s what makes displaying it mean something.
Get this report by email, then a heads-up when weather.com’s agent-readiness actually changes: a check regressing, or the composite dropping.
Two things never trigger an email: a score change caused by us publishing a new rubric, because that is a different instrument rather than a change to your site, and movement explained only by timing-derived checks, which shift run to run on a site nobody touched.
Want it on your own schedule, with Slack or a webhook instead of email? Score-drift monitoring is on every paid plan.
This score says whether an AI agent could read weather.com. AI traffic analytics says whether one actually came, and which pages turned it away.
Full breakdown
19 pass3 warn5 fail10 skip
Discoverability
pass
robots.txt present
A robots.txt is reachable at the site root.
100/100
warn
robots.txt allows AI agents
2 access agent(s) blocked: PerplexityBot, Amazonbot. These keep your site out of AI answers. 11 training-data crawler(s) blocked (GPTBot, ClaudeBot, anthropic-ai, Google-Extended, Applebot-Extended, Bytespider, CCBot, cohere-ai, Diffbot, Meta-ExternalAgent, FacebookBot), a licensing choice, not scored.
78/100
fail
Content-Signal directivesemerging
No Content-Signal directives. Add e.g. `Content-Signal: ai-train=no, ai-summarize=yes` to declare granular AI usage policy beyond binary allow/disallow.
0/100
pass
llms.txt present
Found /llms.txt (8776 bytes), H1 + at least one link present. Not scored.
100/100
pass
Sitemap present
sitemap.xml reachable and referenced from robots.txt.
100/100
fail
Link response headers
No HTTP Link: response headers. Emit canonical/alternate/describedby relations so HEAD-only or stream-rendering agents get them without parsing HTML.
0/100
pass
Canonical URL
Canonical points to self: https://weather.com/.
100/100
skip
MCP server cardemerging
Agent-protocol discovery was not probed for this scan.
—
skip
OAuth authorization metadataemerging
Agent-protocol discovery was not probed for this scan.
—
skip
OAuth resource metadataemerging
Agent-protocol discovery was not probed for this scan.
—
skip
API catalog / OpenAPIemerging
Agent-protocol discovery was not probed for this scan.
—
skip
Web Bot Authemerging
Agent-protocol discovery was not probed for this scan.
—
Readability
fail
Markdown negotiationemerging
Accept: text/markdown returns HTML, not markdown. Serve a markdown variant of primary content when requested; agents summarize and cite it more reliably.
0/100
fail
Text-to-markup ratio
Visible text is 0.2% of HTML weight (850/356477 bytes). 76% of the document is inline script, 24% is markup.
2/100
pass
Content without JavaScript
Primary content is present in the initial HTML — agents without JS can read it. (Measured statically; a rendered comparison was not run because the static reading sufficed.)
90/100
pass
Heading hierarchy
Hierarchy is well-formed (1 H1, no skipped levels, 2 headings total).
100/100
pass
Cookie wall blocks content
No common consent-modal markers detected.
100/100
pass
Declared language
Declared lang="en-US".
100/100
warn
Page title
<title> is 116 characters; keep it under 70 so it isn't truncated.
60/100
pass
Meta description
A well-sized meta description is present (149 chars).
Page intent unclear; no specific schema.org type expected.
100/100
Coherence: do your channels agree?
skip
JSON-LD price matches the visible priceemerging
No price declared in JSON-LD, so there is nothing to cross-check. (Whether a price SHOULD be declared is the structured-data checks’ question.)
—
pass
JSON-LD name appears on the pageemerging
Declared entity name "The Weather Channel" appears on the visible page.
100/100
pass
Declared language matches the contentemerging
Declared lang="en-US" matches the detected content language (en).
100/100
Actionability
pass
Paywall / login wall
No paywall or login-wall detected on the landing URL.
100/100
skip
Agent Skills manifestemerging
Agent-protocol discovery was not probed for this scan.
—
fail
WebMCP actionsemerging
No WebMCP detected. On pages with first-class actions (cart, support, account), embed a WebMCP server so on-page agents invoke tools directly.
0/100
skip
Commerce protocol manifestemerging
Agent-protocol discovery was not probed for this scan.
—
pass
Primary action reachable
A primary offering and a price or call to action were both identifiable from the served HTML, so an agent can tell what this page sells and what it costs without running scripts.
100/100
pass
Reachability
HTTP 200.
100/100
pass
Primary action is clickableemerging
1 clickable element(s) carry an action an agent can take: <a> announcing "sign up". An agent driving the page has something real to click.
100/100
skip
Filters reachable by URLemerging
This page does not offer to filter or sort a list, so there is nothing to address by URL. Pages without facets are not penalised here.
—
Performance
pass
Time to first byte
TTFB 150ms. Healthy.
100/100
skip
Full render time
Full-render time unavailable (browser pass skipped or failed).
—
pass
Page weight
Initial document is a lean 348 KB.
100/100
Beyond the scan
Running AI agents of your own? AgentSpeed also monitors them in production: every run, its latency, cost, and failures, with alerts when something breaks. Free for 10,000 events a month, no credit card.
Run reports like this for every client: white-label PDFs and weekly monitoring for up to 15 domains, $149/mo. AgentSpeed for Agencies →
Methodology & limitations
This score is a measurement of machine-readability for automated agents under the published rubric 2026.11.0. AgentSpeed does not assess security, legitimacy, privacy, or financial trust.
Open yoursite.com/robots.txt in a browser to see what you have today.
Look for "User-agent: *" followed by "Disallow: /". That single pair blocks everything, AI agents included.
Add an explicit allow block for each AI agent you want (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot), placed after the wildcard rule so it wins.
The generated robots.txt in the Fix it section below has the exact lines for this site. Copy it, or merge the blocks into your existing file.
→ https://weather.com/robots.txt
Paste into Claude Code, Cursor, or any assistant with your repository open. It carries this finding, the steps, and the shortcuts to avoid.
Or copy it written for your tool:
What to check
After deploying, open yoursite.com/robots.txt and confirm the new blocks are live, not cached.
Hit Re-scan above: "robots.txt allows AI agents" should flip to PASS.
The free robots.txt checker at /tools/robots-txt-checker shows per-agent allow/block for your live file.
What to avoid
Don’t delete rules you added deliberately. If you block a specific scraper for a reason, keep that block; just don’t let a blanket rule catch the assistants too.
Allowing crawlers you actually want blocked just to raise a score is backwards. The score measures readiness for agents you WANT; decide the policy first, then express it precisely.
Don’t serve a different robots.txt to different visitors. Inconsistent answers get you treated as unreliable by every crawler.
No HTTP Link: response headers. Emit canonical/alternate/describedby relations so HEAD-only or stream-rendering agents get them without parsing HTML.
Show me exactly what to do
What this means
HTTP responses can carry a Link header: the same canonical and alternate-language information your HTML declares, but delivered in the response envelope itself. Agents that only send a HEAD request, or that decide what to do while the page is still streaming, read the header without parsing any HTML. Yours is missing (or carries none of the relations agents use: canonical, alternate, describedby).
Why it matters
An agent triaging fifty URLs doesn’t want to download and parse fifty pages to learn which are duplicates of which. The Link header answers at the cheapest possible layer. Sites that provide it get correctly de-duplicated and correctly language-routed even by the most minimal fetchers.
Do this
Emit a Link header alongside each page, mirroring what your HTML head already declares. Example: Link: <https://yoursite.com/page>; rel="canonical".
Most stacks set this in one place: Next.js headers() config, an nginx add_header line, or a CDN response-header rule.
If you serve translations, add rel="alternate" entries with hreflang, matching your HTML.
Paste into Claude Code, Cursor, or any assistant with your repository open. It carries this finding, the steps, and the shortcuts to avoid.
Or copy it written for your tool:
What to check
Run: curl -sI https://yoursite.com | grep -i "^link:" and confirm the relations appear.
Hit Re-scan above: "Link response headers" flips to PASS.
What to avoid
The header must AGREE with the HTML. A Link header pointing one place while the in-page canonical points another gives machines two contradictory answers, which is worse than one missing answer.
Don’t inject the same site-wide canonical on every page via a blanket CDN rule; each page names its own clean URL, exactly as in the HTML tag.
Visible text is 0.2% of HTML weight (850/356477 bytes). 76% of the document is inline script, 24% is markup.
Show me exactly what to do
What this means
Of everything your server sends for this page, very little is actual readable text. The rest is code, styling, and markup wrapper. Agents fetched a lot of bytes and found few words.
Why it matters
Agents work with retrieval budgets. A page that is 2% text either gets skimmed (and mis-summarised) or skipped. More signal per byte means more of your actual message survives into the agent’s summary.
Do this
Confirm your real copy is server-rendered. See the JavaScript check; these two usually fail together.
Cut boilerplate wrappers: deeply nested divs, inline SVG logos repeated per section, base64 images inlined into HTML.
Move large inline styles and scripts into external files.
Check how much of the document is inline script: your report now says. If most of it is, that is your framework serializing state for hydration, often a second copy of text already in the HTML. Send only the props a component cannot recompute, and render statically where a page does not need to hydrate.
Paste into Claude Code, Cursor, or any assistant with your repository open. It carries this finding, the steps, and the shortcuts to avoid.
Or copy it written for your tool:
What to check
Re-scan: "Text-to-markup ratio" improves, and "Page weight" often improves alongside.
What to avoid
Don’t pad the page with keyword text to inflate the ratio. The ratio is a proxy for substance, and stuffing is the opposite of substance.
<title> is 116 characters; keep it under 70 so it isn't truncated.
Show me exactly what to do
What this means
The <title> tag is the page’s name: the text in the browser tab. Yours is missing, empty, or generic (e.g. just "Home"), so machines have no usable name for this page.
Why it matters
The title is the single most-quoted identifier for a page; agents cite it verbatim when referencing you. A missing one gets replaced with a guess assembled from headings, which amounts to your page’s name being written by someone else.
Do this
Give every page a unique, descriptive title in the shape "What the page is | Brand", under about 60 characters.
Set it in your framework’s metadata mechanism (Next.js Metadata API, WordPress SEO plugin, Shopify page settings) so it is per-page, not hardcoded site-wide.
→ <head>
Paste into Claude Code, Cursor, or any assistant with your repository open. It carries this finding, the steps, and the shortcuts to avoid.
Or copy it written for your tool:
What to check
Each page shows its own title in the browser tab.
Re-scan: "Page title" flips to PASS.
What to avoid
Don’t stuff keywords. The title is quoted verbatim in agent answers, and a keyword string reads as spam exactly where customers see it.
Don’t reuse one title across all pages; identical names make pages indistinguishable in citations.
Your page HAS structured data, but it is malformed: broken JSON syntax, or types and fields that don’t exist in the schema.org vocabulary. Machines can see the block but can’t parse it.
Why it matters
Invalid JSON-LD is worse than it looks. The agent spends its attention budget on the block, fails to parse it, and falls back to guessing from prose anyway. You pay the cost of structured data without getting the benefit.
Do this
Paste the failing page into validator.schema.org. It names the exact line and field that breaks.
The most common breaks: a trailing comma (invalid JSON), a typo’d type name, or a price written as "$79" instead of "79" with a separate priceCurrency.
Fix in place, redeploy, revalidate.
Paste into Claude Code, Cursor, or any assistant with your repository open. It carries this finding, the steps, and the shortcuts to avoid.
Or copy it written for your tool:
What to check
validator.schema.org reports zero errors for the page.
Re-scan: "Structured data validates" flips to PASS.
What to avoid
Don’t delete the block to make the error disappear. That trades "invalid" for "absent" and fails the presence check instead; fix it.
Don’t hand-edit generated JSON-LD in the page. Fix the template or plugin that generates it, or the error returns on the next publish.
No Content-Signal directives. Add e.g. `Content-Signal: ai-train=no, ai-summarize=yes` to declare granular AI usage policy beyond binary allow/disallow.
Show me exactly what to do
What this means
Content-Signal is a newer robots.txt directive (pushed by Cloudflare) that lets you state a granular AI policy: for example "don’t train on my content, but summarising it in answers is fine". Classic robots.txt only offers all-or-nothing per crawler; this adds the middle ground. We looked for a Content-Signal line in your robots.txt and response headers and found none.
Why it matters
Without a granular signal, crawlers infer your intent from blunt allow/block rules, and sites often block more than they mean to just to avoid training use. A declared signal lets you keep answer visibility (summaries, citations) while opting out of what you object to. This is an emerging standard: honoring is voluntary and adoption is early, which is why it is flagged as emerging and weighs little.
Do this
Decide your actual policy first: are you fine with AI training on your content? With summarisation in answers? With search indexing?
Express it as one line in robots.txt, for example: Content-Signal: ai-train=no, ai-summarize=yes.
Keep your existing User-agent rules; the signal adds nuance on top, it doesn’t replace them.
Paste into Claude Code, Cursor, or any assistant with your repository open. It carries this finding, the steps, and the shortcuts to avoid.
Or copy it written for your tool:
What to check
Open yoursite.com/robots.txt and confirm the Content-Signal line is live.
Hit Re-scan above: "Content-Signal directives" flips to PASS.
What to avoid
Don’t declare signals that contradict your User-agent rules, such as ai-summarize=yes while blocking every AI crawler; conflicting instructions get you treated as unreliable.
Don’t add the line just for the score. It is a public policy statement; say what you mean, because compliant crawlers will act on it.