FeaturedMeasurementAugust 202610 min

Measuring agent readiness

Beyond AI visibility

Most tools that claim to measure how ready your website is for AI answer one question: can a model retrieve and read this page. It is a fair question and it is not the one your business runs on. The question that matters is whether an agent your customer sent can finish the job, safely, and at a cost that still makes the product worth selling.

In BFSI the two answers diverge sharply. Discovery is usually the strongest layer we measure, and on its own it moves completion by nothing.

A readable page is the entry condition, not the result.

Visibility measurement grew out of search. It looks at whether your content is crawlable, whether it is cited by assistants, whether the markup is clean, and whether the site publishes the files a model expects to find. All of that is worth doing. None of it tells you what happened after the agent arrived.

The gap shows up as soon as you follow a real journey. An agent can be served a perfectly indexed rates page and still be unable to say which rate applies to a nine month tenure for a customer of that profile. It can find the application form and still lose everything it entered at the identity step. It can submit and receive nothing it can hand back. A visibility score reads all three of those as success.

Readiness is not a property of your content. It is a property of your journey.

Visibility asks
Can a model read this?

Crawlability, citations, markup, published policy files. Measured on the page.

Readiness asks
Can an agent finish this?

Completion, recovery, consent, evidence and cost. Measured on the journey.

Why both matter
One is a gate. One is the outcome.

Invisible means the agent never arrives. Visible without ready means it arrives and leaves empty.

Five layers, and an agent has to clear each to reach the next.

An agent has to clear each layer to reach the next, so the score is not an average of independent factors. A property can be excellent at the first two and return nothing, which is exactly the pattern we keep finding in banking and insurance.

01 · Discoverability

Was it found

Whether the product, its terms and its entry point can be retrieved as facts, and whether your access policy lets a permitted agent in at all.

retrievable · permitted
02 · Understanding

Was it read correctly

Whether the agent can tell which rate, fee, tenure or exclusion governs its case, rather than collecting every number on the page.

conditions resolved
03 · Interaction

Could it act

Identity, conditional forms, one-time passwords, step-up authentication and vendor hand-offs, each measured for whether accumulated state survives it.

state survival · steps
04 · Task completion

Did it finish

The share of benchmark tasks completed, the errors hit, and how many of those errors the agent could recover from without human rescue.

completed · recovered
05 · Outcome

Did it leave with proof

Whether a verifiable result came back: a reference, a status, a receipt the caller can hand to the customer who sent it.

observable · cost per outcome

The five roll up into one reading per journey, and the journey readings roll up into one reading for the property. Both are reported with the trajectory evidence attached, because a score without the run behind it is an opinion with a number on it.

If tuning the agent can move the score, the score measures the agent.

This is the whole methodological argument, and it is why the measurement is built the way it is. Agents vary: the same model, given the same instruction twice, can take different paths, spend different tokens and fail in different places. A measurement that moves with that variance is useless for a regulated release process.

01 · FROZEN FLEET
The probe is versioned and held still

A benchmark fleet and task library, frozen between runs. When the number rises, the property changed, not the probe. It also means the improvement survives the next model release.

02 · REPEATED RUNS
Variance is reported, not averaged away

Each task runs repeatedly. A journey that completes sometimes is a different risk from one that completes reliably, and a spread that wide is itself a finding for risk and operations.

03 · PERMITTED GROUND
Synthetic identities, no live instruments

Runs happen in an environment you nominate, rate-bound and logged, stopping before anything irreversible. If your access policy refuses the agent, that refusal is recorded as a finding.

What comes back is not a grade. It is a reading per layer, the failure point for every incomplete task, the token and step cost of the ones that completed, and a remediation list ranked by score movement per unit of engineering effort.

One reading a board and a backlog can both act on.

The reason to compress five layers into one number is not simplicity, it is shared language. Risk needs to know whether a delegated journey is safe. Engineering needs to know which surface to fix first. Marketing needs to know whether the traffic it is buying can convert when the visitor is software. The same reading answers all three, as long as the evidence travels with it.

Risk
Is a delegated journey safe to allow?

Consent scope, confirmation gates, audit trail and the failure modes an agent can reach. Refusals are recorded rather than routed around.

Engineering
Which surface do we fix first?

A backlog ranked by score movement per unit of effort, with the failing trajectory attached to each item so it can be reproduced locally.

Product
Where does the journey lose value?

The step that ends most attempts, and what it costs in tokens and retries to reach it. Usually not the step teams expect.

Executive
Are we better than last quarter?

The same frozen suite, re-run. One number per journey and one for the property, comparable across releases because the probe did not change.

Two limits worth stating plainly. The reading decays: agentic readiness is not a certification you hold, and a release that reworks a form or swaps a payments vendor can move it materially in one sprint. And it is not a league table. We do not publish client readings, journeys or trajectories, including on our own site.

Get a reading on two of your own journeys.

Two weeks, a permitted environment, synthetic identities, and a w0 score with the trajectory evidence behind it. No change to your stack.

Read next