The measureOutcome-basedGovernance-gated

Webzero Agent Readiness Index

The Webzero Agent Readiness Index (WARI) is one number for a question no existing metric answers: when a customer sends an agent to your bank, insurer or platform, does it leave with the outcome it was sent for, safely, and at a cost that makes sense?

Visibility scores tell you an agent could see your property. The w0 score tells you whether it left with an outcome.
The Webzero Agent Readiness Index began as Bridge AI research and is maintained by Webzero, formerly Bridge AI.
14.4%
Frontier-model task success on WebArena, the public benchmark of real multi-step web journeys.
78.2%
Human success on the identical task set. The gap is not intelligence, it is the infrastructure.
63.8
Points of addressable gap. The property is the only side of it you control.

Four questions, in the order that matters.

Earlier readiness frameworks averaged a checklist, so a property could score well on preparation while completing nothing. The w0 score only counts what an agent actually finished. Everything else explains the result, constrains it, or prices it.

01
Can it find it?
diagnostic · discovery
Crawl policy, structured data, semantic clarity, permission to be there at all. In BFSI this is almost always the strongest layer, and on its own it moves the score by nothing. Being findable is table stakes, not readiness.
Role in score
explains
02
Can it do it?
the score itself
Measured completions of real, value-weighted customer tasks, deposit booking, claims intake, eligibility and servicing, under clean and deliberately perturbed conditions. This layer, and only this layer, produces the number.
Role in score
drives the score
03
Can it do it safely?
governance gate
Scoped consent, prompt-injection resistance, audit trail, and human confirmation before anything irreversible. A gate does not contribute points; it decides whether the points survive.
Role in score
gates the score
04
And is it worth it?
economic modifier
Cost and latency per completed outcome. A journey that succeeds at ten times the token cost is a different commercial result from one that succeeds, and at agent volumes, the difference is the business case.
Role in score
adjusts the score

Three rules make the number usable by a risk committee.

The scoring model, weightings and benchmark task library are Webzero methodology and are shared under engagement. The rules that govern how the number behaves are not a secret, they are the reason it can be trusted.

01 · NON-COMPENSATORY
A failed governance gate takes the score to zero.

Consent, injection resistance, audit trail and human confirmation on irreversible actions are gates, not categories. No volume of completed journeys averages past a missing one.

02 · CONFIDENCE-ADJUSTED
Nine of ten is not ninety of a hundred.

Small samples are penalised for their own uncertainty, and thin evidence withholds a score rather than inflating one. What you get is the floor of what agents can do, never the ceiling of what one did once.

03 · INSTRUMENT-FROZEN
The agents are the instrument. Your property is measured.

The benchmark fleet and task library are versioned and frozen between runs. A score you could move by tuning the agent would measure the agent. The only way this number rises is that the property got better.

COUNTS TOWARD THE SCORE
Completed, value-weighted customer tasks
Recovery from perturbed conditions
Action surfaces an agent can call
Consent, audit trails and confirmation contracts
Observability of the outcome
Cost and latency per completed outcome
DOES NOT COUNT
Per-client prompt tuning
Custom agent memory or fine-tuned probes
Model-specific workarounds
Hidden retries
Post-hoc patches to the run
Being findable, on its own

What the number obligates you to do.

A band is a decision, not a grade. Each one implies a different owner, a different budget line and a different conversation with the risk function.

90 / 100
Excellent
Agent-native. Tasks complete reliably under perturbation with governance evidence to match. Maintain, and monitor for drift as journeys change.
75 / 89
Good
Production-viable for most journeys. Remaining losses are concentrated in specific tasks, fix those, not the platform.
50 / 74
Needs improvement
Agents complete the easy half. Structural work is required on action surfaces and recovery paths before agentic traffic can be relied on.
25 / 49
Poor
Major intervention required. Typically excellent discovery and unusable execution: every agent visit is a cost with no outcome attached.
0–24
Critical
Effectively closed to agents, or gated to zero by a governance failure. Treat as a platform-level programme, not an optimisation.
Most BFSI properties we simulate land between 25 and 45.
Almost always with excellent discovery: the product pages are immaculate, the schema is valid, the crawl policy is clean. The loss is concentrated in identity, conditional forms, third-party hand-offs and outcomes that are never returned to the caller. That is a fixable profile, it is engineering work on a small number of steps, not a platform rewrite.

A score, the evidence under it, and a ranked list of what to fix.

A WARI engagement is delivered as a reading your board can quote and a backlog your engineers can start on the same week.

01
The reading

One w0 score per journey and one for the property, with the band and the gate status behind each.

02
Trajectory evidence

Every run, step, retry and dead end, the recording of what your customer's agent experienced.

03
Ranked remediation

Fixes ordered by score movement per unit of engineering effort, scoped to your stack and controls.

04
Re-simulation

The identical frozen suite, re-run after the work, to prove the delta rather than assert it.

Scoring model, weightings, benchmark task library and perturbation sets are Webzero methodology, disclosed to clients under engagement. No client score, journey or trajectory is published, including on this site.

Get your w0 score.

We start with two live journeys and return a w0 score with the trajectory evidence behind it. Two weeks, no change to your stack.