I spent two days at apidays London listening to people building the infrastructure for an agentic economy.
The programme described APIs as the “legal and functional hands” of AI. That feels exactly right. Models can reason and agents can plan, but APIs and MCP tools are how those decisions become consequential actions: moving money, changing records, accessing data or triggering another automated workflow.
Across the sessions, one question kept recurring:
How do we know these systems are actually doing what we intended?
Not whether the model is available. Not whether an API returned 200 OK. Not even whether the request was authenticated and technically valid.
Did the complete production workflow remain reachable? Did it use the right tool? Did it return a meaningful result? Did it stay within its controls? And can we prove that from outside the systems responsible for producing the answer?
That is the assurance gap.
A valid API call can still produce an invalid outcome
Afsha H. of Capgemini captured the issue neatly in her talk, A Valid API Call, An Invalid Outcome.
Traditional infrastructure is good at answering two questions:
- Was the request technically valid?
- Was the caller authorized?
Agentic systems add a third:
- Was the resulting action appropriate?
An authenticated request can reach a valid endpoint, satisfy its schema and receive a successful response while still producing the wrong human or business outcome.
That problem is not confined to AI. We have always had APIs that were technically healthy but commercially broken. A payment may be accepted but reconciled incorrectly. A customer journey may complete while returning unusable data. A service may be operational from inside the provider’s network but unreachable to customers in a particular geography.
AI makes this existing gap much harder to ignore.
Internal observability has an external blind spot
Logs, traces and infrastructure telemetry are essential. They tell teams a great deal about what happened inside the systems they control.
But customers, partners and agents do not experience your internal architecture. They experience a connected production journey across systems and organizational boundaries.
From the outside, the questions are different:
- Does DNS resolve from the network where the customer operates?
- Can the caller establish a connection through the CDN, gateway and security controls?
- Does authentication work with a real identity and realistic credentials?
- Can an agent discover the tools it is supposed to use?
- Does the advertised schema still match production behaviour?
- Does the complete workflow produce the expected outcome?
- Is that outcome consistent across locations, providers and repeated runs?
A green internal dashboard cannot prove these things.
This is why reachability matters. Before evaluating whether a system behaved correctly, you must establish that the intended consumer can reach and use it at all.
Outside-in synthetic execution provides that independent signal. It approaches the service as a customer, partner or agent would: through the public network, across the production security boundary and through the authenticated workflow. That is the core of API monitoring from outside the systems responsible for producing the answer.
That evidence is complementary to observability. It answers a question internal telemetry cannot answer on its own:
What is the system actually like to use from out here?
Agents multiply the consequence of small failures
Yu Zhen Koh of Wise described agents as capable of moving confidently and quickly while making wrong turns.
That is the operational change that matters.
A human encountering an ambiguous response may stop, investigate or ask for help. An agent can consume that response, make another decision, select another tool and continue the workflow. Agents can also retry, operate in parallel and initiate further agents.
A small ambiguity therefore stops being a single bad interaction. It can become the input to a chain of additional actions.
The same effect applies to scale. An integration defect that once affected the pace of human activity can now be exercised continuously at machine speed. An agent may discover and consume capabilities dynamically rather than following a journey that was anticipated and manually tested in advance.
Failures consequently compound in three directions:
- Speed: more actions can be attempted in less time.
- Reach: one result can influence multiple downstream systems.
- Variation: the same goal may produce different paths and outcomes.
Static certification and pre-production testing remain valuable, but they cannot establish that a changing production ecosystem continues to behave correctly.
A successful response is no longer enough
One apidays discussion on APIs and AI made the transition particularly explicit.
Historically, teams monitored availability, latency, throughput and error codes. Agentic workflows also require visibility into tool selection, decisions, business outcomes and quality. The speakers’ conclusion was that evaluation must become an engineering discipline: every failure should become a future test case.
That creates a practical assurance model.
First, execute a realistic authenticated scenario from outside the system.
Then evaluate the resulting evidence against a defined expectation:
- A deterministic assertion
- A security or identity control
- A schema or data boundary
- A business rule
- An industry profile
- A probabilistic evaluator or judge
Finally, retain the verdict over time so that teams can see drift, receive alerts and demonstrate what happened.
The synthetic execution generates independent evidence. Controls define what is permitted or expected. Evaluators determine whether the evidence is acceptable. Assurance shows whether the production system remains within those expectations.
MCP makes the need more immediate
Goutam Verma of Expedia Group explained why MCP cannot simply be treated as another stateless API call. MCP introduces sessions, changing tool inventories, bidirectional interactions and protocol-level context.
Other sessions explored the unresolved questions around agent identity and authority: who owns the agent, whose credentials it uses, which tools it may invoke, whether authority was delegated and whether that authority can be revoked during a workflow.
Gateways and control planes will play a crucial role in enforcing these boundaries. Chris Wood of Ozone API also showed how reusable, layered security profiles could make policy more deterministic across implementations.
But enforcement and verification are different jobs.
A gateway can enforce the policy it has been configured to enforce. It cannot independently prove that the complete external workflow remains reachable, that every dependency behaved correctly or that the final result was acceptable.
The more infrastructure we add to control agentic execution, the more important it becomes to test the resulting system independently.
Outside-in assurance is not AI red teaming
There is a temptation to place every form of AI testing under a broad security or red-teaming category. I think that misses the larger operational requirement.
The everyday problem is more fundamental:
Can this production system reliably do its job, and can the organization prove it?
That includes security, but also reachability, performance, resilience, semantic correctness, tool behaviour and business outcomes.
It is closer to continuous assurance than periodic testing.
For high-stakes actions, assurance does not replace inline controls, human approval or runtime enforcement. It tests that those mechanisms—and the systems surrounding them—continue to work as intended.
The emerging production model
My main conclusion from apidays is that APIs are not becoming less important in an agentic world. They are becoming more consequential.
As agents become API consumers, production paths become more dynamic. As systems become less deterministic, successful execution becomes harder to infer from status codes and internal telemetry. As actions happen at machine speed, the cost of discovering problems through customers rises sharply.
The answer is not another isolated AI dashboard.
It is an independent assurance layer spanning APIs, MCP tools and the downstream services they orchestrate:
- Continuous outside-in execution
- Real production identities and workflows
- Reachability from the places that matter
- Reusable controls and policies
- Deterministic and probabilistic evaluation
- Historical evidence, drift and alerts
APIContext already continuously tests digital services from around the world so enterprises can see when their APIs, websites and workflows fail or slow down. See how this applies to agentic AI.
The next step is to verify something broader:
Whether APIs, MCP tools and agentic services remain reachable—and continue to behave correctly—when the outside world depends on them.
That need existed before AI. Agentic systems are causing it to compound at machine speed.

