
Summary
In 2026 the dominant production failure for AI agents is no longer hallucination or reasoning — it is tool-call reliability. Anthropic, OpenAI, Latitude, Waxell, and Stackpulsar all converge on the same diagnosis: agents call tools with wrong arguments, missing fields, malformed JSON, or semantically adjacent (but incorrect) values, and the resulting corruption silently propagates through multi-step pipelines. This post breaks down the three classes of tool-call failure, the observability stack required to catch them, the experiments that prove the gap between benchmark scores and production reality, and how Anthropic, OpenAI, and the open-source community are responding.
Key Takeaways
- Tool misuse is the #1 agent-specific failure mode in production, ahead of hallucination, reasoning failures, and context rot — per Latitude's 2026 observability framework report 1.
- Three failure classes require different handling: schema mismatch (retry-loop trap), partial data (retryable with different params), and semantic garbage (invisible to schema validation) 1.
- Final-output evaluation is structurally blind to tool-call fidelity. A single malformed argument at step 2 silently corrupts every downstream step without producing a hard failure 1.
- Anthropic's "Writing effective tools for agents" (Sept 2025, still the canonical 2026 reference) argues tools are a new kind of software contract — between deterministic systems and non-deterministic agents — and must be designed differently from APIs for human developers 2.
- OpenAI's
strict: truemode for function calling is the closest thing to a vendor-level schema-enforcement guarantee; it locks tool calls to the JSON schema instead of treating compliance as best-effort 3. - The SWE-bench Verified leaderboard is 99% self-reported as of June 16, 2026 — only 1 of 100 entries is independently verified, and the same model can score 50.2%–55.4% on SWE-bench Pro depending purely on the scaffold 4.
- Production observability requires OTel-style trace hierarchies (
agent.task→agent.step→llm.call+tool.call) 5, built on OpenTelemetry's GenAI semantic conventions, which are still marked Development 12.
Problem Background
Why tool calls are the new bottleneck
In 2024–2025 the AI engineering community obsessed over prompt engineering. In 2026 the conversation has shifted decisively to context engineering and, more specifically, to the tool interface — the contract between a non-deterministic agent and the deterministic APIs, databases, and file systems it invokes.
Anthropic's framing in Effective context engineering for AI agents is explicit: "Context, therefore, must be treated as a finite resource with diminishing marginal returns… Every new token introduced depletes this [attention] budget by some amount." 6 Tools are the primary mechanism by which an agent pulls new tokens into its context, which makes the tool interface the highest-leverage point of failure in any multi-step agent.
The structural reason is straightforward. A well-designed agent isn't just a model generating text — it's a chain of tool invocations connected by model reasoning. Each link in that chain has specific schema expectations: exact field names, exact data types, exact nesting structures. The model must satisfy every one of these on every call. At context window lengths typical of production multi-step workflows, models fail at schema compliance more often than benchmark scores suggest — because benchmarks typically measure final-output accuracy, not intermediate tool-call fidelity 1.
What "tool argument rot" actually looks like
The clearest documented case is Gabriel Anhaia's May 2026 production trace, reproduced in Waxell's analysis 1:
A customer service agent called a database lookup tool. The tool returned a truncated JSON blob — an upstream gateway had a 4KB response cap nobody had documented. The model correctly identified the response as broken and decided to retry. The tool returned the same truncated payload. The model retried again. And again. Seventeen times. Each turn a full prompt round-trip with growing context, accumulating tokens, running up costs — all because nothing between the tool and the model translated "this response is malformed" into actionable signal.
The model wasn't wrong. It was doing exactly what it was trained to do: retry on apparent transient failure. The failure wasn't the model. It was the gap between what the tool returned and what the agent could reason about. That gap is where most AI agent output quality problems originate — and it's the last place most teams look.
The three failure classes
Not all tool-call failures look the same, and the agent's right move depends heavily on which class it's dealing with 1:
| Class | Symptom | Right response | Wrong response |
|---|---|---|---|
| Schema mismatch | Tool returned malformed JSON, truncated payload, wrong field types | Stop retrying the same call; surface the structural error to the harness | Retry with identical args (the Anhaia 17× loop) |
| Partial data | Valid JSON, but pagination cut off / rate-limited / timeout | Retry with smaller page, narrower filter, different time window | Treat as complete; propagate truncated result downstream |
| Semantic garbage | Valid JSON, complete data, but wrong in ways schema can't detect | Trajectory-level evaluation; re-query with corrected semantics | Pass through silently; downstream step inherits the wrong value |
The third class — semantic garbage — is the hardest to catch. An agent passes a customer name in a field that expected an ID. A search API returns zero results because a filter value was semantically adjacent to the right one but not syntactically correct. The response looks fine. The content makes no sense relative to what was asked.
Each class produces a different downstream failure pattern. Without tool-level instrumentation that captures the argument, the schema context, and the response in sequence, distinguishing these three classes after the fact requires manual reconstruction of execution traces — which doesn't scale.
Proposed Experiments / Examples
Experiment 1 — Reproduce the Anhaia retry loop
Hypothesis. A naive agent loop will retry a structurally broken tool response indefinitely because the model has no way to distinguish "transient network error" from "structurally broken response."
Setup. Build a tool stub that returns a truncated JSON payload (4KB cap, same as the production gateway). Wrap it in a vanilla ReAct loop with no schema validation. Run 50 trials of a customer-lookup task.
Expected result. Median retries before context budget exhaustion: 12–17. Total tokens wasted per failed task: ~40k. Pass rate: 0%.
Fix. Add a JSON-schema validator between tool and model. On ValidationError, emit a structured error message to the model: "The previous response was truncated by an upstream 4KB cap. Do not retry. Ask the user for a narrower query or paginate client-side." Re-run. Expected retries: 0–2. Pass rate: ~85%.
Why this matters. This is the single highest-ROI instrumentation change for production agents. The cost of adding a validator is ~50 lines of code. The cost of not adding it is unbounded token spend on retry storms.
Experiment 2 — Schema compliance vs. context length
Hypothesis. Tool-call schema compliance degrades as context window length grows, even with identical instructions.
Setup. Take a fixed tool schema (10 fields, 3 nested objects). Generate 100 evaluation prompts. Run the same model with context windows of 2k, 8k, 32k, and 128k tokens. Measure: (a) schema validity rate, (b) argument accuracy rate, (c) final task success rate.
Expected result (illustrative numbers, not published measurements; Latitude and Waxell describe the degradation, not these figures 1, 7).
- Schema validity at 2k: ~98%
- Schema validity at 128k: ~71%
- Argument accuracy at 2k: ~94%
- Argument accuracy at 128k: ~52%
- Final task success at 2k: ~89%
- Final task success at 128k: ~38%
The cliff between argument accuracy and final task success is the production failure surface: agents look like they're working, but the intermediate values are subtly wrong.
Experiment 3 — Trajectory evaluation vs. final-output evaluation
Hypothesis. Final-output LLM-as-judge evaluation passes agents that trajectory evaluation correctly flags.
Setup. Build 50 tasks where the correct final answer can be reached via two paths: (a) a clean trajectory that queries the right tool with the right arguments, (b) a degenerate trajectory that infers an intermediate value rather than querying for it. Run both paths. Evaluate with (i) final-output LLM-as-judge, (ii) trajectory-level comparison against expected tool-call sequence.
Expected result. Final-output judge: both paths pass. Trajectory judge: path (b) fails on the "expected tool call not made" check. In production, path (b) is the silent failure mode — correct output, incorrect process — that surfaces only when the same shortcut produces a wrong answer on different data.
Example code — minimal tool-call validator
import json
from jsonschema import validate, ValidationError
TOOL_SCHEMA = {
"type": "object",
"properties": {
"customer_id": {"type": "string", "pattern": "^C[0-9]{6}$"},
"include_orders": {"type": "boolean"},
"limit": {"type": "integer", "minimum": 1, "maximum": 100}
},
"required": ["customer_id"],
"additionalProperties": False
}
TOOL_RESPONSE_SCHEMA = {
"type": "object",
"properties": {
"customer_id": {"type": "string"},
"name": {"type": "string"},
"orders": {"type": "array", "items": {"type": "object"}}
},
"required": ["customer_id", "name"]
}
def safe_tool_call(tool_fn, args, model, max_retries=2):
"""Wrap a tool call with schema validation and structured error feedback."""
for attempt in range(max_retries + 1):
try:
validate(instance=args, schema=TOOL_SCHEMA)
except ValidationError as e:
if attempt == max_retries:
return {
"error": "schema_invalid",
"detail": str(e),
"advice": "Do not retry. Ask the user to clarify the input."
}
args = model.repair_args(args, error=str(e))
continue
result = tool_fn(**args)
try:
parsed = json.loads(result) if isinstance(result, str) else result
validate(instance=parsed, schema=TOOL_RESPONSE_SCHEMA)
return parsed
except (ValidationError, json.JSONDecodeError) as e:
if attempt == max_retries:
return {
"error": "response_malformed",
"detail": str(e),
"advice": "Tool is broken upstream. Do not retry with same args."
}
# Different args, not same args — this is the key fix
args = model.repair_args(args, error=f"upstream: {e}")
The critical line is the last one: when the response is malformed, we do not retry with the same arguments. We either give up or ask the model to repair the arguments based on what it learned about the upstream failure. This is the structural fix that breaks the Anhaia loop.
How Big Companies Solve This
Anthropic — Tools as a new software contract
Anthropic's Writing effective tools for agents (Sept 2025) is the canonical 2026 reference 2. The core thesis: tools are a new kind of software that reflects a contract between deterministic systems and non-deterministic agents. "When a user asks 'Should I bring an umbrella today?,' an agent might call the weather tool, answer from general knowledge, or even ask a clarifying question about location first. Occasionally, an agent might hallucinate or even fail to grasp how to use a tool."
Anthropic's published principles for tool design 2:
- More tools don't always lead to better outcomes. "A common error we've observed is tools that merely wrap existing software functionality or API endpoints — whether or not the tools are appropriate for agents."
- Search over listing. Prefer tools that let agents query for what they need rather than returning giant enumerations.
- Consolidation. Reduce tool count by combining related operations.
- Namespacing. Group related tools under clear prefixes.
- Response design. Return token-efficient, structured responses.
- Error handling. "If a tool call raises an error (for example, during input validation), you can prompt-engineer your error responses to clearly communicate specific and actionable improvements, rather than opaque error codes or tracebacks."
- Description engineering. Treat tool descriptions with the same care as system prompts.
Anthropic's own evaluation methodology is rigorous: "We relied on held-out test sets to ensure we did not overfit to our 'training' evaluations. These test sets revealed that we could extract additional performance improvements even beyond what we achieved with 'expert' tool implementations — whether those tools were manually written by our researchers or generated by Claude itself." 2
OpenAI — Strict mode and structured outputs
OpenAI's function-calling docs explicitly recommend strict: true 3: "Setting strict to true will ensure function calls reliably adhere to the function schema, instead of being best effort." This is the closest thing to a vendor-level schema-enforcement guarantee and is the single highest-leverage flag any OpenAI agent developer can flip.
OpenAI's New tools for building agents announcement frames the broader strategy: "Today, we're releasing the first set of building blocks that will help developers and enterprises build useful and reliable agents. We view agents as systems that independently accomplish tasks on behalf of users." 8 The release includes the Responses API, the Agents SDK, and built-in tracing — all aimed at closing the observability gap that makes tool-call failures invisible.
UC Berkeley — the function-calling leaderboard
The Berkeley Function Calling Leaderboard (BFCL v3/v4) has become a standard stress test. Berkeley's BFCL v3 multi-turn pass rate for frontier models sits at ~55–65% 9 — meaning even the best models fail roughly 4 out of 10 multi-turn tool-use tasks. This is the structural baseline that production agents must engineer around.
Open-source ecosystem — Llama-based models
The open-source ecosystem has made dramatic progress on tool-calling capability in 2025–2026, with Llama-based tool-calling models closing the gap to frontier proprietary systems on standard benchmarks 10. The practical implication: tool-call reliability is no longer a model-quality problem alone — it's a harness, schema, and observability problem.
The observability vendors — Latitude, Waxell, Stackpulsar, Maxim
A new category of AI agent observability platforms has emerged specifically to address the tool-call reliability gap:
- Latitude (open-source, MIT-licensed) models agent execution as a causal trace, automatically clusters related failures, and generates evals from production data. Their 2026 report is the primary source for the "tool misuse is the #1 failure mode" claim 1, 7.
- Waxell Observe captures every model call and tool use in order, with 50+ policy categories that operate at the individual step level. Policy checks run at 0.045ms p95 latency 1.
- Stackpulsar documents four production failure modes (silent loops, context overflow, tool call cascades, credential drift) and the OTel span hierarchy required to surface them 5.
- Maxim AI focuses on trajectory-level evaluation and the gap between "answers a question" and "answers correctly" 11.
The benchmark problem — SWE-bench Verified is 99% self-reported
A critical caveat for anyone choosing models based on tool-use benchmarks: as of June 16, 2026, the llm-stats SWE-bench Verified leaderboard listed 100 models, but only 1 result was independently verified 4. The same Claude Opus 4.5 model scored 50.2%–55.4% on SWE-bench Pro across three different scaffolds — a 5.2-point spread from scaffold differences alone. Scale AI's own analysis puts the swing from harness choices at 10–20 points.
The practical lesson: a SWE-bench number is a measurement of a system (model + scaffold), not of a model alone. When you wire a model into your own tooling, you are building a harness — and your harness will produce a different number than any leaderboard.
Discussion
Why this problem is harder than hallucination
Hallucination is a model-quality problem. You can fine-tune against it, you can RAG around it, you can detect it with reasonable accuracy post-hoc. Tool-call reliability is a systems problem — it sits at the intersection of model behavior, schema design, error handling, observability, and harness configuration. There is no single fix.
The Anhaia trace is the canonical example: the model did the right thing (retry on apparent failure). The tool did the wrong thing (returned truncated data). The harness did nothing (no schema validation). The result was a 17× retry storm that produced no value. Fixing any one of the three layers would have prevented it. Fixing only the model would not.
The "correct output, incorrect process" failure mode
The most insidious tool-call failure is the one that doesn't fail. The agent references last year's report instead of this year's. It queries a deprecated API endpoint that happens to return data that was valid six months ago. It infers an intermediate value — one that happens to be right in this instance — rather than querying for it. The output passes final-output validation. It gets used. The failure is invisible until the same behavior on different data produces a wrong answer 1.
This class of failure requires trajectory evaluation to detect. You need to see not just what the agent produced, but how it navigated there: which tool calls it made, in what order, with what arguments, and what each returned. Without that trace, process failures are invisible by definition.
What "agent-friendly" means for this content
This post is structured for both human engineers and downstream AI agents:
- Explicit key takeaways at the top so an agent can extract the thesis without parsing prose.
- Numbered sources with stable identifiers for citation graphs.
- Code blocks that are syntactically valid and copy-pasteable.
- Tables for failure-class taxonomy and benchmark comparisons.
- Hypothesis / Setup / Expected result structure for experiments so an agent can reproduce them.
- Inline citations in
[Source N]format that map to the Sources section at the bottom.
The goal is that an agent reading this post can: (1) extract the core claim, (2) cite the primary source, (3) reproduce the experiment, (4) implement the fix.
Conclusion
Tool-call reliability is the defining AI engineering problem of 2026. It is not a model problem, it is not a prompt problem, and it is not solvable by a single vendor feature. It is a systems problem that requires:
- Schema validation at the tool interface (not just at the model).
- Structured error feedback that distinguishes transient from structural failures.
- Trajectory-level evaluation that catches "correct output, incorrect process" failures.
- OTel-style trace hierarchies (
agent.task→agent.step→tool.call) for post-hoc debugging. - Honest benchmarking that reports scaffold alongside model, and uses held-out test sets.
Anthropic's tool-design principles and OpenAI's strict: true mode are the two most actionable vendor-level primitives. Latitude, Waxell, Stackpulsar, and Maxim are the four most mature observability platforms. The SWE-bench Verified leaderboard should be treated as a pass/fail tier filter, not a rank order.
The teams that win in 2026 are the ones that treat the tool interface as a first-class engineering surface — with the same rigor they would apply to any other production API contract.
Sources
- Waxell — AI Agent Tool Call Failures: #1 Production Problem [2026] — https://waxell.ai/blog/ai-agent-tool-call-failures-output-rot
- Anthropic — Writing effective tools for agents — with agents — https://www.anthropic.com/engineering/writing-tools-for-agents
- OpenAI — Function calling | OpenAI API — https://developers.openai.com/api/docs/guides/function-calling
- Digital Applied — SWE-bench in 2026: Benchmarks vs Scaffolding Reality — https://www.digitalapplied.com/blog/swe-bench-verified-june-2026-benchmark-vs-scaffolding-analysis
- Stackpulsar — AI Agent Reliability 2026: Failure Modes + Observability — https://stackpulsar.com/blog/ai-agent-reliability-monitoring/
- Anthropic — Effective context engineering for AI agents — https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Latitude — Detecting AI agent failure modes in production: A framework for observability-driven diagnosis — https://latitude.so/blog/ai-agent-failure-detection-guide
- OpenAI — New tools for building agents — https://openai.com/index/new-tools-for-building-agents/
- Future AGI — OpenAI Agent SDK: FutureAGI Reliability Guide (2026) — https://futureagi.com/glossary/openai-agent-sdk/
- Zylos Research — Tool Use and Function Calling in AI Agents — Standards, Benchmarks, and Emerging Patterns — https://zylos.ai/research/2026-04-07-tool-use-function-calling-standards-benchmarks/
- Maxim AI — The State of AI Hallucinations in 2025: Challenges, Solutions, and the Maxim AI Advantage — https://www.getmaxim.ai/articles/the-state-of-ai-hallucinations-in-2025-challenges-solutions-and-the-maxim-ai-advantage/
- OpenTelemetry — Semantic conventions for generative client AI spans — https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md


