The Harness Eats The Model

Why context engineering is now the most important AI engineering skill

Summary

In 2026 the AI engineering problem that everyone is talking about is no longer "which model is smarter" — it's "what do you put around the model." Anthropic's July 2026 guidance capped a run of posts from OpenAI, Google, and Meta earlier in the year, all reframing agent reliability as a harness engineering problem: the scaffolding — context delivery, tool interfaces, compaction policy, memory, sandboxing, and verification loops — that wraps a language model determines whether it ships to production or gets stuck in a demo. Anthropic removed over 80% of Claude Code's system prompt for the Claude 5 generation with no measurable loss on coding evals 1. OpenAI, Google (ADK), and Meta (REA) have all shipped or documented production harnesses earlier in 2026 2, 3, 4. This post unpacks what changed, runs a reproducible experiment, surveys how each frontier lab is solving it, and gives an agent-friendly checklist you can paste into your own repo.

Key Takeaways

  • The "model is the product" era is over. Anthropic cut 80%+ of Claude Code's system prompt for Opus 5 with no eval regression 1. The differentiator is now the harness — context, tools, memory, sandboxing, verification 1, 5, 6.
  • Context engineering ≠ prompt engineering. It's the discipline of deciding what the model sees on every single call: what stays in the window, what gets compacted, what lives outside in files or a retrieval index 5.
  • Compaction is conditional, not default. Under modern prompt caching, keeping the full history beat every summarization strategy on cost, latency, and recall in production evals; cap tool outputs first, compact only when you must 5.
  • Progressive disclosure beats upfront prompts. Skills, deferred tool loading, and tree-shaped CLAUDE.md files are now the consensus pattern across Anthropic, OpenAI, and Google ADK 1, 2, 7.
  • The harness is measurable. Vercel moved an agent from 80% → 100% success by removing 80% of its tools; LangChain moved 52.8% → 66.5% on Terminal Bench by changing the harness alone; Princeton's CORE-Bench recorded 42% → 78% on the same model under different scaffolds 8.
  • Four frontier labs are converging on the same architecture: OpenAI's Harness Engineering framing 2, Anthropic's "unhobbling" + Building Effective Agents 1, 9, Google's ADK for long-running agents 3, Meta's REA hibernate-and-wake harness 4.

Problem Background

Why this is the problem of the moment

If you read one AI engineering essay this week, make it Anthropic's "The new rules of context engineering for Claude 5 generation models" by Thariq Shihipar (July 2026) 1. The thesis is blunt: most of what we wrote into agent system prompts for older Claude models was hobbling — guardrails for worst-case behavior — and the newest models don't need it.

"We removed over 80% of Claude Code's system prompt for models like Claude Opus 5 and Claude Fable 5 with no measurable loss on our coding evaluations." — Anthropic, July 2026 1

That sentence landed on top of a wave of harness posts from across the industry earlier in 2026:

  • OpenAI published Harness Engineering as a discipline and Unrolling the Codex Agent Loop 2, 10.
  • Google shipped the Agent Development Kit (ADK) update emphasizing persistent sessions, webhook resumption, and state_delta for long-running agents 3, 11.
  • Meta published the Ranking Engineer Agent (REA) — a production harness that "hibernates and wakes" through multi-day ML pipelines using checkpointed state 4.
  • Birgitta Böckeler, writing on martinfowler.com, synthesized it as the central discipline of AI-assisted software engineering, splitting harness work into feedforward guides (prevent bad output) and feedback sensors (catch it after the fact) 6.

The reason this all converged within a few months is the same problem: agents that look brilliant in a 5-turn demo fall apart in a 50-turn production session. The model isn't getting dumber — its context is.

The mechanics of context rot

Louis Bouchard's Context Engineering in 2026 workshop report, run against a production AI tutor, documents the failure mode precisely 5:

  1. The context window is finite, and everything competes for one attention budget: instructions, retrieved lessons, tool outputs, code, history.
  2. The model is stateless — every call starts from zero. Within a session you need context management; across sessions you need memory.
  3. A growing window hurts three times over: quality degrades as facts get buried (the "lost-in-the-middle" or "context rot" effect), cost climbs because the whole history is re-billed each turn, and time-to-first-token climbs because the model reprocesses the window on every call. Bouchard's measurements put median TTFT at ~22s below 100k input tokens and ~76s above 800k — users feel that latency as a UX regression 5.
  4. Most "long agent session" advice is wrong: in their tutor, retrieved tool outputs (not chat history) dominated the input. Trimming the chat was rounding error; trimming the tool outputs was the lever.

What "harness" actually means

Per LangChain's Anatomy of an Agent Harness and Winder.ai's 2026 comparison 7, 12:

"A harness is every piece of code, configuration, and execution logic that isn't the model itself."

The stack has four layers:

  1. Model — reasons.
  2. Harness — runs one agent (loop, tools, sandbox, memory, permissions, context rules).
  3. Framework — composes several harnesses (LangGraph, CrewAI, ADK, OpenAI Agents SDK).
  4. Platform — runs many harnesses across a team over time (durable execution, cost attribution, governance).

Sebastian Raschka's six-component decomposition of a coding agent is the most-cited checklist: live repo context, cached-prefix prompt split, predefined tools with path-based access control, clipping and compression, structured session memory apart from transcript, and bounded subagents with read-only permissions and a recursion limit 12.

Proposed Experiments / Examples

All four experiments below are designed to be agent-friendly (copy-pasteable, deterministic where possible, and OpenAI / Anthropic / Google APIs are interchangeable).

Experiment 1 — Reproduce Anthropic's "unhobbling" finding

Hypothesis. Removing large blocks of restrictive instructions from your agent's system prompt does not degrade (and may improve) eval scores on the newest Claude / GPT-5.6 / Gemini 3.1 models.

Setup. Use the SWE-bench Verified mini split (50 issues) as the harness.

# harness.py — agent-friendly, model-agnostic
import os, json, subprocess
from openai import OpenAI  # swap import for anthropic.Anthropic or google.genai

client = OpenAI()  # or Anthropic() / genai.Client()

HEAVY_SYSTEM = """You are a meticulous senior engineer.
NEVER delete files.
NEVER add comments unless asked.
DO NOT create planning documents.
DO NOT modify tests.
Default to writing no comments. Never write multi-paragraph docstrings.
… (≈ 2,400 tokens of restrictive instructions) …
"""

LIGHT_SYSTEM = "You are a senior engineer working in this repo. Match the surrounding code's style, naming, and comment density."

# Your own harness pieces: tool schemas, the tool-call loop, patch submission, the issue list.
TOOLS = [READ_FILE, EDIT_FILE, RUN_TESTS, SEARCH_CODE]

def run_episode(system_prompt, issue):
    msgs = [{"role": "system", "content": system_prompt},
            {"role": "user", "content": issue["problem_statement"]}]
    for _ in range(80):  # bounded agent loop
        resp = client.responses.create(
            model="gpt-5.6",          # swap: claude-opus-5, gemini-3.1-pro-preview
            input=msgs, tools=TOOLS,
            reasoning={"effort": "medium"},
        )
        msgs += tool_call_loop(resp, TOOLS)
        if resp.output_text and "PATCH_FINAL" in resp.output_text:
            return submit_patch(msgs, issue)
    return None

results_heavy = [run_episode(HEAVY_SYSTEM, i) for i in ISSUES]
results_light = [run_episode(LIGHT_SYSTEM,  i) for i in ISSUES]

Expected result. Per Anthropic's report 1, LIGHT_SYSTEM matches or beats HEAVY_SYSTEM on resolve rate, with lower cost and lower latency. If it doesn't on your workload, that's a signal your tools or task structure are still doing the hobbling — see Experiment 2.

Experiment 2 — Cap tool outputs before compacting history

Hypothesis. Per Bouchard's production data, capping every tool output at a stable size (truncate head + tail + pointer) cut cost per turn by ~38% with no measurable loss in memory recall, because it shrinks context without rewriting the cached prefix 5.

MAX_TOOL_BYTES = 8_000  # tune for your model

def truncate_tool_output(name: str, raw: str) -> str:
    if len(raw) <= MAX_TOOL_BYTES:
        return raw
    head = raw[: MAX_TOOL_BYTES // 2]
    tail = raw[-MAX_TOOL_BYTES // 2:]
    return f"[{name} truncated: {len(raw):,} bytes total; re-fetch with tool if needed]\n{head}\n…\n{tail}"

Expected result. Cost per turn down 30–45%, no quality regression on long-horizon tasks. This is the "trivial tier" of compaction that pays before you ever call an LLM to summarize.

Experiment 3 — Progressive disclosure of skills

Hypothesis. Replacing a single 6,000-token CLAUDE.md with a 200-token index file plus on-demand skill files improves both eval score and time-to-first-token.

Structure to test.

repo/
  AGENTS.md          # 200 tokens: index only
  skills/
    verify.md        # 400 tokens: how to run tests + lint
    db-migrations.md # 300 tokens
    api-style.md     # 350 tokens

AGENTS.md lists the skill names + one-line descriptions; the model loads full skill bodies only when relevant. This is the pattern Anthropic, OpenAI, and Google ADK all converged on in 2026 1, 2, 11.

Expected result. First-turn latency down 15–30% (smaller cached prefix), eval score flat or up (better-fit guidance at the right moment).

Experiment 4 — Measure your harness, not just your model

Hypothesis. Running the same model under two different harnesses produces a larger accuracy delta than running two different models under the same harness.

Reproduce the pattern documented by MongoDB's synthesis 8: run Terminal-Bench 2.0 with both Claude Code's default harness and a memory-first harness (e.g., Letta Code or a custom harness built around the patterns above). The published gap was Letta Code 59.1% vs Claude Code 41.6% on Claude Opus 4.5 — same model, different harness, ~18-point swing 12.

How Big Companies Solve This

Anthropic — "Unhobbling Claude" + progressive disclosure

Anthropic's July 2026 essay 1 formalizes six myth-busting migrations:

Then (old Claude) Now (Claude 5 generation)
Give Claude rules Let Claude use judgement
Give Claude examples Design interfaces
Put it all upfront Use progressive disclosure
Repeat yourself Simple tool descriptions
Memory in CLAUDE.md Auto-memory
Simple specs (markdown) Rich references (HTML artifacts, test suites, code)

The system prompt collapse (80%+ removed) is the headline. The deeper point is that each old rule existed because the model couldn't do something the harness had to do for it — and each new rule expires as the model improves.

Anthropic's earlier Building Effective Agents essay 9 remains the canonical reference for workflows vs. agents — when to use a deterministic chain, when to let the model steer.

OpenAI — Harness Engineering as a discipline

OpenAI's Harness Engineering essay (February 2026) and Unrolling the Codex Agent Loop 2, 10 frame the harness as a first-class engineering artifact: design the scaffolding, instrument it, and iterate. Codex's agent loop exposes the same components as Anthropic's — prompt split, tool dispatch, compaction, subagents — with explicit hooks for harness work. OpenAI's Practical Guide to Building AI Agents (April 2025) layers in single-agent vs. multi-agent orchestration (manager vs. decentralized handoffs), tool design for many-to-many relationships, and layered guardrails (input validation, output filtering, tool-risk ratings, human-intervention triggers) 13.

Google DeepMind / Google Cloud — ADK + long-running agents

Google's Agent Development Kit (ADK) 3 is the production-grade framework; the May 2026 Build Long-running AI agents that pause, resume, and never lose context with ADK guide adds the missing piece for production: DatabaseSessionService for persistent sessions, webhook-triggered state_delta resumption (so containers can scale to zero), and explicit state machines instead of dumping raw JSON into vector databases. The April 2026 Next '26 developer keynote extended this to Skills — small, named, on-demand instruction modules, identical in shape to Anthropic's and OpenAI's 11.

Meta — Ranking Engineer Agent (REA)

Meta's engineering blog on REA 4 documents a production harness for multi-day ML pipeline automation. The interesting trick is hibernate-and-wake checkpointing: when REA launches a training job that runs for hours or days, it delegates the wait to a background system, shuts down, and resumes where it left off when the job completes. This is the pattern every long-running agent will eventually need.

Microsoft — Azure SRE Agent as a benchmark

Microsoft's Azure SRE Agent has handled 35,000+ production incidents autonomously, reducing Azure App Service time-to-mitigation from 40.5 hours to 3 minutes 14. The harness layers MCP and Python tools, telemetry, code repositories, and incident management on top of existing workflows, and Microsoft credits its balance of autonomy and governance for safe operation at scale 14.

Discussion

What this means for AI engineering teams

  1. Stop paying prompt engineers to write 6,000-token system prompts. The 2026 frontier models want concise, judgement-respecting prompts with rich on-demand context. Your senior engineers' time is better spent designing tools, skills, and verification loops.
  2. Treat the harness as a product. It has a roadmap, SLAs, observability, and regression tests. Böckeler's framing on martinfowler.com 6 — feedforward guides + feedback sensors, computational vs. inferential, regulation categories (maintainability / architecture fitness / behavior) — is the right mental model.
  3. Measure with harnesses-in-the-loop benchmarks. Standard model benchmarks under-rate or over-rate agents based on scaffold luck. On CORE-Bench the same model scored 42% under one scaffold and 78% under another 8.
  4. Compaction is a named response to a named constraint, not a default. Window pressure, cache cost above ~$0.55/M input tokens, or measured quality rot — pick the constraint first, then pick the lever 5.
  5. Skills over megafiles. Across Anthropic, OpenAI, and Google the pattern is identical: small, named, discoverable, load-on-demand. Your AGENTS.md should be an index, not a manual.

Open questions

  • Will the harness-model split itself get absorbed back into the model? Anthropic's essay implies yes — every harness component exists "because the model can't do something," and those assumptions expire 1. The interesting 2027 question is which components expire first.
  • How will multi-agent harness orchestration (manager handoffs vs. decentralized swarms) settle out? Anthropic's multi-agent research 15 and OpenAI's decentralized patterns 13 are still empirically divergent.
  • What's the right evaluation harness for an evaluation harness? Published scores for the same model already swing with the harness, from 41.6% to 59.1% on Terminal-Bench 2.0 12.

Conclusion

The most important AI engineering problem of 2026 isn't a smarter model — it's everything around the model. Anthropic proved you can delete 80% of the system prompt 1. Bouchard proved that with prompt caching, compaction should be conditional 5. OpenAI, Google, Meta, and Microsoft all shipped production harnesses earlier in 2026 that treat the scaffolding as a first-class engineering artifact 2, 3, 4, 14. For the first time in the LLM era, the harness moves more numbers than the model.

The implication for engineering teams is concrete: spend your next sprint on tool design, progressive-disclosure skills, cap-before-compaction, and verification loops — not on writing longer prompts. The frontier labs already did.

Sources

  1. Shihipar, T. The new rules of context engineering for Claude 5 generation models. Anthropic, July 24, 2026. https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models
  2. OpenAI. Harness engineering: leveraging Codex in an agent-first world. OpenAI Index, February 2026. https://openai.com/index/harness-engineering/ (referenced via the awesome-harness-engineering curated index — direct fetch returned 403, citation retained from index)
  3. Google. Build Long-running AI agents that pause, resume, and never lose context with ADK. Google Developers Blog, May 2026. https://developers.googleblog.com/build-long-running-ai-agents-that-pause-resume-and-never-lose-context-with-adk/
  4. Meta Engineering. Ranking Engineer Agent (REA): Meta's Autonomous AI System for Ads Ranking. March 2026. https://engineering.fb.com/2026/03/17/developer-tools/ranking-engineer-agent-rea-autonomous-ai-system-accelerating-meta-ads-ranking-innovation/
  5. Bouchard, L. Context Engineering in 2026: Why We Stopped Compacting Our Agent's Context. August 2026. https://www.louisbouchard.ai/context-engineering-2026/
  6. Böckeler, B. Harness engineering for coding agent users. martinfowler.com, April 2026. https://martinfowler.com/articles/exploring-gen-ai/harness-engineering.html
  7. LangChain. The Anatomy of an Agent Harness. LangChain Blog, March 2026. https://www.langchain.com/blog/the-anatomy-of-an-agent-harness
  8. MongoDB. Agent Harness: Why the LLM Is the Smallest Part of Your Agent System. MongoDB Engineering Blog, April 2026. https://www.mongodb.com/company/blog/technical/agent-harness-why-llm-is-smallest-part-of-your-agent-system
  9. Anthropic. Building Effective Agents. Anthropic Research, December 2024. https://www.anthropic.com/research/building-effective-agents
  10. OpenAI. Unrolling the Codex Agent Loop. OpenAI Index, January 2026. https://openai.com/index/unrolling-the-codex-agent-loop/
  11. Google. Next '26 Keynote: Building ADK Agents with Skills and Tools. Google Codelabs, April 2026. https://codelabs.developers.google.com/next26/dev-keynote/building-agents-with-skills
  12. Winder, P. A Comparison of AI Agent Harnesses in 2026. Winder.ai, August 2026. https://winder.ai/ai-agent-harness-comparison/
  13. OpenAI. A Practical Guide to Building Agents. OpenAI, April 2025. https://openai.com/business/guides-and-resources/a-practical-guide-to-building-ai-agents/
  14. Microsoft. How we build and use Azure SRE Agent with agentic workflows. Microsoft Tech Community, April 2026. https://techcommunity.microsoft.com/blog/appsonazureblog/how-we-build-and-use-azure-sre-agent-with-agentic-workflows/4508753
  15. Anthropic. Patterns and problems in emerging multiagent systems. Anthropic Research, August 2026. https://www.anthropic.com/research/multiagent-systems

Curated index used for cross-validation: ai-boost/awesome-harness-engineering — https://github.com/ai-boost/awesome-harness-engineering

Keep reading

All posts →

Be first in when doors open.

Get early access