
Summary
In 2026, the hardest problem in AI engineering is no longer "how smart is the model" — it is "how do you keep an agent coherent, safe, and productive across hours or days of autonomous work." Frontier labs have converged on the same diagnosis: context windows leak attention, raw tool calls bleed tokens, and unconstrained execution is unsafe. Their answers differ in detail but agree in shape: Anthropic ships context engineering (compaction, structured note-taking, multi-agent harnesses) and OS-level sandboxing for Claude Code 1, 2, 3, 4; OpenAI ships a model-native harness with native sandbox execution in its Agents SDK, backed by seven sandbox providers and standardized primitives like MCP, AGENTS.md, skills, and apply_patch 5, 6. This post walks through the problem, the experiments, and how each lab is solving it.
Key Takeaways
- Context is a finite resource. Chroma's "Context Rot" research tested 18 frontier models in 2025 and found every one degrades as context length grows — including models with 1M-token windows 7.
- Compaction is the primary lever. Anthropic treats compaction, structured note-taking, and multi-agent architectures as the three techniques that make long-horizon agents viable, and shipped a server-side compaction API (beta header
compact-2026-01-12) on January 12, 2026 1, 8. - A fresh context window every session is the killer constraint. Anthropic's own long-running-agent harness uses an initializer agent on first run and a coding agent on every subsequent session, each working one feature at a time against a shared
feature_list.jsonandclaude-progress.txt2. - Sandboxing is now a first-class primitive, not a tax. Anthropic open-sourced its OS-level sandbox runtime (Linux bubblewrap, macOS seatbelt), reporting an 84% reduction in permission prompts internally 3, 9. OpenAI shipped sandbox execution built into the Agents SDK with seven provider integrations 5, 6.
- Code execution with MCP beats direct tool calls. Anthropic measured a 98.7% token reduction when agents load MCP tools on demand from a filesystem-shaped code API instead of stuffing all tool definitions into context 4.
- The agent stack is consolidating. OpenAI shipping MCP,
AGENTS.md, skills, andapply_patchfirst-class in their SDK means the cross-vendor agent vocabulary is settling 5, 6.
Main Content
Problem Background: Why Long-Horizon Agents Are Hard
Three forces converge to make autonomous, long-running agents the dominant engineering headache of 2026.
1. Context is not free, even in a 1M-token window. Chroma's "Context Rot" research paper formalized a phenomenon practitioners had been reporting: as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases 7. Anthropic's framing is even sharper — context must be treated as a finite resource with diminishing marginal returns, because the transformer creates n² pairwise relationships for n tokens, and trained attention patterns are biased toward shorter sequences 1. A 1M-token window does not mean an agent reasons reliably across 1M tokens; it means the cliff arrives later, not that it disappears 1.
2. Each new session starts with zero memory. Anthropic describes the long-running-agent problem vividly: imagine a software project staffed by engineers working in shifts, where each new engineer arrives with no memory of what happened on the previous shift 2. Without an explicit mechanism to bridge sessions, agents either one-shot the entire app (exhausting context mid-implementation) or prematurely declare the job done after seeing partial progress 2.
3. Unconstrained execution is unsafe. Agents that edit files, run shell commands, and call external services can be hijacked by prompt injection, exfiltrate credentials, or destroy data. Approval prompts were the default defense, but they create "approval fatigue" — the more dialogs a user clicks through, the less they read them 3, 10.
These three forces are why "context engineering" has displaced "prompt engineering" as the central concern of applied AI in 2026 1.
Proposed Experiments / Examples
Below are three reproducible scaffolds any engineering team can run to measure their own posture on the long-horizon-agent problem, modeled on what Anthropic and OpenAI have described.
Experiment 1 — Context-Rot Curve for Your Stack
Reproduce the Lost in the Middle methodology (Liu et al.) on your own model:
- Build a 20-document context where exactly one document contains a "needle" fact (e.g., a unique ID and value).
- Vary the needle's position across runs: positions 0–4, 5–9, 10–14, 15–19.
- Vary the total context size from ~2K tokens to your model's max.
- Measure exact-match retrieval accuracy.
- Plot accuracy vs. needle position at each total size.
Expected outcome (per Liu et al.): accuracy drops by more than 20% when the needle lands in the middle of a 20-document context, sometimes below the model's no-documents baseline 16; Chroma found that every model it tested degrades as total input length grows 7. Use this curve to set your compaction trigger (commonly around 70% of the context budget) 11.
Experiment 2 — Compaction vs. Sliding Window vs. Raw Append
For a long-horizon coding task (e.g., incrementally building a 200-feature CRUD app):
# Pseudo-code — three harness variants
def run_session(agent, history, strategy):
if strategy == "raw":
return agent.continue_with(history) # OOM / context rot
elif strategy == "sliding_window":
return agent.continue_with(history[-N_TURNS:])
elif strategy == "compaction":
if token_count(history) > THRESHOLD:
history = compact(history) + history[-K_TURNS:]
return agent.continue_with(history)
Track (a) features completed end-to-end, (b) reverted commits, (c) wall-clock time, (d) tokens spent per session. Anthropic's internal harness shows that compaction alone is insufficient — you also need the initializer-agent / feature-list / clean-state-commit pattern to avoid one-shotting and premature-completion failures 2.
Experiment 3 — Sandbox Boundary Measurement
Wrap your agent's bash tool in an OS-level sandbox (Linux bubblewrap, macOS seatbelt) and measure:
- Permission-prompt frequency before/after (Anthropic reports an 84% internal reduction 3, 9).
- Worst-case blast radius: can a prompt-injected agent reach
~/.ssh, write to/etc, or open an outbound TCP socket to a non-allow-listed host? - Recovery time when a sandbox is destroyed mid-session (OpenAI's SDK restores state from a snapshot in a new container 6).
For multi-MCP deployments, also test Code Mode for MCP — agents load only the tool files they need by exploring a ./servers/<server>/<tool>.ts filesystem, cutting the example Anthropic measured from 150,000 tokens to 2,000 tokens (a 98.7% reduction) 4.
How Big Companies Solve This
Anthropic — Context Engineering as a Discipline
Anthropic positions context engineering as the natural progression of prompt engineering, focused on curating the optimal set of tokens during inference rather than crafting the perfect instruction 1. The company's playbook across its engineering posts converges on three techniques:
- Compaction — distill current context into a high-fidelity summary to avoid exhausting the window. Anthropic shipped a server-side compaction API (beta header
compact-2026-01-12) on January 12, 2026, available across the Claude API, AWS Bedrock, Google Vertex AI, and Microsoft Foundry, with Zero Data Retention support 8, 15. - Structured note-taking — keep a
claude-progress.txtlog and afeature_list.jsonchecklist that survives across sessions; agents only edit thepassesfield 2. - Multi-agent architectures — specialized subagents that work in clean, isolated context windows and return condensed summaries to a lead agent 1.
Anthropic's complementary lever is Code Execution with MCP, which presents MCP servers as code APIs in a filesystem, lets the model load tool definitions on demand, filter large results in code before they reach context, and tokenize sensitive data so PII never enters the model's context window 4.
OpenAI — The Model-Native Harness and Native Sandbox
On April 15, 2026, OpenAI shipped the largest Agents SDK overhaul since launch, adding 5, 6, 14:
- Native sandbox execution — first-class primitives for running agents in isolated compute environments. Developers can bring their own sandbox or use built-in integrations for Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, and Vercel.
- Model-native harness — a runtime layer with configurable memory, sandbox-aware orchestration, Codex-style filesystem tools, and standardized agent primitives.
- MCP tools,
AGENTS.md, skills, shell tool, andapply_patch— the same vocabulary Anthropic and the community converged on. - Provider-agnostic models — the SDK now works with 100+ non-OpenAI LLMs via the Chat Completions API 14.
OpenAI's Manifest abstraction describes an agent's workspace so the same code runs across sandbox providers. State is externalized: if a sandbox container is destroyed, the SDK restores state in a new container from the last snapshot 6. The strategic signal is that OpenAI is shipping the primitives every serious agent team was already reinventing, and is signaling convergence on MCP, AGENTS.md, skills, and apply_patch as a shared cross-vendor vocabulary 5.
Smaller / Adjacent Labs
- E2B is purpose-built for AI-agent code execution, with Python and TypeScript SDKs and Firecracker microVM isolation 12.
- Windsurf's Cascade agent uses a "Flows" model combining RAG-based automatic context retrieval with integrated terminal access, rather than requiring manual
@mentionof files 13. - MCP has seen rapid adoption as the de-facto standard for tool integration, with a growing set of community-built servers 4.
Discussion
Three observations cut across both labs.
First, the bottleneck moved from capability to infrastructure. Both Anthropic and OpenAI explicitly frame their 2026 work as solving the gap between flashy agent demos and production. The fact that the model is no longer the limiting factor is itself a signal — competitive advantage is moving up the stack into harness design, sandbox boundary definition, and evaluation loops 5.
Second, the agent vocabulary is converging. Anthropic introduced MCP and skills, AGENTS.md is a cross-vendor convention OpenAI helped start, and OpenAI's SDK now ships all three. Within 12 months the cross-vendor agent stack may look less like a battle of frameworks and more like Linux distros on a shared kernel 4, 5.
Third, "context window" is a misleading number. Practitioners planning 2026 systems should budget context the way they budget memory in a managed runtime: with a working set, a compaction trigger, and a recovery path. If your agent harness doesn't have all three, it will fail on the first task that takes more than an hour.
Conclusion
The long-horizon-agent problem is the defining AI engineering problem of 2026, and the frontier labs have settled on the same shape of answer: treat context as a finite resource, compact and structure it deliberately, run the agent inside an OS-level sandbox, and ship a model-native harness that handles session continuity, tool loading, and subagent isolation. Anthropic's contribution is the context engineering discipline (compaction, structured note-taking, multi-agent harnesses, code execution with MCP). OpenAI's contribution is the productized harness (native sandbox, standardized primitives, provider-agnostic models). Teams building in this space should pick a side on the harness question (or build their own) but cannot afford to skip any of the four primitives: compaction, sandboxing, structured note-taking, and code-mode tool loading.
Sources
- Anthropic — Effective context engineering for AI agents — https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Anthropic — Effective harnesses for long-running agents — https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents
- Anthropic — Making Claude Code more secure and autonomous with sandboxing — https://www.anthropic.com/engineering/claude-code-sandboxing
- Anthropic — Code execution with MCP: building more efficient AI agents — https://www.anthropic.com/engineering/code-execution-with-mcp
- OpenAI — The next evolution of the Agents SDK — https://openai.com/index/the-next-evolution-of-the-agents-sdk/
- Help Net Security — OpenAI updates Agents SDK, adds sandbox for safer code execution (April 16, 2026) — https://www.helpnetsecurity.com/2026/04/16/openai-agents-sdk-harness-and-sandbox-update/
- Chroma — Context Rot: How Increasing Input Tokens Impacts LLM Performance — https://www.trychroma.com/research/context-rot
- Fazm — Anthropic API release notes, 2026: the changes that actually reach a Claude Code wrapper (server-side compaction API beta header
compact-2026-01-12, Jan 12, 2026) — https://fazm.ai/t/anthropic-latest-api-release-notes-2026 - Cryptobriefing — Anthropic introduces local sandbox mode for Claude Code desktop (84% reduction in permission prompts) — https://cryptobriefing.com/anthropic-claude-code-sandbox-mode/
- Anthropic — How we built Claude Code auto mode: a safer way to skip permissions (93% manual-prompt acceptance rate) — https://www.anthropic.com/engineering/claude-code-auto-mode
- Zylos Research — AI Agent Context Compression: Strategies for Long-Running Sessions — https://zylos.ai/research/2026-02-28-ai-agent-context-compression-strategies/
- Northflank — Best sandboxes for coding agents in 2026 (E2B Firecracker) — https://northflank.com/blog/best-sandboxes-for-coding-agents
- Zylos Research — Context Window Management and Session Lifecycle for Long-Running AI Agents (Windsurf Cascade Flows) — https://zylos.ai/research/2026-03-31-context-window-management-session-lifecycle-long-running-agents/
- AI Automation Global — OpenAI Agents SDK 2026: Native Sandbox and Subagents — https://aiautomationglobal.com/blog/openai-agents-sdk-sandbox-native-agent-primitives-2026
- Anthropic — Compaction at a token threshold (Claude Platform docs; platform availability and ZDR eligibility) — https://platform.claude.com/docs/en/build-with-claude/compaction-threshold
- Liu, N. F. et al. — Lost in the Middle: How Language Models Use Long Contexts (TACL) — https://arxiv.org/abs/2307.03172


