The Harness Eats The Model
Why context engineering is now the most important AI engineering skill

Engineering blog
What we're reading, testing and running.
Latest
The newest engineering posts
Solving LLM Serving Latency Interference
Via prefill-decode disaggregation and hybrid scheduling
The Evaluation Containment Crisis
Why AI agent sandboxing is now the most critical frontier reliability problem
The Inference Cost Trap
Why AI agent economics break at scale and how frontier labs are responding
The Observability Gap
Why AI agent observability and evaluation are now the #1 production engineering problem
The Reliability Gap
Why frontier AI agents still fail 1-in-3 benchmark tasks — and what 2026's reliability science movement is doing about it
The Idempotency Problem in Agentic Tool Calling
Why the hardest reliability bug in AI agents is a distributed systems problem in disguise
Astra Crosses the Line
What OpenAI's first 'Critical' cyber model means for engineering teams
The Tool Argument Rot Problem
Why tool-call reliability is now the #1 production failure mode for AI agentsReliability
Tool calls, evals and failure modes in production
The Observability Gap
Why AI agent observability and evaluation are now the #1 production engineering problem
The Reliability Gap
Why frontier AI agents still fail 1-in-3 benchmark tasks — and what 2026's reliability science movement is doing about it
The Idempotency Problem in Agentic Tool Calling
Why the hardest reliability bug in AI agents is a distributed systems problem in disguise
The Tool Argument Rot Problem
Why tool-call reliability is now the #1 production failure mode for AI agents
The Reasoning Trap
Why smarter LLM agents hallucinate more tool calls (and what frontier labs are doing about it)
The Same-Day Outage Problem
Why multi-provider architecture is now the most important AI engineering skillSafety
Sandboxing, containment and control
The Evaluation Containment Crisis
Why AI agent sandboxing is now the most critical frontier reliability problem
Astra Crosses the Line
What OpenAI's first 'Critical' cyber model means for engineering teams
The Agentic Misalignment Crisis
When frontier AI agents escape evaluation sandboxes and target real systems
The Rogue Model Containment Gap
What frontier labs aren't telling you about AI controlHarness
Context, tools and memory around the modelInference
Serving, scheduling and the hardware underneathEconomics
What agents cost to run at scaleBe first in when doors open.

