Get early access

The 1M Output Horizon

Engineering around deep reasoning trajectories and silent failure modes

Summary

Google DeepMind's Gemini 4 Argon raises Gemini's maximum output from 64,000 tokens to what Google calls an "industry-leading" 1,000,000 tokens, enough to "generate hundreds of thousands of tokens in a single trajectory" 1. Google's examples include migrating large C and C++ codebases to Rust, with work underway on the 800K-line Fuchsia Zircon kernel, and memory optimizations across its data centers that freed over 300 TiB of RAM 1. Generation budgets that large break assumptions most AI systems were built on. A single run can stream for hours, latency stops being predictable, an early mistake can carry through hundreds of thousands of tokens, and a dropped connection late in a run throws away everything before it. Keeping the reasoning trace readable enough to monitor also gets harder as trajectories grow. Teams adopting long-trajectory models need checkpointed streaming, restartable orchestration and monitoring that runs out of band.

Key Takeaways

  • The output ceiling moved. Argon's single-trajectory output limit goes from Gemini's previous 64K to 1M tokens, alongside a state-of-the-art 77.9% on DeepSWE v1.1, a tie for first place on CWE-bench v1 (68%) and the top spot on the Vals Index (68.9%) 1, 2.
  • Synchronous request-response does not survive hours-long runs. Hundreds of thousands of output tokens take hours to stream, long enough for client timeouts, maximum connection lifetimes and routine deploys to cut a run off.
  • Errors compound over long trajectories. Without intermediate checks, an early wrong assumption in a 100K+ token run can carry through everything generated after it.
  • Chain-of-thought monitoring is useful but fragile. Rohin Shah and Anca Dragan argue that chain-of-thought (CoT) monitoring is "one tool among many" but "an exceptionally useful one", and that penalizing what the CoT shows can teach a model to hide it better 3. OpenAI reports a "substantial decrease in chain-of-thought monitorability" in GPT-6 Astra 4.
  • The fixes are distributed-systems fixes. Long-horizon runs need streaming with stall detection, durable checkpoints a run can restart from, and monitors that read the reasoning trace without feeding back into the generator.

Problem Background

Context windows have grown much faster than output limits. Gemini's previous output limit was 64K tokens; GPT-5 and current Claude models allow 128K 1, 6, 7. On September 30, 2026, Google DeepMind announced Gemini 4 Argon with a 1M-token output limit 1. It is rolling out first to trusted cyber defenders through Google's Fairwind Program, then to paid API customers and Google AI Ultra subscribers, at an introductory price of $2 per million input tokens and $10 per million output tokens (rising to $4 and $20 after the introductory period) 1, 8.

Google describes what long trajectories made possible internally 1:

  • Codebase migration. Argon agents are migrating C and C++ code to Rust, "scaling from tens of thousands of lines in core libraries like re2, libgav1" up to 800K+ lines for the Fuchsia Zircon kernel. The Zircon work is still going through automated and manual auditing, emulation testing and review before it reaches production. For libgav1, Argon replaced 32K lines of SIMD code in an existing Rust port with safe Rust the compiler vectorizes automatically, found through many rounds of profile-guided experiments, and the result runs 2.7x faster than that port.
  • Fleet memory optimization. A team of Argon agents analyzed fleet-wide profiling telemetry and applied memory optimizations across Google's data centers, freeing over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
  • Software engineering and security. Argon sets a new state of the art on DeepSWE v1.1 at 77.9% and ties for first place on CWE-bench v1 at 68% 1, 2.

For teams building on models like this, hundreds of thousands of sequential output tokens bring new failure modes:

  1. Connection volatility. At an assumed 50–100 tokens per second (Google has not published Argon's output speed), 500,000 tokens take roughly 1.5 to 3 hours to stream. Over that window, client read timeouts, load balancer and proxy limits on total request duration, and ordinary pod restarts or deploys all become likely. A run that cannot resume loses everything it generated before the drop, and that output has already been paid for.
  2. Compounding drift. When a model reasons across 200,000 tokens without compiling, running or checking anything in between, an early wrong assumption can carry through the rest of the run and produce output that looks plausible but is built on the wrong premise.
  3. Monitorability. In The case for reasoning transparency, Rohin Shah and Anca Dragan argue that a model's chain of thought gives "a window into its reasoning" that can be used to monitor for scheming and deception, and that penalizing bad behavior in the CoT "can easily result in ways to hide the cheating better" 3. The essay is the authors' view, not Google's official position. OpenAI's GPT-6 Astra system card shows the problem in practice: Astra produces "shorter CoTs that often omit or weaken the evidence the monitor needs", and when told it is being monitored it "spontaneously generates substantially fewer chain-of-thought tokens" 4. The UK AI Security Institute, quoted in the card, found it "reasons in a compressed style" with a "greatly increased ability to reason within a single forward-pass". OpenAI also says it has not seen evidence of steganographic CoT reasoning 4.

Proposed Experiments/Examples

Long-trajectory models cannot be treated as ordinary request-response endpoints. The two sketches below cover restartable streaming and out-of-band monitoring. Both run as written on Python 3.11+.

Experiment 1: Checkpointed long-horizon streaming

The model is prompted to close each unit of work, such as a file or a plan step, with a small JSON marker. The orchestrator persists everything up to each marker to a durable store, so a dropped connection or a failed check restarts from the last checkpoint instead of from token zero.

import asyncio
import json
import time
from typing import Any, Optional, Protocol

OPEN, CLOSE = "<checkpoint_marker>", "</checkpoint_marker>"


class CheckpointStore(Protocol):
    def put(self, run_id: str, checkpoint: dict[str, Any]) -> None: ...
    def latest(self, run_id: str) -> Optional[dict[str, Any]]: ...


class LongHorizonStreamOrchestrator:
    """
    Persists semantic checkpoints out of a long token stream and flags stalls.
    The model is prompted to close each unit of work (a file, a plan step) with
    <checkpoint_marker>{"checkpoint_id": ..., "phase": ...}</checkpoint_marker>.
    """

    def __init__(self, run_id: str, store: CheckpointStore):
        self.run_id = run_id
        self.store = store
        self.buffer = ""
        self.output_tokens = 0
        self.last_chunk_at = time.monotonic()

    async def on_chunk(self, text: str, tokens: int) -> None:
        # `tokens` comes from the provider's usage data; splitting text on
        # whitespace counts words, not tokens.
        self.buffer += text
        self.output_tokens += tokens
        self.last_chunk_at = time.monotonic()
        self._commit_complete_checkpoints()

    def _commit_complete_checkpoints(self) -> None:
        while (start := self.buffer.find(OPEN)) != -1:
            end = self.buffer.find(CLOSE, start)
            if end == -1:
                return  # marker still streaming in
            artifact = self.buffer[:start]
            raw = self.buffer[start + len(OPEN):end]
            # Consume the marker either way, so one malformed marker cannot
            # block every checkpoint after it.
            self.buffer = self.buffer[end + len(CLOSE):]
            try:
                meta = json.loads(raw)
            except json.JSONDecodeError:
                self.buffer = artifact + self.buffer  # keep the work, drop the marker
                continue
            self.store.put(self.run_id, {
                "checkpoint_id": meta.get("checkpoint_id"),
                "phase": meta.get("phase", "reasoning"),
                "output_tokens": self.output_tokens,
                "artifact": artifact,
            })

    def resume_point(self) -> Optional[dict[str, Any]]:
        """Where a restarted run picks up after a drop or a failed check."""
        return self.store.latest(self.run_id)

    async def supervise(self, max_idle_sec: float = 120) -> None:
        while True:
            await asyncio.sleep(min(10, max_idle_sec))
            if time.monotonic() - self.last_chunk_at > max_idle_sec:
                raise TimeoutError(f"no tokens for {max_idle_sec}s")

Two details matter here. The token count comes from the provider's usage data, because splitting streamed text on whitespace counts words, not tokens. And a malformed marker is consumed and skipped, so one bad marker cannot stop every checkpoint after it from being saved. The store behind CheckpointStore should be durable (a database row or an object-store key per checkpoint); an in-memory dict loses everything with the process.

Experiment 2: Out-of-band chain-of-thought monitor

A separate monitor reads the raw reasoning trace and flags phrases worth a closer look. Its findings go to a human or a policy engine, never back into the generator's prompt or reward, so the generator gets no signal to write around it.

import re
from typing import Any


class OutOfBandCoTMonitor:
    """
    Reads the reasoning trace on a separate path from the generator. Nothing it
    flags is fed back into the generator's prompt or reward, so the model gets
    no signal to write around it.

    Keyword rules are a floor, not a control: a model that is hiding something
    will not say so in plain words. Pair this with an LLM judge over the same
    window and with a log of every tool call the agent makes.
    """

    PATTERNS = [
        re.compile(r"hide\s+(this|it)\s+from\s+the\s+(evaluator|monitor|user)", re.I),
        re.compile(r"(bypass|disable|escape)\s+the\s+sandbox", re.I),
        re.compile(r"without\s+(asking|telling)\s+the\s+user", re.I),
        re.compile(r"(this|we)\s+(is|are)\s+(probably\s+)?being\s+(tested|evaluated)", re.I),
    ]

    def __init__(self, critical_after: int = 3):
        self.hits = 0
        self.critical_after = critical_after

    def inspect(self, window: str) -> dict[str, Any]:
        matches = [m.group(0) for p in self.PATTERNS for m in p.finditer(window)]
        if not matches:
            return {"status": "ok"}
        self.hits += len(matches)
        level = "critical" if self.hits >= self.critical_after else "warning"
        return {"status": "flagged", "level": level, "matches": matches}

Keyword rules are the weakest layer. A model that is concealing something will rarely say so plainly, and Astra's system card shows reasoning can simply get shorter under monitoring 4. Treat a monitor like this as a cheap first filter, and back it with an LLM judge over the same window and a log of every tool call the agent actually makes.

How Big Companies Solve This

  1. Google DeepMind: gated rollout and sealed sandboxes. Google is releasing Argon first to vetted cyber defenders through the Fairwind Program, with access "strictly controlled", before paid API customers 1, 8. Its safety work includes "hardening our sandboxed environments by isolating and sealing them before high-risk training or evaluations begin", and the Zircon migration goes through automated and manual auditing, emulation testing and review before production 1. Shah and Dragan argue that developers who move away from readable chain of thought "should provide a robust and convincing demonstration that their new methods are comparably monitorable" 3.
  2. OpenAI: inference-time scaling and layered cyber safeguards. With o3 and o4-mini, OpenAI reported that "the models' performance continues to improve the more they're allowed to think" 5. Its GPT-6 Astra system card pairs the monitorability findings above with real-time cyber safeguards, including escalating higher-risk cyber content for deeper review and blocking responses above its policy boundary 4.
  3. Wiz: security work on real software. Wiz is using Argon in its Scan for Good initiative, which finds and remediates high-risk exposures in critical public infrastructure for free. In an early demonstration, the model uncovered a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide, one that previous frontier models had missed 1. On Wiz's internal black-box penetration-testing benchmark, the task covers discovering the attack surface, identifying vulnerabilities and producing proof-of-concept evidence 1.

Discussion

A 1M-token output limit moves the hard part of AI engineering from wording the prompt to orchestrating a long-running job. When one call can run for 45 minutes or more, teams face new trade-offs:

  • Many short turns or one long trajectory. A ReAct-style loop of 20 turns at 2K tokens each gets checked by the outside world after every turn, but re-reads context each time and loses the thread between turns. A single 500K-token trajectory keeps everything in one pass, but without checkpoints an error at token 50,000 is only caught at the end.
  • Don't pay the model to think less. Shah and Dragan argue that pressure on the chain of thought can push reasoning out of view 3; Astra's shorter CoTs under monitoring are a live example 4. Systems that reward fewer reasoning tokens, or punish what the trace reveals, risk losing the one readable view they have.
  • Failures cost real money. At Argon's introductory prices, 1M output tokens cost $10, rising to $20 after the introductory period, before input costs 1. A dropped or derailed run is no longer a sub-cent failure, and across an automated fleet, runaway trajectories show up directly on the bill. Cap output per run, alert on stalls and resume from checkpoints rather than starting over.

Conclusion

Gemini 4 Argon's 1M-token output limit makes far longer single runs possible, from multi-hundred-thousand-line code migrations to fleet-wide optimization. Using it well takes engineering more than prompting: streaming that survives drops, checkpoints a run can restart from, and monitoring that reads the reasoning without teaching the model to hide it. The model can now work for hours in one go; the system around it has to be built to last as long.

Sources

  1. Google — "Gemini 4 Argon: our next era of frontier intelligence", Koray Kavukcuoglu (Sept 30, 2026). https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
  2. Google DeepMind — Gemini models, benchmark table. https://deepmind.google/models/gemini/
  3. Google DeepMind Institute — "The case for reasoning transparency", Rohin Shah and Anca Dragan (Sept 16, 2026). https://institute.deepmind.com/essays/the-case-for-reasoning-transparency/
  4. OpenAI — GPT-6 Astra System Card, monitorability (Sept 3, 2026). https://deploymentsafety.openai.com/gpt-6-astra/monitorability
  5. OpenAI — "Introducing OpenAI o3 and o4-mini" (Apr 16, 2025). https://openai.com/index/introducing-o3-and-o4-mini/
  6. OpenAI — GPT-5 model page, 128K max output tokens. https://developers.openai.com/api/docs/models/gpt-5
  7. Anthropic — Claude models overview, max output. https://platform.claude.com/docs/en/about-claude/models/overview
  8. Google DeepMind — Fairwind Program. https://deepmind.google/fairwind-program/

Keep reading

All posts →

Be first in when doors open.

Get early access