The Streaming Illusion: Why AI Coding Harnesses Lag
On Developer Ergonomics, Inner Loops, and the Perception Lag in Agent Tooling
Sam Saccone · October 2026
1. The Developer Critical Path
For many of us, the terminal is our primary feedback loop. We make a change, we run a test suite, and output streams by. The moment we see a failure we can jump into action, we see, we perceive, and we can act.
The Human-Tool Foundation: The Incremental Perception Gap
The tools our agents are using; test runners, linters, compilers, and build systems were written for humans. Over decades of software development, tooling ergonomics were tuned specifically around humans and cognitive feedback loops: live progress spinners, streaming output, and incremental progress displays..
A core principle of good CLI design for humans is incremental discovery UX: the tool assumes that someone may be actively watching, discovering output as it arrives, and ready to intercept at the moment of emission. When a test suite emits a failure at second 2 of a 200 second run, the human can cull the dead time and jump into action without waiting the additional 198 seconds.
With agents becoming the primary users of these tools here lies the problem: the UX that is good for humans at times falls drastically short from what agents need and from what the harnesses around the agents assume.
Incremental discovery UX from tools assumes that, like humans, agents will be able to see and intercept the opportunity, but tokens are not free and harnesses wrapping the agents treat tools less like streaming data pipes, and more like batched sleep and wait data sources.
The CLI tools assume an active entity is watching and ready to act. The harness assumes the tool is a silent batch job that will notify it when it's done.
This mismatched assumption ends up costing agent productivity greatly.
To put this plainly assume the following:
An agent runs ./run-test. The suite has 12 files. At second 2, src/auth.test.ts emits a fatal assertion failure. But the runner has 10 more suites that take another 20 seconds to finish:
[00:00] RUNS src/index.test.ts
[00:01] PASS src/utils/math.test.ts (142 ms)
[00:02] FAIL src/auth.test.ts
● Auth > Token verification > should refresh expired JWT token
AssertionError: Token expired (status 401)
Expected: 200 OK
Received: 401 Unauthorized
at verifyToken (src/auth.ts:88:13)
at Object.<anonymous> (src/auth.test.ts:42:15)To the human watching the terminal, the failure happened at second 2. Humans can start fixing things right away. To the agent harness, the failure does not exist until the child process exits at second 22.
The Metric We Are Missing: Time to Awareness (TTA)
In engineering productivity, we often measure Time To Debug (TTD), how long it takes from observing a problem to having suspect code running under inspection.
When evaluating agent harnesses, we focus almost exclusively on output metrics: benchmark solve rates, final pass percentages, and total token burn. But we overlook the fundamental input to the system: Time to Awareness (TTA).
Time to Awareness (TTA)
Definition: The time elapsed from when an actionable signal is emitted to when the model's reasoning loop can start working.
My pitch is simple: treating subprocess execution as an atomic, blocking request-response turn breaks the developer and agent critical path. It reduces an agile feedback loop into a high latency batch job.
And the cost isn't linear. In a single interactive session, a 20-second delay is nothing - but agents depend on many cycles and rapid iteration at scale. What feels like a one-off delay of just a few seconds trivially costs years and years of wasted wall-clock agent time against making real progress.
2. What Does This Look Like in Practice?
To understand how widespread this problem is, I built a simplistic test runner (run-test) with the following behavior:
Harness Executions
To ensure an apples-to-apples comparison of harness execution overhead rather than model reasoning latency, we standardized all harnesses on the same model: gpt-4o-mini. For Anthropic's Claude Code (which exclusively interfaces with Anthropic's API), we tested Anthropic's matching lightweight tier (claude-haiku-4-5):
| Harness / Environment | MIN TTA | Latency | Latency vs Human |
|---|
| Human Developer (Baseline) | 2.03 s | 0.00 s | N/A |
| OpenAI Codex CLI | 13.24 s | 11.20 s | 6.5× slower |
Smolagents (v1.26.0) | 24.28 s | 22.24 s | 11.9× slower |
Aider CLI (v0.86.2) | 24.38 s | 22.34 s | 12.0× slower |
| Vercel FX | 25.77 s | 23.73 s | 12.6× slower |
Pi Coding Agent (v1.1.0) | 27.40 s | 25.36 s | 13.4× slower |
| Anthropic Claude | 28.72 s | 26.68 s | 14.1× slower |
| Anomaly OpenCode | 30.29 s | 28.25 s | 14.9× slower |
*Note: Claude Code exclusively targets Anthropic's Messages API, so Anthropic's equivalent low-cost tier (claude-haiku-4-5) was selected to match the parameter and latency profile of gpt-4o-mini.
During those 20 to 28 seconds, every single harness was waiting. The human saw the failure in real time at second 2, but the agent harnesses were forced to wait until the process finished before becoming aware.
3. Where Does the Time Go? Tracing the Code
Let's look at the process tree and trace the actual source code paths across these runtimes to see how this happens.
Case 1: Vercel FX
| 580 | const started_ms = io_mod.milliTimestamp(); |
| 581 | const zio = io_mod.getIo(); |
| 582 | while (true) { |
| 583 | entry.mutex.lockUncancelable(zio); |
| 584 | const terminal = entry.isTerminal(); |
| 585 | entry.mutex.unlock(zio); |
| 586 | if (terminal) break; |
| 587 | |
| 588 | const elapsed = io_mod.milliTimestamp() - started_ms; |
| 589 | if (elapsed >= input.yield_time_ms) break; |
| 590 | |
| 591 | // Thread sleeps: no pattern inspection, no awareness |
| 592 | io_mod.sleep(10 * std.time.ns_per_ms); |
| 593 | } |
| 594 | |
| 595 | const prepared = |
| 596 | try self.prepareSnapshot(alloc, entry.execution_id); |
| 597 | return prepared; |
Look closely at what is happening:
entry.isTerminal() evaluates to false at second 2 because the child process is still running.elapsed >= yield_time_ms (30,000ms) won't trigger for another 28 seconds.- So the dispatch thread simply sleeps in a 10ms loop.
Meanwhile, a separate background worker thread streams stdout chunks to stderr for human viewing. The human sees the failure immediately. But the LLM turn is completely asleep in managed_execution.zig until second 22 when the process terminates.
Case 2: Anomaly OpenCode (TypeScript / Effect Runtime)
| 260 | // Foreground command BLOCKS the Effect Fiber: |
| 261 | const result = yield* jobs |
| 262 | .block({ id: job.id, sessionID: context.sessionID }) |
| 263 | .pipe(Effect.onInterrupt(() => |
| 264 | jobs.cancel(job.id).pipe(Effect.ignore) |
| 265 | )) |
| 266 | |
| 272 | return yield* Deferred.await(settled) |
Subprocess output is streamed chunk-by-chunk to disk in sh_<id>.out. The OpenCode TUI polls this file at 1 Hz via setInterval, so the developer sees the failure within 1 second.
Yet the Effect fiber executing the agent turn has no tap into this stream. It remains suspended on Deferred.await(settled) until handle.exitCode settles:
Human TUI Polling Path (1 Hz):
Proc -> sh_*.out (t=2s) -> setInterval poll -> Dev spots FAIL at t=2.5s
Agent Fiber Execution Path (Synchronous Barrier):
Tool -> jobs.start -> jobs.block -> Fiber suspended on Deferred.await(settled)
[... 20.2s of passing test suites running in background ...]
Proc exits (code 1) -> handle.exitCode settles -> Fiber resumes at t=22.2s -> LLM reads file
When we inspected the trace from our headless OpenCode tester, look at the timeline:
t = 2.000s: FAIL src/auth.test.ts written to sh_*.out.t = 22.197s: handle.exitCode resolves (code 1).t = 22.208s: tool_use completes.t = 23.855s: The model emits its next tool call: read("src/auth.test.ts").
The model knew exactly what file to inspect from the error emitted at second 2. But the harness delayed that action by 21.85 seconds.
Case 3: Anthropic Claude Code CLI (@anthropic-ai/claude-code)
| 1 | { |
| 2 | "name": "Bash", |
| 3 | "description": "execute shell commands", |
| 4 | "input_schema": { |
| 5 | "type": "object", |
| 6 | "properties": { |
| 7 | "command": { |
| 8 | "type": "string", |
| 9 | "description": "The command to execute" |
| 10 | }, |
| 11 | "timeout": { |
| 12 | "type": "number", |
| 13 | "description": "Timeout in milliseconds" |
| 14 | } |
| 15 | } |
| 16 | } |
| 17 | } |
In Claude Code's agent loop, stdout is rendered to the terminal TTY in real time. But under the Anthropic Messages API, the LLM turn cannot proceed until a tool_result message is appended to the message array.
The runtime holds this message until the child process terminates (exitCode != null). In the test with the real claude binary, Claude Code blocked synchronously for 23.31 seconds before dispatching the tool_result callback back to /v1/messages.
Case 4: OpenAI Codex CLI (@openai/codex)
| 1 | { |
| 2 | "name": "exec_command", |
| 3 | "description": "Runs command in PTY session", |
| 4 | "parameters": { |
| 5 | "properties": { |
| 6 | "cmd": {"type": "string"}, |
| 7 | "yield_time_ms": { |
| 8 | "type": "number", |
| 9 | "description": "Defaults to 10000 ms" |
| 10 | } |
| 11 | } |
| 12 | } |
| 13 | } |
Note the key parameter: yield_time_ms, which defaults to 10,000 ms (10 seconds).
Codex is architecturally superior to harnesses that block until exit: if a command runs longer than yield_time_ms, Codex cuts an interim chunk and yields while leaving the process running in the background.
However, because yield_time_ms is fixed to 10 seconds, Codex sat on the failure for another 8 seconds after it appeared at second 2, unblocking at t = 10.17s. While Codex avoided waiting the full 22 seconds, it was still 5.0× to 6.5× slower than the human lower bound
Case 5: Aider & Smolagents (Python Runtimes)
| 75 | output = [] |
| 76 | while True: |
| 77 | chunk = process.stdout.read(1) |
| 78 | if not chunk: break |
| 79 | print(chunk, end="", flush=True) |
| 80 | output.append(chunk) |
| 81 | |
| 82 | process.wait() # <-- Blocks until exit! |
| 83 | return process.returncode, "".join(output) |
Line 4 prints stdout directly to the terminal with flush=True. The human watching the screen sees the error within 2 seconds. But line 7 calls p.wait(). The function cannot return until the remaining 20 seconds of passing tests are complete.
| 139 | // Stream stdout and stderr: |
| 140 | child.stdout?.on("data", onData); |
| 141 | child.stderr?.on("data", onData); |
| 142 | if (signal) { |
| 143 | if (signal.aborted) onAbort(); |
| 144 | else signal.addEventListener( |
| 145 | "abort", onAbort, { once: true }); |
| 146 | } |
| 147 | // Wait for child process to terminate: |
| 148 | const exitCode = |
| 149 | await waitForChildProcess(child); |
| 150 | return { exitCode }; |
While onData forwards output buffers to scheduleOutputUpdate() for rendering in the developer's terminal, the tool's execute() method is suspended awaiting waitForChildProcess(child).
When executed with gpt-4o-mini against ./run-test, Pi's TUI displayed the test failure in real time at second 2. But the agent loop was idle for another 25.36 seconds while the remaining test suites finished, only becoming model-aware at t = 27.40s (13.4× slower than human response time).
4. The Real Cost of Agent Lag
Why does a 20-second delay matter so much? In isolation, 20 seconds is nothing. But in the inner loop, this lag generates compounding friction:
1. The Flywheel Penalty at Scale
Agent driven software engineering is fundamentally a game of rapid cycles and high-frequency iteration loops. An agent attempting to solve a subtle bug or navigate an unfamiliar codebase routinely goes through many cycles of thinking, patching, and testing.
As a one-off command in an interactive terminal, waiting an extra 20 seconds might seem like just a few seconds. But these agents depend on many cycles and rapid iteration at scale:
- In a single 30-turn agent session, a 23-second blind wait burns 11.5 minutes of pure, unrecoverable idle latency.
- At human speed, that entire diagnostic loop concludes in under 1 minute.
- Across thousands of agents running millions of iterations across an engineering organization, what seems like just a few seconds as a one-off ends up accounting for years and years of wasted wall-clock agent time stranding compute infrastructure, driving up cloud costs, and stalling execution.
2. Context Window Pollution
Because the command runs to completion, the output buffer captured includes all 12 test suites, timing logs, snapshot summaries, and passing banners. Instead of returning a compact 10-line error snippet, the tool payload contains hundreds of lines of noise, bloating context windows and increasing prompt cache churn.
3. The Broken Assumption of Incremental Discovery
Incremental discovery UX works for humans because our eyes and cognitive loops operate out-of-band from the shell process. We don't freeze our brain waiting for the shell to finish. But when an agent harness wraps a tool in a synchronous wait loop, incremental discovery transforms into dead wait time.
5. The path forward
Why are agent harnesses trapped waiting for subprocesses to terminate?
A simplistic diagnosis is: "The agent's terminal is non-interactive." Developers assume that if we just wrap subprocess execution in a pseudo-terminal (PTY) or full TTY emulation, the problem will solve itself.
But that misses the flaw. It is not just that the terminal is non-interactive.
It is that process exit is treated as the only meaningful signal to dump the buffer to consumers.
In the classical POSIX paradigm inherited by AI harnesses:
- The tool spawns a process.
- The process runs, streaming stdout or buffering it internally.
- The harness hits a hard synchronization barrier:
waitpid(), proc.wait(), or child.on('exit'). - Only when the process exits does the harness treat the output buffer as ready to dump to the model.
Making process death the prerequisite for data consumption is the fundamental bottleneck. If a command runs for 30 seconds or 5 minutes, the agent cannot consume data until the process terminates.
The Pragmatic Precedent: --bail
Before designing speculative solutions, it is worth asking: don't test frameworks already have a built-in mechanism to abort on failure?
Yes. In CI and developer workflows, we have a long-standing production standard for this: fail-fast flags.
| Framework | Fail-Fast Flag | Behavior on First Failure |
|---|
| Jest | jest --bail (or --bail=1) | Aborts worker pool and exits immediately (<150ms) |
| Vitest | vitest run --bail 1 | Cancels remaining suites and exits immediately |
| Pytest | pytest -x | Halts execution instantly upon first failure |
| Cargo Test | cargo test -- --no-fail-fast | Halts multi-target runs on initial failure |
| Go Test | go test -failfast ./... | Halts execution of further packages on failure |
When a test suite is executed with --bail, the subprocess self-terminates at second 2 natively, unblocking every single harness immediately without needing runtime rewrites.
However, relying solely on --bail is insufficient as a general solution for agentic workflows:
- Agents frequently generate ad-hoc commands (
npm test, make test, custom shell scripts) where fail-fast flags are omitted. - Non-test tools (compilers, bundlers, long-running database migrations) lack uniform fail-fast flags.
- Most fundamentally,
--bail still bows to the original constraint: it relies on killing the process to deliver information.
Solving this at scale requires dual remediation: harnesses must get better, and tools need to evolve for more execution contexts.
Part 1: Harnesses Must Evolve
Harnesses cannot remain passive wrappers waiting on waitpid() with passive checks. They need to start actively inspecting in-flight execution and treat subprocesses as reactive event streams. Rather than relying purely on time-based timeouts, harnesses should allow agents or tools or skills to declare failure watchpoints:
| 1 | { |
| 2 | "command": "./run-test", |
| 3 | "early_yield_patterns": [ |
| 4 | "FAIL ", |
| 5 | "AssertionError", |
| 6 | "panic:", |
| 7 | "BUILD FAILED" |
| 8 | ] |
| 9 | } |
As output chunks flow through the runtime stream pump, a sliding-window regex scanner matches patterns. If FAIL is detected, the harness unblocks the tool turn immediately with status early_yield, backgrounding the child process at t = 2.03s rather than waiting for the exit.
/btw Branch Spawning on output Match
When a failure pattern is matched in a slow test runner:
- The harness leaves the test runner executing in the background to collect full run info.
- Concurrently, the harness forks a speculative investigation with the partial logs.
6. In Recap
The tools that our agents are using were written for humans. But with agents as the new primary users, the UX that was good for humans falls drastically short from what agents need and from what the harnesses wrapping those agents assume.
To build agents that truly code at the speed of tools, we must rethink both sides of the contract: harnesses must stop treating streaming subprocesses as blocking batch turns, and developer tooling must evolve beyond human-centric incremental discovery toward agent-first reactive event protocols. Giving agents the agility to see failures as they happen and act before the terminal goes quiet is the single highest-leverage UX unlock in modern AI tooling.