Image: leehanchung.github.io · rights & removal
Executive Summary
Agentic AI evaluation requires building an entire system, not just measuring a final output. The focus must shift from single-turn metrics to evaluating the entire process—the data plane that runs the agent alongside the control plane that governs experimentation and decision-making. This infrastructure encompasses Rollouts (agent episodes), traces capturing step-by-step actions, memory management, environment changes, and mechanistic interpretability. Evaluating agents involves five surfaces: the final output, the trace of execution, memory usage, the resulting state changes in the environment, and internal mechanistic properties.
The core difficulty lies in moving from simple metrics to actionable insight. A single score is insufficient because agent failure can occur at any step—a bad tool call, a polluted memory update, or an incorrect state delta. Therefore, evaluation must capture these granular details through structured data, such as OpenTelemetry-style spans for traces and detailed state diffs. The infrastructure serves the entire development lifecycle, from training and experimentation to production monitoring, allowing for continuous learning and debugging.
Facts Only
* Evaluation infrastructure requires a control plane (task definition, configuration) and a data plane (agent execution and recording).
* Rollouts are core artifacts, representing an agent's interaction with an environment via tools, memory, and workflows.
* The data plane includes the model, harness, runtime, tools, memory, environment state, traces, outputs, logs, snapshots, and state deltas.
* Evaluation surfaces include Output, Trace, Memory, Environment, and Mechanistic Interpretability.
* Traces must be structured records of every step (tool calls, observations, state deltas) to enable failure attribution.
* Memory evaluation is necessary because memory pollution can silently change agent behavior by polluting context with unintentional information.
* Environment evaluation requires capturing state deltas (file changes, database updates) at each step, not just the terminal state.
* Evaluation infrastructure should support experimentation-driven methods like perturbation tests and ablation tests to isolate variable contributions.
* Checkpointing is necessary for long-horizon evaluations to allow for resuming, replaying, or branching experiments.
Full Take
The transition from single-turn evaluation to agent evaluation necessitates a fundamental shift from static measurement to dynamic system control. The central implication is that agentic systems are coupled environments where the output is an emergent property of the entire system configuration, not just the model itself. This demands treating agent development as an experimental science governed by causal inference rather than simple correlation.
The concept of "agent worlds"—where memory and mutable state exist—reveals the brittle nature of relying solely on extrinsic scores. The accumulation of evaluation debt occurs when convenience masks this coupling: a single aggregate score fails to capture the distributed failure modes across the agent's operational history. The focus must therefore be on building durable infrastructure that resolves these causal relationships by treating state and traces not as secondary artifacts, but as the primary experimental variables.
The pattern being established is the necessity of an internal language for system observability—a structure where external performance metrics are directly traceable back to specific components (model vs. harness vs. memory) and specific transitions (state deltas). This pushes evaluation beyond simple benchmarking into mechanistic oversight, forcing a consideration of how state mutations cause behavior shifts. The cost of avoiding this is the accumulation of "cargo cult evaluation," where surface-level scores mask deeper systemic regressions, which ultimately erodes user trust and requires a system capable of evidence-based decision-making over mere aggregated results.
Bridge Questions: If the agent learns an unintended but highly effective internal state (memory), how can external verification methods reliably detect that learned behavior without access to the model's internal computations? What are the necessary constraints for defining a "safe" or "justified" memory write within the control plane? How can organizations operationalize this system to transition from reactive debugging to proactive, continuous alignment driven by replayable state infrastructures?
From the original · Han, Not Solo
Most conversations about evals collapse into which SaaS tool to buy, which metrics to track, which LLM-as-a-judge prompt to slap on, or which single headline benchmark score to worship. SWE-bench percentage.Read the full story at leehanchung.github.io
Sentinel — Human
The article presents a highly structured, expert-level argument detailing the necessity of building comprehensive evaluation infrastructure for agentic AI systems, emphasizing state tracking over single metrics.
