Image: cdn-uploads.huggingface.co · rights & removal
The Agent Said It Was Done. The Database Disagreed.
Reporting by Hugging Face BlogRead the original at huggingface.co
Executive Summary
Facts Only
* An agent executed nine tool calls in handling a customer refund request.
* The required end state for the ticket was "hold," but the final status recorded was "solved."
* An AI grader of tool calls found nine well-formed calls, and a separate check on database writes also found correct.
* A gap exists between the agent's output and the actual terminal backend state, which is measured by ThinkingBox.
* Across 121,680 valid trials across 12 LLM models, 79,853 attempts failed executable checks.
* In these failures, 77.61% involved wrong field values, 43.30% involved unintended side effects, and 25.36% involved missing required effects.
* Pass@1 scores vary widely across models, with Claude Opus 5.5 leading overall at 67.16%.
* Consistency is measured by pass@20; GPT-6 Astra retains 78% of its single-attempt rate, and Claude Opus 5.5 and 5 retain 71%.
* Cost per successful task attempt was calculated by dividing the cost of one run by the number of successful attempts.
* The lowest cost per dependable task is $6.80 for GPT-5.4, while Claude Opus 5.5 costs $7.80 per dependable task.
Full Take
From the original · Hugging Face Blog
Figure 1: ThinkingBox runs an agent against isolated MCP tool sessions, then grades the terminal backend state and side effects it leaves behind. From our ThinkingBox paper.Read the full story at huggingface.co
Sentinel — Human
This text appears to be a legitimate technical report or blog post detailing an experimental framework (ThinkingBox) and its empirical findings on AI agent reliability, exhibiting the structure and depth of human-authored research communication.
