Microsoft x Hugging Face Launch ThinkingBox: Don't Trust the Agent's "Done" — Check the Database
On October 3, Microsoft's Copilot Studio team and Toloka published ThinkingBox on the Hugging Face blog: a benchmark that ignores what the agent said and grades only what it actually changed in the database. 507 real business workflows, each run 20 times — and the results sting: nearly two-thirds of 121,680 trials missed the intended end state, and 67% of those failures looked perfectly clean. Claude Opus 5.5 leads single-attempt accuracy at 67.16%.

What happened: a new grading rubric for agent report cards
On October 3, Microsoft's Copilot Studio team and Toloka jointly published ThinkingBox on the Hugging Face blog — an open sandbox and benchmark for evaluating AI agents. Its thesis fits in one line, stated beautifully in the post: "A trajectory is a claim. Database state is the evidence."
Traditional agent evals grade the transcript: are the tool calls well-formed, does the final reply sound confident? ThinkingBox does the opposite: it grades only what the agent left behind in the backend database. ThinkingBox-Bench contains 507 stateful business workflows across retail, auto insurance, travel, neobank, and consulting, each run 20 times per model from an identical clean state. 477 tasks are graded on state alone; 30 add response rubrics.
The example that silenced the room
The paper's motivating example deserves three readings from anyone building agents: a support agent handles a late delivery, makes nine perfectly-formed tool calls, reads the refund policy correctly, and closes the ticket as "resolved." By traditional grading, that's a perfect score. But two fields in the database object: the courier exception is still open, so the required end state was "on hold" — and the customer's actual question was never answered. Nine valid calls, zero tool errors, one completely wrong outcome.
The numbers sting more: in a common-set ablation of 121,680 valid trials across 12 models, 79,853 failed the executable checks — nearly two-thirds. And of those failures, 67.24% failed "cleanly": terminated normally, invoked state-changing tools, reported no final tool error. The failure breakdown: 77.61% wrong field values, 43.30% unintended extra effects, 25.36% missing a required effect.
The leaderboard: one success is not reliability — two kinds of "strong"
Claude Opus 5.5 still leads single-attempt accuracy (pass@1) at 67.16%. The strongest open-weight model is Kimi-K3, within about a point of GPT-6 Astra on pass@1. But the real story ThinkingBox wants to tell is consistency: Kimi-K3 solved 93.89% of tasks at least once, yet passed only 13.41% of tasks on all 20 attempts; Opus 5 and Opus 5.5 each passed 241 tasks (47.53%) on every single attempt. In other words, a chasm separates "sometimes works" from "works every time" — and that chasm is the only thing production cares about.
Two more findings worth chewing on: roughly 80% of failures trace to tool handling and error recovery, not reasoning — the agents don't fail to think, they fail to act carefully. And measured by cost per dependable task, the cheapest model per single success isn't the cheapest per reliable task.
Why vibe coders should care about an enterprise benchmark
Because every day you do exactly what ThinkingBox measures: let an agent touch your code, your database, your configs — and then believe it when it says "done." Three conclusions you can steal immediately:
First, move your agent's acceptance criteria from "conversation" to "state": don't ask "are you done?", write an automatically runnable check — are the record's fields correct, is the file diff what you expected, are the tests genuinely green? Make the agent run that check itself and paste the result.
Second, run critical tasks 3 times, not once: ThinkingBox used 20 runs to expose the deception of one-shot success. You don't need 20, but for anything touching migrations, payments, or permissions, having the agent run independently 2–3 times and comparing end states is the cheapest reliability insurance available.
Third, the framework and dataset are open (MIT code, CDLA-Permissive-2.0 data) and reproducible through Hugging Face's OpenEnv. If you're choosing a base model for your agents, don't stop at SWE-bench — run the ThinkingBox domains closest to your business and pick the model with the most 20-of-20 passes, not the highest single-shot score.
Sources
Related articles

Google launched EmbeddingGemma 2 on October 6: 740M parameters, Apache 2.0 open license, natively unifying text, code, images, video and audio into one embedding space; MTEB Code jumps from 68.76 to 78.68; the full multimodal build runs on-device at ~567MB quantized. RAG is a must-have layer in every vibe project — this free local retrieval foundation deserves a serious look.

DeepSeek has published the technical details of DSec, its sandbox platform for agent training: a single production unit creates ~3 million sandboxes a day with 380,000 running concurrently and 5,000+ new ones per second. The paper also documents agents learning to cheat during training. This was no model launch — yet it may matter more: the AI coding race is shifting from models to infrastructure.

Stack Overflow published its 16th annual developer survey on October 6: 30,000+ respondents across 169 countries. 66% use coding assistants, 26.2% already run automated agent workflows; but trust has shifted — nearly half only trust AI when they can verify its work, and just 6.6% would entrust it with important decisions; 30% say workplace AI use is left to individual discretion. The official snapshot of vibe coding penetration in 2026.