Microsoft's ThinkingBox grades AI agents on the database they leave, 20 runs each
Microsoft has published ThinkingBox, a benchmark for AI agents that ignores what the agent says it did and checks what is actually in the backend afterwards. A joint post with Hugging Face, published on 3 October, describes 507 stateful business workflows, each run 20 times from an identical clean backend, and makes the benchmark available through Hugging Face's OpenEnv.
The example the authors lead with. A customer's $745 appliance is stuck in a courier exception fifteen days late. The agent makes nine sensible tool calls, opens a ticket, reads the refund policy correctly, then closes the ticket as solved. The required end state was on hold, because the carrier exception was still open. A grader looking at tool calls sees nine valid calls; the only thing that disagrees is one field in the database.
How common that is. In an ablation over 121,680 valid trials on 12 models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still finished cleanly, called a state-changing tool and reported no final error. The checks found wrong field values in 77.61% of them, unintended extra effects in 43.30% and missing required effects in 25.36%; one failure can have several.
Once versus every time. The post reports three numbers per model: the single-attempt score, tasks solved at least once in 20 runs, and tasks solved in all 20. Claude Opus 5.5 leads the single-attempt score at 67.16%, two thirds of a point ahead of Claude Opus 5. Kimi-K3, the strongest open-weights model, solves 476 of 507 tasks at least once but only 68 in all 20 attempts. Opus 5 solves fewer tasks at least once, yet completes 241 every time, and Opus 5.5 also completes exactly 241, so the newer model's higher headline bought no extra dependability. GPT-6 Astra keeps 78% of its single-attempt rate across repeats; GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro keep about 8%.
What it costs. Priced at list rates, GPT-5.6 Sol is cheapest per successful attempt at $0.127. Per dependable task, meaning the full 20-run campaign divided by tasks passed 20 out of 20, GPT-5.4 is cheapest at $6.80 for 128 tasks, GPT-6 Astra reaches 231 at $7.45 and Opus 5.5 reaches 241 at $7.80. The authors also classify failures and say roughly four in five are tool handling rather than reasoning.

What it means
For anyone putting an agent in front of real records, the useful column is the one most leaderboards leave out: how often the same task comes out right on every run. Tests that check the agent's transcript, or only that it wrote something, would have passed the failed ticket above. The check that caught it compared the final state with the one the task required.