
Microsoft has released ThinkingBox, an open-source evaluation framework for AI agents available through Hugging Face. The tool measures whether agents successfully modify backend databases and terminal states across hundreds of stateful business workflows. It requires agents to demonstrate reliability by performing tasks correctly twenty times in a row rather than relying on single-outcome success.
Verifying backend state and side effects
ThinkingBox evaluates AI agents based on the actual records they leave behind in a system rather than the text they generate or the validity of their tool calls. The framework uses isolated sessions to grade terminal backend states and side effects resulting from agent actions. This approach addresses cases where an agent might produce well-formed tool calls or sound correct to a user while failing to update the underlying database accurately.
The benchmark identifies gaps between an agent's reported actions and the final system state. For example, an agent might appear to resolve a customer issue while leaving a carrier exception open in the database. By focusing on these discrepancies, the tool provides a measure of how models perform when tasked with stateful operations across various business scenarios.
Evaluating reliability through repeated trials
The framework measures consistency by testing agents against hundreds of workflows and requiring successful completion twenty times in a row. This repetition aims to differentiate between accidental success and true reliability. Microsoft's research involved testing various large language models to map the relationship between execution costs and the consistency of the resulting backend updates.
ThinkingBox is now accessible via Hugging Face and utilizes the OpenEnv server and Model Context Protocol servers to run agents in controlled environments. Users can install the tool to score agent episodes and identify failure signatures. The current release focuses on grading agents against a set of five hundred and seven stateful business workflows to determine their dependability for production use.
Original source
This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.
Read the original at Hugging Face