terminal-bench
Everything on Ground Truth tagged “terminal-bench” — 2 items.
A runbook, not a model, hit 95 percent on a live agent benchmark for 15 dollars News
StateM reaches 95.3 percent raw accuracy on Terminal-Bench 2.1 across 445 trials without changing any model weights, using a state-machine runtime and a reusable runbook, at about 15 dollars of final-score API spend against 574.68 dollars for the reference run.
A new terminal benchmark drops the best agent from 84 percent to 34 News
Terminal-Bench 3.0 launched with 74 tasks across seven domains, and the top agent scores 34.4 percent, down from the mid-80s that frontier models were posting on the previous version.