Ground Truth.
AI, checked against the source.

← All topics

terminal-bench

Everything on Ground Truth tagged “terminal-bench” — 2 items.

A runbook, not a model, hit 95 percent on a live agent benchmark for 15 dollars News

StateM reaches 95.3 percent raw accuracy on Terminal-Bench 2.1 across 445 trials without changing any model weights, using a state-machine runtime and a reusable runbook, at about 15 dollars of final-score API spend against 574.68 dollars for the reference run.

A new terminal benchmark drops the best agent from 84 percent to 34 News

Terminal-Bench 3.0 launched with 74 tasks across seven domains, and the top agent scores 34.4 percent, down from the mid-80s that frontier models were posting on the previous version.