Ground Truth.
AI, checked against the source.

← All topics

benchmark

Everything on Ground Truth tagged “benchmark” — 5 items.

A giant benchmark tested 24 optimizers - and AdamW's edge held up News

OmniOpt ran a controlled bake-off of more than two dozen modern training optimizers across model sizes from 60M to 1B parameters, and its main lesson is deflating: no challenger cleanly dethrones AdamW, because an optimizer's advantage depends heavily on scale, task, and tuning budget.

Turn around, and the world disappears News

AI video models that are supposed to "understand" a 3D scene only remember what's on screen — pan away and back, and things have reset. Bigger models are worse at it.

Terminal-Bench 3.0 Tool

Continuously versioned agent benchmark with 74 tasks across seven domains, from databases and CUDA to Lean proofs, CAD, and music notation. Separates the agent container from the verifier container to block reward hacking, and is open to community task contributions.

AgentDojo Tool

Independent benchmark for prompt-injection resistance in tool-using agents, used this week as the external check on whether adversarially generated alignment data actually transfers rather than overfitting to its own test set.

ADR Tool

Uber's runtime detector for coding agents, watching what agents actually do on developer machines rather than filtering prompts. Reported 206 credential exposures at 97.2 percent precision across 7,200 hosts, and ships with ADR-Bench, a 300-task benign-versus-malicious evaluation set.