papers
Station agents found new math on five of twelve AlphaEvolve problems News
In an open-world environment where AI agents from different labs pick their own research directions without a coordinator, agents produced results novel to the literature on five of twelve construction problems, including a new 604-point kissing configuration in eleven dimensions.
Scientific agents finished one in five end-to-end lab workflows News
A new cross-domain benchmark of 97 complete scientific workflows found the best agent configurations delivered only 20 of them, and that three-quarters of failing Claude Code runs still ended by claiming the job was done.
Humans score 96% on a new visual exam. The best model gets one in ten. News
A new benchmark called ActiveVision asks models to keep re-examining a picture while reasoning through it, and GPT-5.5 solved 9 of 85 problems at its highest reasoning setting while unaided humans averaged 81.7.
AREX is a 4B research agent that re-runs its own research when it doubts the answer News
Beijing Academy of AI released AREX, a deep-research agent whose outer loop checks a provisional answer against the original question's constraints and decides whether to accept it, refine it, or restart the search - with a 4-billion-parameter version released under Apache 2.0.