inference-speed
Everything on Ground Truth tagged “inference-speed” — 2 items.
Inception ships Mercury 2.5, a language model that writes in parallel instead of left to right News
Inception released Mercury 2.5, which it calls the largest diffusion language model ever trained, reporting 1,107 tokens per second on standard NVIDIA hardware — roughly an order of magnitude faster than typical autoregressive models — with a 260,000-token context window and no open weights.
Mercury 2.5 Tool
Inception's diffusion language model, which refines a whole draft in parallel rather than writing left to right, reporting 1,107 tokens per second on standard NVIDIA GPUs with a 260K context window. Closed weights, available through Inception's API, Baseten and OpenRouter with 100 million free tokens.