Ground Truth.
AI, checked against the source.

← All topics

inference-speed

Everything on Ground Truth tagged “inference-speed” — 2 items.

Inception ships Mercury 2.5, a language model that writes in parallel instead of left to right News

Inception released Mercury 2.5, which it calls the largest diffusion language model ever trained, reporting 1,107 tokens per second on standard NVIDIA hardware — roughly an order of magnitude faster than typical autoregressive models — with a 260,000-token context window and no open weights.

Mercury 2.5 Tool

Inception's diffusion language model, which refines a whole draft in parallel rather than writing left to right, reporting 1,107 tokens per second on standard NVIDIA GPUs with a 260K context window. Closed weights, available through Inception's API, Baseten and OpenRouter with 100 million free tokens.