fast-inference
Everything on Ground Truth tagged “fast-inference” — 1 item.
Cerebras Inference (Qwen 3.8 27B) Tool
Serves the open Qwen 3.8 27B at roughly 1,500 output tokens per second, with a free tier at 64k context and paid at 128k. Automatic prompt caching cuts time-to-first-token. Read the rate limits first -- the free tier's 90,000 tokens per minute lands almost exactly at the model's own output rate.