free-tier
How to run AI on free APIs (with 9router) Lesson
The big AI labs and a few hardware makers give away real model access for free, if you can juggle the rate limits. Here is how we run this whole site's research on zero dollars, and how you can too.
Retriever Free Mode Tool
A public zero-credit mode for everyday AI and cloud-browser tasks, with fair-use limits and a clearly labeled sponsored card beside results.
Ox Alpha on OpenRouter Tool
A free anonymous stealth model with a 1,048,576-token context window, up to 131,072 output tokens, and text, image, and video input. Genuinely usable and genuinely free right now, with a caveat worth reading first: OpenRouter states it is not the developer or provider, its model page says prompts and completions are retained by the anonymous provider, and other documentation for the same model claims zero retention.
OpenCode Zen (Ox Alpha free tier) Tool
OpenCode's Zen gateway serves 'Ox Alpha,' a free, unlimited, reasoning-mandatory coding model with a load-tested one-million-token context window and no authentication required. Measured at about one second to first token and 35-46 tokens per second. Independent forensics attribute it to the GLM family; the operator has not identified itself, so treat everything you send it as disclosed to an unknown party.
Google Antigravity Tool
Google's agentic coding environment, free during its public preview, built to run Gemini models as autonomous agents across an editor, a terminal and a browser rather than as an inline autocomplete.
Cloudflare AI crawler controls Tool
Free-tier controls that split AI crawler traffic into Search, Agent and Training, each set independently to allow, block site-wide, or block only on ad-bearing pages. These are edge blocks on classified traffic, not robots.txt requests. New domains change default on September 15, 2026.
Cerebras Inference (Qwen 3.8 27B) Tool
Serves the open Qwen 3.8 27B at roughly 1,500 output tokens per second, with a free tier at 64k context and paid at 128k. Automatic prompt caching cuts time-to-first-token. Read the rate limits first -- the free tier's 90,000 tokens per minute lands almost exactly at the model's own output rate.