Ground Truth.
AI, checked against the source.

← All topics

api

Everything on Ground Truth tagged “api” — 70 items.

Sakana's Fugu Max sells a model that hands your request to other models, at $2 per million tokens in News

Sakana AI launched Fugu Max on 11 September 2026, a model trained to route each task to a pool of open-weight and specialist models and combine the results, priced at $2 per million input tokens and $6 per million output tokens; the pool is not published and all quality claims are Sakana's own.

OpenAI turns the Codex harness into a product with an Agents API News

OpenAI put an Agents API into public beta on 10 September 2026 that lets developers build on the same managed Codex harness its own products use, with OpenAI handling session orchestration, context compaction and recovery - making the scaffolding around the model, rather than the model, the thing being sold.

OpenAI ships ChatGPT Images 2.5 with a drawing tool News

OpenAI released ChatGPT Images 2.5, cutting image generation latency by up to half against the previous version and adding Sketch, which lets users draw directly in ChatGPT as a reference for a generated image.

DeepSeek gave its cheapest model eyes and did not change the price News

DeepSeek shipped an experimental vision version of its V4-Flash model that accepts images by base64, URL, or file upload, and bills it at exactly the same rate as the text-only model.

Gemini Omni 1.1 Flash can extend a scene instead of restarting it News

Google's updated video model reads up to ten seconds of a clip's prior context before continuing it, up from one second, and adds keyframe control, cheap 360p drafts and 4K upscaling through the Gemini API.

Google's new transcription model edits what you said News

Gemini 3.5 Transcribe removes filler words, silently resolves speakers' self-corrections, and can make function calls out of the transcription layer -- which makes it excellent for voice agents and unusable as a verbatim record.

OpenAI cut Sol's price, and OpenRouter cut it again News

OpenAI dropped GPT-5.6 Sol to $4 per million input tokens and $20 per million output on August 21, 2026, a 33 percent cut on output, and OpenRouter is separately listing the same model from OpenAI at half that.

DeepSeek gave its cheap model eyes, then capped them at 384 tokens an image News

DeepSeek released deepseek-v4-flash-vision-exp, an experimental multimodal version of its cheapest model that accepts images directly in the same agent loop as text, but budgets each image to at most 384 tokens after resizing toward roughly 800 by 800 pixels.

OpenAI put its most intelligent model on Cerebras chips at 750 tokens a second News

OpenAI is previewing Ultrafast, a service tier that runs GPT-5.6 Sol on Cerebras hardware at up to 14 times the speed of standard processing and up to 750 output tokens per second.

DeepSeek starts charging rush-hour prices on August 17 News

DeepSeek is replacing flat API pricing with peak and off-peak rates on August 17, and the steepest change hits cached input on its Pro model, which goes up twelvefold during Beijing business hours.

DeepSeek warns of a significant API price rise, five days after being called 100 times cheaper News

DeepSeek added a footnote to its official pricing page warning that it plans to raise API prices significantly in the near future with no figure and no date attached, five days after an independent benchmark study priced its model at roughly 100 times less per task than Western frontier models.

Qwen3.8-Max Shipped as a Paid API, Not as Open Weights News

Alibaba put Qwen3.8-Max live as a hosted API at $2 per million input tokens and $6 per million output tokens, a fifth cheaper than the model it replaces, while the open weights it promised for Max and a 27B sibling have not shipped.

Ant's Ling-3.0-flash goes live free: 124 billion parameters, 5 billion doing the work News

Ant released Ling-3.0-flash on July 23, a 124-billion-parameter model that activates only about 4% of itself per token, with a 256,000-token context and free access on OpenRouter and Vercel's AI Gateway.

Meta opens its first paid model API with Muse Spark 1.1 News

Meta launched a public preview of the Meta Model API built around Muse Spark 1.1, its first paid, developer-facing model service, with a million-token context window, drop-in OpenAI-compatible access, and zero-shot support for new tools.

fal H3 Max Director Tool

fal's API for continuous AI-generated video streams with live prompts and chunked playback controls, built around MiniMax H3 Max Director.

fal H3 Max Tool

A post-trained MiniMax H3 that renders a five-second 768p clip in under three seconds through fal's API, at 480p or 768p and up to 15 seconds. Hosted only -- there are no downloadable weights -- and priced at $3.60 per minute of generated video. There is a browser sandbox for trying it without writing code.

WeatherNext models on Google Cloud Tool

Google DeepMind's AI weather forecasts, available as a developer API and as raw forecast data in Earth Engine, BigQuery and Vertex AI, with ensemble scenarios out to 15 days.

Vercel AI Gateway (Ling-3.0-flash) Tool

Vercel added Ling-3.0-flash to its AI Gateway with bring-your-own-key support and failover routing, free through August 3. Useful if you want the model behind a single gateway alongside other providers rather than wiring a second API.

Tinker Tool

Thinking Machines Lab's hosted fine-tuning service, now serving Inkling alongside its other models. It is the managed path to customizing Inkling if you do not want to provision the GPUs yourself -- with the caveat that the API caps context at 256K tokens, versus 1M for the open weights you run yourself.

Stack Exchange API Tool

The free, key-less API behind Stack Overflow and its sister sites, which will return exact question, answer and user counts for any date range -- useful for checking claims about the site's decline yourself.

Sakana Fugu Ultra v2 Tool

The higher-priced sibling of Fugu Max, also released 11 September, at $5 per million input tokens and $30 per million output tokens (rising above 272,000 tokens of context). Sakana says Claude Fable 5, Fable 5.1 and GPT-6 Astra are not in its model pool. Available through OpenRouter as sakana/fugu-ultra-v2.

Sakana Fugu Max Tool

Launched 11 September: a model trained to split each request across a pool of other models and combine the answers, sold as one OpenAI-compatible API at $2 per million input tokens and $6 per million output tokens, with a 1 million token context. The pool is not published beyond NVIDIA's Nemotron family, all quality claims are Sakana's own, and it is not yet available in the EU or EEA. Listed on OpenRouter and Vercel AI Gateway.

Sakana Fugu Tool

A single OpenAI-compatible endpoint that dynamically routes each request across several frontier models, so you call one API and get a coordinated multi-model answer.

Sakana AI Fugu Max and Fugu Ultra v2 Tool

Orchestration models that route each task across a pool of other models, available from 11 September through an OpenAI-compatible API. Fugu Max is priced at $2 per million input tokens and $6 per million output tokens; existing Fugu users upgrade with a one-line parameter change.

Qwen3.8-Max-0902 Tool

Alibaba Cloud's 2.4T-parameter MoE flagship with native vision, long-horizon-task support, and a one-million-token context window.

Qwen3.8-Max Tool

Alibaba's new flagship multimodal model, live today as a paid API at $2 per million input tokens and $6 per million output tokens, with a one-million-token context, function calling, structured output, and prompt caching that drops repeated input to $0.25 per million. Weights are promised but not published.

Opentrons Python API Tool

The mature, vendor-supported way to script a liquid-handling robot today: a documented Python and HTTP interface for the Flex and OT-2 platforms, pipettes and modules. Worth knowing as the existing baseline that Anthropic's new hardware standard is being measured against.

OpenRouter discounted models Tool

A live collection of models currently carrying provider discounts on OpenRouter. GPT-5.6 Sol from the OpenAI provider is listed at roughly half OpenAI's own promotional rate, against $5 and $30 for the same model via Azure.

OpenRouter Auto router Tool

Single endpoint that picks a model per request using the past seven days of aggregate platform spend on similar tasks, with a cost_tier parameter to set how much you want to spend and account-level guardrails respected.

OpenRouter Tool

A production gateway to hundreds of models behind one API, with public rankings built from real usage and the ability to sort by price, throughput, latency and popularity.

OpenAI Agents API Tool

Public beta released 10 September. Build agents on OpenAI's managed Codex harness, with OpenAI handling session orchestration, context compaction and recovery; supports durable sessions, streamed progress, custom tools and MCP servers.

Nano Banana 2 Lite Tool

Google's fastest, cheapest Gemini image model - a text-to-image picture in about four seconds for roughly three cents per thousand images, built for high-volume use.

Muse Spark 1.1 Tool

Meta Superintelligence Labs' multimodal reasoning model built for agentic work - tool and computer use, coding, a 1M-token context window, and subagent orchestration; live in the Meta AI app's Thinking mode and on meta.ai, with a Meta Model API in public preview.

MiniMax Music 3.0 Tool

Production music model that takes a creative concept and optional lyrics and composes, arranges, performs and produces a complete song in a single generation, with instrumental-only support. Callable through MiniMax's platform API as model music-3.0, with open weights also published.

Meta Muse Spark Tool

Meta's natively multimodal reasoning model, updated to version 1.3 on September 2, 2026. Reads images, charts, and text together, and offers a Contemplating mode in which multiple agents reason in parallel before answering. Hosted and proprietary at $1.25 per million input tokens and $4.25 output; Meta says an open-weights release is on the roadmap but has not given a date.

Meta Model API (Muse Spark 1.1) Tool

Meta's first paid, hosted model API, built around the Muse Spark 1.1 multimodal reasoning model -- a million-token context window with active context compaction, zero-shot tool and MCP support, and an OpenAI-compatible interface so existing code drops in with little more than an endpoint change.

Mercury 2.5 Tool

Inception's diffusion language model, which refines a whole draft in parallel rather than writing left to right, reporting 1,107 tokens per second on standard NVIDIA GPUs with a 260K context window. Closed weights, available through Inception's API, Baseten and OpenRouter with 100 million free tokens.

Mercury 2 (Inception Labs) Tool

An API-only diffusion language model pitched on raw speed, claiming to out-pace open diffusion models on tokens-per-second for latency-sensitive generation.

LongCat-2.0 Tool

Meituan's 1.6T-parameter MoE model tuned for coding and agentic work, MIT-licensed weights plus a cheap hosted API (launch promo $0.30/$1.20 per million tokens) that self-hosts to avoid data-jurisdiction concerns.

Ling-3.0-flash (free API) Tool

Ant's 124B-parameter mixture-of-experts model that activates only 5.1B parameters per token, with a 256K context and OpenAI- and Anthropic-compatible endpoints. Currently free on OpenRouter as inclusionai/ling-3.0-flash:free; aimed at long-horizon agent workflows and tool calling.

Grok 4.6 Tool

xAI's new frontier model, tuned for long-running agents and available day one in Cursor, Grok Build, and the xAI API. Two dollars per million input tokens and six per million output, with a faster variant at double the price.

Grok 4.5 Tool

SpaceXAI's new 1.5-trillion-parameter model, available in Grok Build, Cursor, and the API at $2 per million input / $6 per million output tokens, with a full public release on July 9.

Gemini Robotics ER 2 Tool

The embodied-reasoning half of Google DeepMind's new robotics family, and the only part available now - it reasons about physical scenes and plans robot tasks via the Gemini API and AI Studio, while the models that actually drive motors stay in private preview.

Gemini Omni Flash Tool

Google's new video model offering developers programmable conversational editing - generate and revise clips up to ten seconds by describing changes in words.

Gemini Omni 1.1 Flash Tool

Google's production-ready generative video model, now able to extend an existing clip using up to ten seconds of prior context, generate between specified first and last frames, draft at 360p for about a third the cost, and upscale finals to 4K. API-only through Google AI Studio.

Gemini Developer API pricing page Tool

Google's own current rate card, including the context-caching rates that determine whether a long-running agent is cheap or ruinous. Worth reading before assuming a headline per-token price describes what you will actually pay, since caching rates and introductory-period expiry dates do most of the work.

Gemini 3.8 Flash in Google AI Studio Tool

Google's newest Flash-tier model, aimed at long-horizon coding and agent work, with a one-million-token context window and 64,000-token output. Free to try in AI Studio; API pricing is $0.75 per million input tokens and $3.75 output through the end of 2026. It deliberately spends more tokens on hard tasks, so budget by cost per finished job rather than per token.

Gemini 3.6 Flash Tool

Google's newly GA fast model streams output nearly twice as fast as 3.5 Flash and costs less per task while holding the same intelligence-index score, tuned for high-volume agent loops that use fewer tokens and tool calls.

Gemini 3.5 Transcribe Tool

Google's new transcription model, shipped as two endpoints: a bidirectional streaming version for live voice agents and a batch version with speaker attribution and word-level timestamps. Handles 85+ languages with mid-stream language switching, cleans filler and self-corrections automatically, and can make function calls to other Gemini models. Try it in AI Studio.

GPT-Image-2.5 Flare and Sunburst Tool

Two new OpenAI API image models: Flare for high-volume generation where speed matters, Sunburst for detailed creative work needing extra precision at the cost of longer generation times.

GPT-6 Astra API Tool

OpenAI's agentic flagship, aimed at computer use, browsing, coding and long multi-step workflows. Five reasoning effort levels from low to max, with no off switch. $10 per million input tokens and $50 per million output, cached input at $1 -- cache discipline is the difference between an affordable agent loop and an unaffordable one.

GLM-5.3 on OpenRouter Tool

Z.ai's GLM-5.3 with a 1 million token context window and always-on reasoning, billed at $1.40 per million input tokens and $4.40 per million output, with cheaper cache reads. Tuned for long-horizon software engineering and vulnerability discovery.

GLM-5.2 on Baseten Tool

The top trending open-weight model served as a fast hosted endpoint, reported at 280+ tokens/sec on Blackwell-class hardware -- an open model you can call like a closed one.

Firecrawl Tool

A hosted API that crawls, scrapes and structures web pages into clean text for agents and retrieval pipelines, handling the JavaScript rendering and rate limiting you would otherwise build yourself.

ElevenLabs Music v2.5 Tool

Default model in ElevenMusic since 11 September, with more layered arrangements and more natural-sounding instruments than v2, which stays available. API access for paid subscribers via model_id music_v2_5; tracks run 3 seconds to 5 minutes and music generation costs 900 credits per minute.

DeepSeek V4-Flash vision API Tool

Experimental image input for DeepSeek's cheap V4-Flash model, billed at the same rate as the text-only version. Accepts images as inline base64, as a URL the model fetches, or as a file uploaded through the Files API. The experimental tag is real, so treat the interface as unstable, but it makes high-volume image reading economically sensible.

DeepSeek V4 Pro (API) Tool

A strong open-weight reasoning and coding model now offered through DeepSeek's own API at a permanently cut, low per-token price, undercutting frontier closed models for high-volume work.

DeepSeek V4 Flash Vision (experimental) Tool

An experimental multimodal version of DeepSeek's cheapest model, live on the DeepSeek API as deepseek-v4-flash-vision-exp. It takes images inline with text via base64, external URL, or the Files API, budgets each image to at most 384 tokens after resizing toward roughly 800 by 800 pixels, and bills at ordinary V4 Flash rates. Good for screenshots, charts, and document layout; not for small type or dense diagrams.

DeepSeek V4 Flash 0731 (API) Tool

The updated V4 Flash checkpoint now serves behind the existing deepseek-v4-flash identifier, with a 1-million-token context, 384K maximum output, tool calls, and an OpenAI-, Anthropic- and Responses-API-compatible interface. Fresh input runs $0.14 per million tokens, output $0.28, and cached input $0.0028 - a fiftyfold discount on repeated prefixes.

DeepSeek V4 Tool

DeepSeek's latest model family (a 1.6T-parameter Pro and a 284B Flash, both with a 1-million-token context by default), available as an API and as open weights on Hugging Face.

DeepSeek Harness Tool

Protocol-aware adapter for DeepSeek V4-Pro and V4-Flash that handles the wire-level quirks a plain OpenAI client drops, including preserving reasoning_content across tool-calling turns and aggregating interleaved parallel tool-call chunks by index. Ships as a Python library, CLI, MCP server, and skill.

Dahl Inference Tool

Third-party inference router reselling top open-weight models (Kimi K2.6, MiniMax M2.7, GLM 5.2) at low per-token prices, currently running a 100M-free-token promotion.

Claude computer use tool Tool

Anthropic's desktop-control toolset reached general availability, giving a model screenshot capture plus mouse and keyboard control for driving real applications. A separate browser-use toolset acts on page structure rather than pixels, which is usually the better choice for web work.

Claude Skills API Tool

Now generally available on the Claude Platform. Skills are folders of instructions, scripts, and templates managed as first-class API objects with create, list, get, delete, and version endpoints. They attach to a request by identifier, execute inside the code-execution sandbox, and up to twenty can ride along on a single call.

Claude Files API Tool

Generally available alongside the Skills API. Upload a file once, get an identifier, and reference it across later requests instead of re-sending contents; download files produced by skills or code execution; list, retrieve, and delete. Files are scoped to the workspace rather than to an end user.

Claude Fable 5.1 Tool

Anthropic's new generally available model for long-running agentic coding and knowledge work, live on the Claude API as claude-fable-5-1 and on AWS, Google Cloud and Azure. Cache reads dropped 75 percent to $0.25 per million tokens; base rates unchanged at $10 in and $50 out. Now permitted to find software vulnerabilities in source code.

Cerebras Inference (Qwen 3.8 27B) Tool

Serves the open Qwen 3.8 27B at roughly 1,500 output tokens per second, with a free tier at 64k context and paid at 128k. Automatic prompt caching cuts time-to-first-token. Read the rate limits first -- the free tier's 90,000 tokens per minute lands almost exactly at the model's own output rate.

Anthropic prompt caching pricing reference Tool

Anthropic's documentation of how cached tokens are billed: cache hits at 10% of standard input, five-minute writes at 1.25 times base input, one-hour writes at 2 times. This is the page that explains why Claude Fable 5.1 can be substantially cheaper per prompt for long agent sessions while costing exactly the same for one-shot calls.

AlphaGenome API client Tool

Google DeepMind's open-source client for querying AlphaGenome variant-effect predictions from Python, for researchers who want programmatic access rather than the web portal.

1F916 Tool

Public discussion forum whose citizens are AI agents, reachable only by JSON API or MCP. Registration issues a secret key, posting is capped at one per day, and the whole ledger is a checkable hash chain. Useful as a working reference design for agent-to-agent coordination.