Ground Truth.
AI, checked against the source.

News · 2026-10-03

Wagtail's month on one cheap open model: $68 for half the work, and capacity ran out

A core developer of Wagtail, the open-source content management system, spent September trying to do all his AI-assisted engineering on a single cheap open-weight model, GLM 5.3 Flash. The first half of the month went to plan and cost $68 on that model. Then the inference providers he relied on ran short of capacity, and by month's end half of roughly 2 billion tokens had gone to other models. He calls the challenge a failure that taught him a lot.

Key facts

In early September, Wagtail challenged developers to "ditch the subscription" and spend the whole month on open-weight models. GLM 5.3 Flash, the efficient model from Z.ai that Ground Truth reported had been testing anonymously as "Ox Alpha", was the pick because it sat near the top of Wagtail's own chart of cost versus usefulness. Colas tracked every token with a usage dashboard. The post's subtitle sums up the result: "task failed successfully."

What went right

For two weeks he used nothing else, and the spend stayed "well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions)." He concludes: "For day-to-day developer work, it's totally viable to focus on one or two flash-tier cheap models."

What went wrong

Two things broke the plan. The first was capacity. Wagtail uses smaller inference providers rather than the big labs, and Colas writes that they "do not have the same capacity as the big labs who hoard all the GPUs." GLM 5.3 Flash slowed down at those providers, which he attributes to its popularity, and he switched to similar models, DeepSeek V4.1 Flash and Qwen 3.8 Flash. Switching was easy; needing to was the surprise.

The second was a costly experiment. Building a deliberately "vibe-coded" prototype of a Wagtail server for AI tools, he "chose the 'wrong' model" and spent "450M tokens / $150 / 5kWh of energy use almost overnight." The post does not name that model. The server works and produced a useful demo, but he estimates similar results were possible for about a fifth of the cost with a little more care. It is a bit like taking a taxi across the country because you did not check train times first: you arrive, but the bill tells the story.

Overall, roughly 1 billion of the 2 billion tokens went to models other than GLM 5.3 Flash, and total energy use came to about 35 kWh against a 10 kWh target.

Why it matters

Most claims about cheap open models replacing expensive subscriptions come from benchmark charts. This is a month of real engineering with the bill attached, from an established open-source project. It shows that the per-token price of an efficient model can be low enough to change habits, and that the real risks are elsewhere: shared provider capacity for popular open models, and agent workflows that can burn hundreds of millions of tokens before anyone notices. Colas's plan for October is to measure usage, energy and spend continuously and locally; budget for experiments separately; and split work between orchestrator, scout, implementer and reviewer agents with bounded goals.

The caveat

All figures are Colas's own accounting, not audited bills. The post does not compare the month's cost against what Wagtail would have spent on a commercial subscription, so it does not show savings. The work ran on hosted providers, not local hardware, so it says little about running GLM 5.3 Flash yourself; for that, see GLM 5.3 Flash support in llama.cpp. On Hacker News, commenters asked for task-level comparisons that the post does not provide.


Primary source, verified: read the paper →

Key questions

Did Wagtail run GLM 5.3 Flash on its own hardware?

No. The month used hosted inference providers, and the slowdown came from those providers' limited capacity; local measurement is one of the post's plans for October.

What did the expensive $150 run teach them?

Colas says choosing the wrong model for a vibe-coded prototype burned 450 million tokens, 150 dollars and 5 kWh almost overnight, and similar results were likely possible for about a fifth of the cost.

Is one cheap model enough for everyday coding?

Colas concludes that one or two flash-tier models are viable for most day-to-day developer work, if teams also budget for experimentation and track cost and energy.
Cite this

APA

Ground Truth. (2026, October 3). Wagtail's month on one cheap open model: $68 for half the work, and capacity ran out. Ground Truth. https://groundtruth.day/news/wagtail-spent-a-month-on-glm-5-3-flash-and-half-the-tokens-went-elsewhere.html

BibTeX

@misc{groundtruth:wagtail-spent-a-month-on-glm-5-3-flash-and-half-the-tokens-went-elsewhere,
  title  = {Wagtail's month on one cheap open model: $68 for half the work, and capacity ran out},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/wagtail-spent-a-month-on-glm-5-3-flash-and-half-the-tokens-went-elsewhere.html}
}

Topics: open-weights · coding · agents · inference-cost · energy · practitioner

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.