News · 2026-08-08
Walmart put a token allowance on its in-house coding AI
Walmart has replaced unlimited access to its in-house AI coding tool with a fixed token allowance per employee. The tool, built internally and called Code Puppy, went from unmetered to rationed -- and the company's chief technology officer says the reason is not the invoice. It is that employees kept asking the same things over and over, and rebuilding software the company already had.
Key facts
- Walmart moved Code Puppy from unlimited access to a fixed per-employee token allotment.
- Stated rationale: cutting duplicated requests and encouraging reuse, not cost alone.
- Employees retain access to outside tools including Claude and ChatGPT alongside the internal one.
- Primary sources: Business Insider's report and SupplyChainBrain's coverage of Bloomberg's reporting.
The framing matters because it is not the one you would predict. When a large company meters a tool, the assumption is a finance decision -- somebody saw the run rate and pulled the lever. Walmart's stated reason is a management observation: the same question was being asked repeatedly across the organisation, and each asking was a fresh billable generation of an answer that already existed somewhere.
That is a real pathology and a slightly embarrassing one. Ask a model the same question a hundred times and you pay a hundred times, whereas the traditional answer -- an internal wiki page, a shared library, a colleague -- gets written once. Unmetered AI access quietly converts an organisation's reusable knowledge into a per-seat consumable. The cap is a crude but effective way of making the cost of not reusing visible to the person doing it.
Uber's version was more straightforwardly financial. Bloomberg Law reported that the company set a $1,500 monthly cap per employee per agentic coding tool, covering products such as Cursor and Claude Code, after running through its annual AI budget in the first few months of the year. Uber's chief technology officer has since described the company as coming out of its "tokenmaxxing" era, pointing to lower per-token costs achieved through better prompt caching, better defaults, clearer usage visibility, and experiments with open-weight models. Meta has been metering employee AI spend too, and GitHub's move to metered Copilot billing pushed the same economics onto individual developers.
Neither company is retreating from AI, and reading these caps as disillusionment gets the story backwards. Walmart's engineering organisation publicly describes itself as all in on agents across customer, associate, partner and developer workflows. Uber is expanding usage while lowering unit cost. What changed is that both stopped treating inference as free and started treating it as a metered utility with a budget owner -- which is what every other significant infrastructure cost went through, usually about eighteen months after it became load-bearing.
The macro numbers explain the urgency. PwC's 2026 Global CEO Survey found that only 26 percent of chief executives reported lower costs from AI, while 56 percent said they had seen neither revenue nor cost benefits at all. That is not a picture of failed technology -- adoption is obviously widespread -- it is a picture of spending that has not yet resolved into a measurable return. When a majority of leaders cannot point to a benefit, the natural next move is to bound the input while the output is figured out. PwC's follow-up work argues that the organisations seeing results are disproportionately the ones with data quality, data management and governance in place first.
The analogy is electricity in a factory, and it is more exact than it sounds. Early industrial adopters ran motors continuously because the plant was wired and nobody was counting. Sub-metering came later, and it did not reduce electrification. It made the difference between a machine that was producing and a machine that was merely running visible for the first time. A token cap does the same thing for a coding agent: it does not stop anyone from using the tool, it makes the hundredth identical query show up on someone's ledger.
The honest caveat: nobody outside these companies can see whether the caps are working. Walmart has not published how much duplication actually fell, and Uber has not disclosed the size of the original overshoot. It is also possible that the metric being optimised is the wrong one -- a developer who stops asking a second clarifying question because they are watching their allowance is cheaper and possibly worse, which is precisely the trade-off recent work on whether a coding agent knows when to stop asking has been trying to measure. Caps make cost legible. They do not make value legible, and the second problem is the harder one.
Key questions
Which tool did Walmart cap?
Was the cap about cost?
Are other large companies doing the same thing?
Cite this
APA
Ground Truth. (2026, August 8). Walmart put a token allowance on its in-house coding AI. Ground Truth. https://groundtruth.day/news/walmart-put-a-token-allowance-on-its-in-house-coding-ai.html
BibTeX
@misc{groundtruth:walmart-put-a-token-allowance-on-its-in-house-coding-ai,
title = {Walmart put a token allowance on its in-house coding AI},
author = {{Ground Truth}},
year = {2026},
month = {aug},
url = {https://groundtruth.day/news/walmart-put-a-token-allowance-on-its-in-house-coding-ai.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.