Skip to content
← All news
3 min read

The US Army Gave Troops Unlimited AI Tokens. They Burned a Year's Budget in Six Weeks.

A $49 million Ask Sage rollout meant to last a year ran out its enterprise pool before summer ended.

The Army lifted its AI token limits in May 2026. The year's enterprise budget was gone within six weeks.

The US Army lifted token limits on its Ask Sage AI platform in May 2026. By mid-June, the year's enterprise allocation was gone, and the Army had to reinstate the limits it had just removed.

What the enterprise pack bought

The Army's Ask Sage contract, a Department of Defense generative AI platform that routes requests to models including OpenAI, Google Gemini, and Meta Llama, included a 100 million token annual enterprise pack as part of a roughly $49 million rollout. Budgeted across the workforce, that works out to about 200,000 tokens per employee per month, a planning number meant to comfortably cover a full year of everyday use across the CIO office's user base. Multi-model routing was part of the pitch: soldiers and civilian staff could pull from whichever provider suited a given task instead of being locked to one vendor's model.

Unlimited lasted six weeks

In May 2026, the Army's CIO office announced unlimited token access on top of that budgeted pool. By mid-June, the year's enterprise allocation was exhausted. An unnamed employee quoted by Wired described the actual mechanism: automatic top-ups kicked in for anyone who hit their initial allocation, which meant the original cap stopped functioning as a real limit the moment unlimited access went live. Limits were reinstated shortly after, and the reversal was reported publicly on July 21 and 22.

The first real stress test of per-token pricing at this scale

A 100 million token annual budget is the kind of number that looks generous on a planning spreadsheet. It didn't survive six weeks of actual unrestricted use. The gap between a per-seat estimate written down in a contract and what people actually consume once the friction of a cap disappears is the whole story here, and it's a gap that shows up at any organization's scale, not just a defense department's budget line. Caps exist precisely because that gap is hard to predict in advance, and removing one without a usage ceiling somewhere else in the system is how a year's budget turns into a six-week budget.

Why a build studio cares

Every product we ship that wires in an LLM has this same planning problem in miniature: a per-seat or per-workspace token estimate that looks comfortable until real usage patterns hit it. We've seen smaller versions of this on client MVPs, a feature that looked lightly used in testing turns out to be the one users hammer once it actually ships to real accounts. Plan token budgets for the spike, not the average, and build the usage alerting to catch the gap before a finance conversation does it for you. An organization the size of the Army finding this out in six weeks is the cheapest version of the lesson it will ever get; a smaller team finds out the same thing from an invoice.

Next step: read the Slashdot summary and TechRadar's coverage. If you're sizing token budgets for an LLM feature and want a sanity check before you commit to a number, write to hello@gattyworks.com.

ai-infrastructuregovernment-techllm-economicsUSArmyAskSageAITokensTokenEconomicsLLMPricingGovTechDefenseTechGenerativeAIPentagonAIInfrastructure

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.