Ramp launched its own AI model router and called it Router
One API endpoint, eight model providers, and three years of internal use before Ramp opened Router to every customer.
Ramp built its own AI model router, named it Router, and claims it cuts AI bills up to 92 percent.
Ramp announced Router on August 20, TechCrunch reported: a proprietary AI model router that sends every request to whichever of eight providers fits the job on cost, latency, or performance. The corporate card and expense fintech says it built and ran the system internally for three years before turning it into a product. For now, it is available in the US only.
What Router does
One API endpoint replaces direct calls to OpenAI, Anthropic, DeepSeek, Moonshot, Minimax, Nvidia, xAI, and Z.ai. A customer optimizes for cost, latency, or performance, and Router picks the model per request instead of per integration. A dashboard shows spend, latency, and how often a request fell back to a second-choice model.
- Flex tier switching: move a request to a cheaper pricing tier when it still clears the quality bar.
- Shadow testing: run a candidate model against live traffic before trusting it with real requests.
- Benchmark based routing: score models against Ramp's own benchmarks instead of provider marketing claims.
- Nvidia Switchyard mode: send only the hardest requests to pricier models and keep the rest on cheaper ones.
From tracking spend to controlling it
Ramp already sells AI token spend tracking to its card customers: visibility into what a company's AI bill looks like. Router is the next step, from watching the number to changing it. Ramp CTO Rahul Sengottuvelu is the company's spokesperson on the launch, and Ramp says it built and ran Router on its own AI spend for three years before turning it into a product anyone can buy.
The company can afford to build infrastructure like this: Ramp raised $750 million in June at a $44 billion valuation. Router itself launches free through the end of 2026, plus a $26 signup credit, though customers still pay the underlying providers for inference. It retains request inputs, outputs, and tool calls for a year by default, with personal data scrubbed before any of it feeds product improvement.
The numbers Ramp is reporting
Ramp's own numbers claim cost cuts of 40 percent, 30 percent, and in one case 92 percent, depending on the customer and the routing mode. Ramp's launch announcement leads with these figures, and they deserve the same skepticism as any vendor's own benchmark: Ramp picked the customers, picked the mode, and published the percentage. Nobody outside Ramp has audited them yet.
Why a build studio cares
We build AI workflows on top of more than one model provider, and every client conversation about that turns into a conversation about cost. Route the easy classification calls to a cheap model, keep the expensive one for anything that needs real reasoning, and log which choice actually got made. Router is that pattern, productized and sold back to the companies who would otherwise build it themselves.
The part worth checking before recommending a vendor router to a client is the same part we check in every audit: what happens to the data in transit. A year of retained inputs, outputs, and tool calls is a year of a client's prompts sitting on someone else's infrastructure, PII scrubbing or not. That is a data flow question, not a pricing question, and it belongs in the vendor evaluation before the savings percentage does.
Next step: if a client's AI workflow already spans more than one provider, write down where each request's inputs and outputs are retained today, and for how long, before adding a router that changes the answer. Questions about routing or cost control across model providers? Write to us at hello@gattyworks.com.