Contributed Article By: Abhi Kumar, Co-Founder ,Voicing AI :
Companies flocked to AI for greater efficiency. A few years into the AI era, they’re realizing how expensive AI-powered efficiency can be.
There are exceptions. According to an August 2026 AXIOS exclusive reported by Yahoo! Finance, Uber increased its AI usage 9.4 times while keeping spending flat. Not “trimmed usage to hold costs down,” but nearly 10 times the usage, with the same bill.
That combination is rare enough to be worth understanding. The following explores why AI spending proved hard to control, how to rein it in, and what a recalibration could mean for the AI industry.
Why are companies seeing costs rise as they use AI?
As AI made its way into the business world, companies bought it like they had been buying software: fixed price, annual budget, done. But AI is different. It bills by the drink.
The change is financially significant. The better your AI-powered product performs, the more it costs you. No company’s tech budget was built to accommodate a line item that punishes you for growth.
AI usage also follows a default engineering habit of sending every request to the largest model on the menu. It’s the low-risk decision for the person making it, and an awful decision for the company paying for it.
Most companies didn’t see the problem during pilots, because pilots hide the truth. At a few hundred calls a month, a 6x difference in model cost is invisible. At 5 million, it’s the entire business. Most teams only ever modeled the first number.
The expense snuck up because very few teams measured cost per transaction on day one. They measured latency and accuracy from the start. Cost per resolved conversation usually showed up around month nine, in an email from the CFO.
And the part people don’t like hearing: if you don’t own any layer of the stack, you don’t really have a cost lever. You can cache, you can trim prompts, you can renegotiate. But your margin is set by someone else’s pricing page.
What does Uber’s accomplishment say about AI’s evolution?
Uber’s result could indicate that AI is getting more economical, or that companies are learning better strategies. It’s both — but they aren’t contributing equally.
Inference has genuinely gotten cheaper. That alone doesn’t produce a flat bill, though, because cheaper AI tends to mean more AI. Every time cost per token drops, someone finds three more places to put a model. Demand expands into the space.
So the deciding factor is discipline. The companies staying flat are the ones who made an actual decision that not every request deserves the same compute.
What makes that decision viable now is that small models have gotten genuinely good — not “acceptable for simple tasks” good, but actually good. An 8B model today handles work that a
70B struggled with a few years ago. The technology made tiering possible. The operators made it happen. Both matter, but the same models were available to every one of Uber’s competitors.
What role does routing play in keeping costs down?
For companies looking to control enterprise AI costs, routing is the highest-leverage factor available. It’s also grossly under-invested in, mainly because there’s nothing exciting about it.
The winning approach routes every request to the smallest model that can handle it properly. Something small and fast reads the request first and decides where it goes. Straightforward retrieval stays on the small tier. Task execution goes mid. Genuine multi-step reasoning escalates. Uber routed its tasks to the model that was best in terms of both cost and intelligence, while smaller tasks were designated for less expensive models.
To optimize the system, don’t route on complexity alone. Route on risk, noting whether an action is reversible or touches regulated data.
Build the escalation path first. A router occasionally sending something to the wrong tier is fine, as long as it gets caught and bumped up. What kills a business is a small model returning a confident wrong answer to a customer with nothing checking it.
The router itself also needs to be cheap and fast, or you’ve defeated the point. If your routing decision burns 200ms and real money, you’ve added a cost center to save costs.
How could more careful AI spending, if widely adopted, impact the industry?
Right now, a big AI spend number reads to a board as commitment. It’s a proxy for ambition. If companies get smarter about spending, that inverts, and a big number starts to look like sloppy engineering.
Smarter spending will also squeeze vendor pricing. The moment a buyer the size of Uber demonstrates publicly that heavy usage doesn’t require heavy spend, every procurement team in the market has something to point at. Buyers will start asking how a vendor routes, what runs where, and what a transaction actually costs. Vendors accustomed to offering a demo and a benchmark chart will have their work cut out for them.
The market will likely split. General platforms for the broad stuff. Purpose-built systems for high-volume, high-consequence work, where owning the stack is what finally lets a company tune the economics.
Zooming out, what Uber’s result points to is just what maturity looks like. Mainframes, cloud, and mobile all went through it. Land grab first, unit economics second. AI is entering the second phase, and that’s a good sign, not a slowdown.
Abhi Kumar is the Co-Founder of Voicing AI, an enterprise voice AI company focused on helping organizations deploy reliable AI solutions in regulated industries. His background spans enterprise software and venture capital, including roles at Oracle, Microsoft, and M12, Microsoft’s venture arm, where he evaluated and invested in enterprise technology companies globally. Through his experience assessing which technologies achieved enterprise adoption, Kumar developed a deep understanding of the challenges organizations face when implementing AI at scale. Today, he focuses on bridging the gap between AI innovation and production reality by building voice AI solutions designed for real-world enterprise environments, including banking, healthcare, aviation, and telecommunications.
Opinions expressed are the author’s own and do not necessarily reflect those of Biz Tech Journals

Leave a Reply