AI agents are becoming one of the biggest developments in artificial intelligence, but there is a problem that could become increasingly important as these systems become more capable: how much will it actually cost to run them? 

 

A new forecast from Gartner says the cost of AI inference per agentic workflow could increase more than fivefold through 2028, creating a surprising situation in which better AI could become significantly more expensive to operate even while the cost of individual AI operations continues to fall.

 

The reason is that future AI agents are expected to do much more than answer a single question. A traditional chatbot might receive a prompt, generate an answer and stop. An AI agent can potentially break a goal into multiple steps, search for information, use software tools, analyze the results, make decisions and continue working until the task is completed. Every additional step can require more computing, which means a single request can consume considerably more AI resources than an ordinary chatbot conversation.

 

This creates what Gartner describes as an inference paradox. AI companies are becoming more efficient at running individual model operations, but businesses are simultaneously asking those models to perform longer and more complicated workflows. The result is that the cost of completing an entire AI-powered task can rise even when the cost of each individual AI operation falls.

 

Imagine asking an AI agent to research ten competitors and write a report. A basic AI assistant might generate a quick response from the information already available to it. A more advanced agent could search websites, collect information, compare pricing, analyze products, create tables, check its findings, rewrite the report and perhaps send the finished document to a manager. The second system is far more useful, but it is also performing many more operations to complete the same general request.

 

That difference could become extremely important for businesses. Companies are increasingly interested in AI agents because they promise to automate work that previously required employees to move between different applications and complete repetitive steps manually. An agent could potentially handle customer support, research, sales administration, software testing, data analysis and other workflows without requiring a person to supervise every individual action.

 

But automation is only attractive if the economics make sense. A company might be willing to spend a few cents or a small amount of money for an AI to complete a simple task, but it may think twice if completing thousands of complicated workflows every day becomes expensive. Businesses will therefore need to measure not just how intelligent an AI agent is, but how much it costs to accomplish a specific business outcome.

 

This could change the way companies compare AI models. Today, much of the attention is focused on benchmark scores, reasoning ability and whether one model is smarter than another. In an agent-driven economy, another measurement could become just as important: how much does it cost for the AI to finish the job?

 

A model that is slightly less capable but completes a task using significantly fewer computing resources could sometimes be more valuable to a business than the most powerful model available. Companies may eventually use different AI models for different stages of a workflow, reserving the most advanced systems for difficult decisions while using cheaper models for routine operations.

 

This could also create a new layer of competition among AI companies. OpenAI, Anthropic, Google, Meta and other AI developers are competing to build increasingly capable systems, but the companies that can make those systems affordable at scale could gain an equally important advantage. Intelligence matters, but businesses ultimately have to pay for the computing required to produce that intelligence.

 

The situation becomes even more interesting when AI agents begin operating continuously. Instead of asking an AI for help once or twice a day, a business could have agents monitoring customer requests, checking inventory, analyzing transactions, watching for security problems and preparing reports around the clock. Such systems could generate enormous value, but they could also generate a large number of inference operations.

 

For consumers, the effect could appear in AI subscription prices and usage limits. Today's AI services often offer a fixed monthly subscription with limits that users may not always notice. If future AI assistants perform much more complicated tasks in the background, companies may eventually need to introduce different pricing structures based on how much work an agent performs rather than simply how many messages a user sends.

 

That could produce a new distinction between AI chat and AI work. Asking an AI a question might remain relatively cheap because the system only needs to generate one response. Asking an agent to research a subject, contact several services, compare information, create files and continue monitoring the situation could be treated more like purchasing a digital service.

 

The economics could also influence how much autonomy companies give their AI agents. A business might allow an agent to handle routine tasks automatically while requiring human approval for expensive or complicated operations. This would not only reduce safety risks but could also prevent AI systems from consuming unnecessary computing resources.

 

There is another possibility: AI agents could become cheaper over time even as their capabilities increase. Improvements in chips, model architecture, software optimization and specialized AI hardware could reduce the cost of individual operations dramatically. 

 

Gartner's forecast does not mean that every AI task will simply become five times more expensive. The important point is that increasingly complicated agentic workflows could require much more computation, potentially overwhelming the savings created by more efficient inference.

 

This is why the future of AI may not simply be a race toward smarter models. It could become a race toward smarter models that know when not to think too much. An efficient AI agent should be able to determine which tasks require advanced reasoning and which can be completed using simpler and cheaper methods.

 

For example, an AI agent handling a customer's address change should not necessarily need the same level of computational power as an agent analyzing a complicated legal document or designing a software system. Future AI platforms could automatically route different parts of a workflow to different models based on difficulty, cost and importance.

 

That approach could make agentic AI much more practical. Instead of running the most powerful model for every single action, an AI system could use a combination of models working together. A small model could handle simple classification, another could search and summarize information, while a more advanced reasoning model could make the final decision when necessary.

 

The development could also create opportunities for companies building AI infrastructure. If millions of businesses begin deploying agents, they will need systems capable of monitoring usage, controlling costs, managing model selection and determining how much computing each workflow should receive. AI infrastructure could therefore become just as important to the agent economy as the underlying models.

 

For developers, this means learning how to build AI agents may eventually involve much more than connecting an application to an AI API. Developers will need to think about token usage, inference costs, model selection, memory, tool calls, workflow design and how to prevent an agent from repeatedly performing unnecessary operations.

 

The biggest surprise may be that the smartest AI is not necessarily the most commercially successful AI. If one company produces a model that is slightly better but dramatically more expensive to operate, another company could potentially win customers by offering a model that is almost as capable while completing the same work at a lower cost.

 

This could become especially important as businesses move from experimenting with AI to depending on it for everyday operations. During the early stages of AI adoption, companies can tolerate expensive experiments because they are trying to understand what the technology can do. Once AI agents become part of critical business processes, however, companies will need predictable costs and measurable returns.

 

The future AI market could therefore develop around a simple question: how much does it cost to get something useful done? That may ultimately matter more to businesses than how many parameters a model has or how impressive its benchmark scores look.

 

AI agents are still developing rapidly, and forecasts about their future costs can change as hardware and software improve. But the underlying issue is unlikely to disappear. The more work we ask AI systems to perform autonomously, the more computing they are likely to require.

 

That means the next stage of the AI race may be about more than creating systems that can think better. It will also be about creating systems that can think efficiently, use the right model for the right task and complete complicated work without wasting enormous amounts of computing power.

 

If AI agents become the digital workers many companies expect them to become, their price will matter just as much as their intelligence. And as these systems move from answering questions to performing entire workflows, understanding the cost of AI could become one of the most important parts of understanding the future of artificial intelligence.