Visualização de leitura

What is sovereign AI? Strategic control of your AI future

Ask IT leaders what sovereign AI is, and you’ll get a wide range of answers. Some will even struggle to define the term.

Sovereign AI is an emerging concept focusing on giving organizations — or countries —control over how they develop, deploy, and govern the technology, often using in-house talent, data, and infrastructure.

But only 13% of respondents in a survey from AI platform provider Cohere and IDC say sovereign AI is widely understood across their organizations, and one in three IT leaders had difficulty describing sovereign AI in their own words.

It’s important for IT leaders to understand the concept, because it can help them control costs, keep internal data private, and avoid vendor lock-in, advocates say.

A solid sovereign AI plan can help organizations avoid disruptions caused by forces outside their control, says Joelle Pineau, chief AI officer at Cohere, which offers an AI platform that enables customers to host AI models on premises.

“Over the past year, enterprises and governments have confronted a hard truth: AI systems that rely on external infrastructure can be disrupted without warning by decisions and actions outside their control,” Cohere says in a recent report. “Recent model access restrictions and several high-profile cybersecurity incidents have become a global wake-up call, exposing how fragile technological dependencies can be.”

Sovereign AI is about giving organizations as much autonomy, choice, and control as possible as they deploy and run AI systems, Pineau says.

“The notion of sovereignty really is about giving users control over their tech stack, the ability to choose how it’s deployed, how it works, what data is fed into the system, and how employees are exposed to the technology,” she adds.

Pineau wasn’t particularly surprised about the lack of understanding about sovereign AI reflected in the survey. Cohere’s accompanying report is an attempt to bring more clarity to the issue, she says.

Many goals under one umbrella

Confusion about sovereign AI in part reflects practitioners’ varying goals. Some users want to maximize their AI model options, some want better control over data ingested into AI systems, some want data to reside within country borders, some want to control costs, and others may want to run AI models optimized to their native language or culture.

For Berk Yilmaz, co-founder and CTO at AI integrated development environment provider Noah Labs, sovereign AI encompasses five characteristics: data sovereignty, legal jurisdiction, model provenance, operational control, and supply chain independence.

“Fulfilling one of those does not mean fulfilling all the others, so two executives can agree with sovereign AI and have little in common,” he says.

Freedom of choice doesn’t always mean a company has to host an AI model on premises or data must reside within a certain country, advocates suggest. Sovereign AI is more about preserving options when something unexpected happens.

“The sovereignty model performs well even in a situation where the vendor breaks off the contract, your model is added to the list of models that are banned for exports, and the connection is off,” Yilmaz says. “Each of these three scenarios has already played out somewhere in the last year.”

Others have different definitions. Confusion over sovereign AI isn’t surprising because it is four separate concepts that were collapsed into one, says Jeet Pattanaik, founder and CTO of AI solutions provider Glokal AI. Those four concepts: where a company’s data physically sits, what country’s law can compel access to it, who controls the AI model, and whether a company could still operate if the relationship with the model provider ends.

“Vendors usually sell you the first and call it sovereignty, because data residency is easy to demonstrate and makes a good diagram,” says Pattanaik, author of the book Sovereign AI: The Enterprise Guide to AI That Is Private by Design, Compliant by Default, and Yours Forever. “The hard one is the second, and it’s a legal question rather than a technical one. A server in Frankfurt owned by a US company is still reachable under US law.”

While the concept is largely about control, few companies want full control of their AI stack, he notes.

“Building your own models is expensive and usually worse,” Pattanaik says. “What CIOs actually want is bounded dependency: knowing exactly what you depend on, what happens if it changes, and having an exit that doesn’t take three years.”

Future impact

But the benefits aren’t always immediate, Pattanaik notes.

“The value arrives in specific moments, not continuously — when a regulator asks who processes this data and under whose jurisdiction, or when a vendor changes terms at renewal,” he says. “Organizations that thought about sovereignty already have an answer. Everyone else discovers the question and the crisis at the same time.”

Still, Pattanaik sees momentum building for the concept, with regulated industries such as banking, healthcare, and the public sector paving the way, treating it as a requirement.

“Most others are still at the stage of asking during vendor selection and accepting whatever answer comes back,” he says. “From where I sit it’s moved from philosophy to a procurement line item over roughly the last 18 months, but unevenly.”

David Wang, COO at enterprise AI gateway provider Tetrate, sees similar adoption trends with regulated industries and defense contractors leading the charge.

Other organizations should focus on a handful of questions to decide whether to investigate sovereign AI, he recommends. Companies that can most benefit include those with more than one regulator or legal entity, including recent acquisitions; those with a huge developer population running coding agents; and those with one AI model vendor that commands more than half of their AI spending.

“At that size, a supplier price change becomes a budget event,” Wang says.

Like Pattanaik, Wang suggests that sovereign AI is part of a long-game strategy rather than immediate gains. A good plan enables organizations to quickly shift to open AI models when a frontier model gets too expensive, he says.

“This work mostly prevents a loss rather than creating a gain, which is why it rarely wins the budget on its own,” he adds.

Cohere’s Pineau sees benefits for a broad range of organizations. With token costs a major concern for many companies, a sovereign AI plan can explore alternatives to current AI vendors, she notes.

“A lot of companies care about it, but they care about different aspects,” she says. “In regulated sectors, they care about the compliance aspects, and in some sectors with tight profit margins, they care a lot about the cost control. The manufacturing, the telecoms, and the energy sectors care about the ability to control their costs.”

IT infrastructure shortages are real and lasting. Here’s how to cope

Lead times of nine to 12 or even 18 months. Costs rising by 35%, 45%, even 50% to 200%. More than halfway through 2026, the market for IT infrastructure that’s crucial for enterprise projects, including those involving artificial intelligence, is strapped.

Memory is at the root of the shortages. Memory prices “have risen by 50% to 200%, resulting in PC prices increasing by 35% to 45% and some server prices rising over 125%,” according to Jon Forest, VP analyst at Gartner. Network switches also need memory, albeit in lesser amounts than servers, so they are not immune, with prices and lead times likewise rising dramatically.

Industry experts agree that most of the issues stem from hyperscalers gobbling up memory capacity, which trickles down to servers, storage systems, and networking devices. But while the source of the problem may be new, supply chain disruptions are far from unprecedented.

As a result, industry insiders are not short on advice on how best to deal with the situation, with tips including making better use of what you have, considering options beyond your usual scope, and lots of planning with your vendors and internal finance teams.

State of the problem

Just how bad is the current supply chain problem? “It’s pretty bad,” says Matt Kimball, vice president and principal analyst with Moor Insights & Strategy. Companies accustomed to 30- to 45-day lead times for various infrastructure are now looking at 6, 12, or even 18 months.

“It’s real, and I’m hearing it from companies of all sizes, from the 1000-server to the 10,000-server shops,” Kimball says.

“Memory costs are expected to rise sharply well into 2027 and will reach up to 25% of network hardware expenses by the end of 2027,” according to an email Gartner’s Forest sent to Network World. The figure below shows the timeline Gartner expects for memory prices, and Forest notes that the same timing applies across networking, storage, and compute infrastructure. 

Gartner NAND DRAM stats

Gartner

“Enterprise network equipment pricing is projected to increase by over 20% in 2026. This upward trend is anticipated to continue with a further rise of 3% to 5% entering 2027, with no signs of price reduction until the end of 2027.”

But “reduction” will likely look more like “stabilization.”

“That’s something a lot of people don’t like to talk about. But let’s say prices went up 40%, they may come down five,” says Phillip Privett, senior vice president of vendor management with the global distributor and value-added reseller TD SYNNEX. “They’re not going to come down 40%.”

Perhaps worse, compared with past disruptions caused by issues such as fires in chip fabrication factories or the Covid pandemic, Kimball says this one is “durable” because its cause—the AI wave—is more long-lasting and just getting started.

“This AI inference wave we’re hitting is just beginning. It’s going to be longer and bigger than the training wave,” he says. “It’s impacting everything, from AI infrastructure to the traditional stuff that’s standing up your virtualization and cloud infrastructure.”

No vendors seem to be immune, not even the likes of Cisco, which makes its own Cisco Silicon One chips. Or, at least, it designs the chips; they’re actually manufactured by the Taiwan Semiconductor Manufacturing Company (TSMC), the same company that makes many of the other chips that are in such demand. And that’s only one component of many that comprise a switch.

On the other hand, the margins Cisco gets from enterprise sales are far greater than those from hyperscalers because Cisco sells mainly just hardware to hyperscalers, whereas enterprise sales generally include software and services as well. So, Cisco has incentive to keep enterprise customers happy and maintain the 66% margins it reported in Q3, its latest quarter.

Still, Cisco must deal with the same shortages as other vendors.

“I wouldn’t say any company is faring better than others,” says Neil Anderson, vice president and CTO for cloud, infrastructure, and AI solutions at World Wide Technology (WWT). “There may be nuances that some suppliers are employing to balance it to some extent, but I fail to recognize a supplier that’s not having almost the same issue.”

Cloud storage vendor Backblaze is one company that’s facing equipment cost and availability issues. “There are different types of shortages occurring in multiple places, all driven by unusual market demands, really by just a handful of very large buyers,” says James Rowell, senior vice president of operations with Backblaze.

Backblaze is constantly forecasting and monitoring demand triggers, Rowell says. That involves close alignment with the sales team to forecast client needs, as well as paying attention to historical trendlines to predict upcoming demand from new deals and growth with existing clients. But the company also looks for “unnatural market-related triggers” that would cause a spike in utilization.

With hyperscalers buying up vast amounts of capacity, “This is definitely an unnatural phase,” Rowell says. “For about for the last 12 months, I would say there’s been somewhere between a 15% and 30% uptick in costs,” especially in terms of servers and compute disks.

On the positive side, at least for Backblaze, the company is also seeing an uptick in business from an interesting source: AI companies. “We reported in the last earnings period a 70% increase in AI companies using our platform,” says Patrick Thomas, vice president of marketing at Backblaze. “That’s massive.”

On top of that, the company is seeing an uptick in deals from enterprises that can’t get the storage capacity they need or want on-prem. “There’s a general market nervousness where we’ve got potential deals coming our way because those organizations are concerned about being able to do it themselves,” Rowell says.

While some expect new chip fabrication plants currently under construction will ease memory supply constraints, Privett doesn’t buy it. “I don’t see it getting better anytime soon,” he says. “Building a new fab is a two-year process.”

Advice: Start with the basics

Enterprises, then, must play the cards they’re dealt. For Moore Insights’ Kimball, who did stints as an IT exec with the states of Florida and Oregon, that starts with making the most of what you have.

Such a strategy is “shockingly not implemented much” across the companies he sees. “A simple capacity planning exercise can free up a lot of resources.” That includes virtualized servers running at just 20% to 30% utilization as well as extending the life of existing servers. While 15 or 20 years ago it was common to refresh every four years or so, companies can often get six or seven years out of today’s servers.

While such strategies won’t solve your AI compute challenges, they can certainly help support your ongoing operations and free up budget for AI and other modernization projects, he says.

“Sweat your assets,” agrees Privett of TD SYNNEX. “Work them as much as you can, add only what you need, get extensions on your licensing, renewals on your services agreements and things like that. Just sweat it out a little longer.”

If you have budget to spend but can’t get the hardware you’re after, buy something else, says WWT’s Anderson. “Look at things that are not tied to those components, like software projects or SaaS licensing,” he says.

Get friendly with finance teams

Numerous experts recommend regular meetings with your CFO or finance teams to keep them apprised of what you’re up against so the company can plan accordingly.

Gartner’s Forest advises using rolling 12- to 24‑month forecasts and engaging early with suppliers to identify constrained components and SKUs. Committing to quarterly or monthly buys can help you avoid long-term agreements that extend past the rapid increases we’re seeing in 2026, he says.

Also engage with the financing arm of your equipment vendors, some of which are offering financing incentives, Privett says. Compute vendors in particular are offering subsidized financing, deferred payments, and low-cost financing for the first year or so. “Those are huge opportunities to take advantage of,” he says.

By engaging with finance teams, IT groups can conduct budget allocation exercises and try to come up with ways to make the financials work. The last thing you want to do is surprise them with additional budget requests out of the blue.

Kimball recalls his days with the state of Florida, when all budget requests were examined by a technical review working group—which was designed to be hostile.

“I can’t imagine going to them and saying, ‘Oh, did I say that was a million dollars? It’s actually $2 million. I need you to write me a bigger check,’” he says. “I would walk into one of the swamps in Tallahassee and get eaten by the alligators instead of doing that.”

Work with your vendors and VARs

As you put plans together, lean on your vendors for help, including channel partners such as value-added resellers (VAR) and national resellers. “Work with them to map things out and understand what your workloads will look like,” Kimball says.

That’s what Backblaze’s Rowell regularly does with his suppliers. He lays out his forecast for the year, with commitments on what Backblaze will definitely buy, as well as scenarios that account for rapid growth, say, 2x. “And they’ll come back with, ‘Well, okay, no problem,’ or maybe they say we need to put in an allocation right away, or we won’t be able to get what we may need,” he says.

Similarly, he sits down with his CFO regularly to map out predictive models that factor in inflation, price hikes, and the like. The idea is to plan out multiple scenarios, so you don’t get blindsided.

“If you don’t do that, you’ll get caught with your pants down, on the upside-down end of spectrum,” he said – meaning not having the capacity to take advantage of market opportunities.

Acquiring the capacity you need to meet project demand may also mean being flexible in terms of your equipment choices. If you’re a Dell shop but can’t get Dell servers, maybe you go with Lenovo, Kimball says.

“You’ve got to figure out how to use all this silicon and infrastructure in a heterogenous way to serve your needs,” he says. That’s especially true when it comes to AI infrastructure. “If you think you’re going to go with 100% Nvidia for everything from RAG [retrieval augmented generation] to inferencing at the edge, you’re kind of crazy, not because of cost but because of availability.”

Look at alternatives, including AMD and cloud solutions, while staying mindful of how it all plays together. You may not be able to get Nvidia GPUs, but AWS, Azure, and Oracle Cloud have them, Kimball notes.

Be strategic, perhaps by using cloud offerings to handle certain tuning or inference workloads, then bringing them back in-house when appropriate. “Have a better understanding of what absolutely has to be on prem and what can be in the cloud,” he says.

That’s good advice, says Backblaze’s Thomas. When it comes to AI, think about performance tiers and the range of use cases you have. They don’t all need top-tier performance.

“People get wrapped around axle of needing the top end. There’s a lot of flexibility in the edges, innovation in different hardware and software,” Thomas says.

Gartner likewise advises companies to increase configuration flexibility and expand sourcing paths. That may include buying from secondary markets and lease-return programs to preserve continuity with existing infrastructure until the shortages pass, Forest says.

Get started somewhere

Even if you can’t acquire or have to wait for the infrastructure you need, don’t let that keep you from getting started with AI or other modernization projects.

Options include public cloud and neocloud providers, Anderson says. WWT also provides capacity in its own lab so customers can get started with proof-of-concept projects. “Don’t just throw your hands up. We can help you find access to capacity,” Anderson says. “Production-scale AI may be delayed, but don’t let that derail your strategy.”

Colocation providers may likewise be an option, especially if enterprises are struggling to acquire high-end networking equipment. Networking is a key value proposition for colocation providers, in that they have built-in connections to various cloud providers and other ecosystem players.

Equinix, for example, has 280 data centers in 77 metropolitan areas, says Phil Read, senior director, colocation product management for the company. If you have the compute infrastructure, Equinix can help you with the high-end connectivity required both intra- data center and at edge facilities.

It also has partnerships with the likes of Cisco and Nvidia for “ready-to-go AI connectivity,” Read says. That means Equinix offers the right infrastructure to meet the requirements of high-end compute solutions in terms of power density and cooling. Such power densities are significant, requiring 120k VA per rack and up. “There’s plenty of talk about a megawatt rack,” he says.

Power is a significant issue in this entire discussion, Privett says. Older installed computing infrastructure likely consumes far more power than newer systems, which is an argument for upgrading as soon as possible.

“If you modernize today, you could substantially reduce the number of servers needed to support the same applications at a much lower power consumption rate,” Privett says. He advises sitting down with folks from the OT side of the house to make sure power is available for whatever you want to do. In many areas, power is at a premium.

If your plans include installing GPU environments in your own data center, WWT advises you not to delay. “We’re telling customers, you need to talk with us and get that designed, get that ordered, because it will take quite a bit of time until it actually ships and we’re able to install it,” Anderson says.

Moor Insights’ Kimball agrees. “You have to order these parts today if you want to see them hitting your dock, your warehouse, or your office 12 months from now.”

Meta minimizes role of token maxing in employee evaluations

Meta won’t judge employees by how much they use AI when it comes to annual performance reviews, despite early efforts to drive AI adoption focusing on so-called token maxing.

The company has told employees that it “will not use AI adoption dashboards or token counts to evaluate impact,” according to  a report by The Information.

The announcement came in an internal memo from executives Maher Saba and Santosh Janardhan, which said that instead of measuring AI usage, “managers should look at output quality, velocity, problem complexity and scope taken on.”

This marks a culture change for Meta, where engineers had previously competed to consume the most AI tokens, displaying their scores on a leaderboard. Meta then discovered that its employees were being diverted from regular work because they were using AI to carry out additional tasks to boost their scores.

Amazon had similar results when it implemented a leaderboard to track AI use; it also found that some employees were trying to game the system by using AI to complete unnecessary tasks, and it has now deleted it.

The company, however, does also monitor employees’ use of AI for training purposes, in a program introduced in April, but this is information was not used to measure employee performance.

Meta had already started to look askance at the concept of using AI metrics as a tool to assess employees. Earlier this year, Chief Technology Officer Andrew Bosworth told employees in a memo that “nobody should be using AI tools just for the sake of using them,” adding that “token usage alone is not a measure of impact of any kind.”

This article first appeared on InfoWorld.

How cost visibility becomes a competitive advantage in FinOps in 2026

As spending on cloud technologies grows, so does waste. The Flexera 2026 State of the Cloud Report found that 27% of organizations expect to spend more on cloud this year, with 17% already exceeding their budgets over the previous 12 months. The estimated share of wasted cloud spend has already crept up to 29%, undoing several years of progress, undoing several years of progress.

Companies are rapidly investing in cloud technology, but often understand less about how to use it fully and efficiently. That is not a coincidence, and it is exactly the gap FinOps is meant to close. It is also why the practice is moving out of the finance department and into the strategy conversation.

What is FinOps?

FinOps is a blended operational framework that maximizes technology value by uniting engineering (DevOps), finance, and business teams. It involves close collaboration to break down silos between tech and finance, with shared ownership of cloud spend across engineering, finance, and business teams.

What distinguishes FinOps is real-time visibility into what is being spent and why. It also treats optimization as continuous work rather than an annual cleanup exercise.

FinOps is important because cloud spending isn’t like a typical budget line. It’s more variable and usage-based, so relying on an annual review doesn’t work. Engineers can quickly create infrastructure, scale it, and tear it back down in a day, making forecasting more challenging than in the past.

What FinOps does is change who sees what. Engineers have more insight into the actual cost of a build. Finance gets numbers it can trust. Business leaders can tie spending directly to business outcomes. It removes much of the guesswork and turns cost data into a shared language rather than a monthly surprise.

As Siarhei Sukhadolski, Chief Delivery Officer & Head of Competence Center at Innowise, puts it, “For a long time, FinOps meant only cutting the bill: find the unused stuff, resize a few instances, and report the savings. That still matters, but it’s not what separates companies today. The ones pulling ahead are using FinOps to make faster, smarter calls about where their tech spend actually pays off. That’s a different job, and it shows up directly in how fast a company can move.”

How FinOps spending has changed

Only a few years ago, FinOps was mostly focused on cloud infrastructure spending. Today, that focus increasingly includes AI-specific investment. The FinOps Foundation 2026 State of FinOps Report found that 98% of organizations now manage AI spend specifically. FinOps has also expanded well beyond cloud infrastructure. It’s more common now to see FinOps coverage extend to licensing (64%), private cloud (57%), and data centers (48%). Around 90% also manage SaaS spend or plan to do so within the next year.

It’s also worth noting that the same FinOps Foundation report found that 78% of teams report directly to the CTO or CIO rather than operating solely within finance departments. To us, that reporting line says a lot. It suggests that companies increasingly see technology spending as a strategic lever rather than simply a line item to reconcile.

How AI and cloud spending are moving in the same direction

The trend toward bringing AI and a broader range of technology spending into FinOps is backed up by a Gartner report, which estimates that global IT spending will hit $6.31 trillion by the end of 2026. That’s up 13.5% from the previous year. Data center systems spending is expected to grow 55.8%, with generative AI model spending more than doubling over the same timeframe. Gartner, in a separate forecast, expects public cloud services to grow by 21.3% in 2026, with the market reaching $1.48 trillion in value by the end of 2029.

We see these figures as two sides of the same shift. AI workloads are also usage-based, which makes them more unpredictable, partly because some teams haven’t had to consider unit economics before. A fine-tuning run or a forgotten inference endpoint can quickly become one of the biggest items on a cloud bill. Most teams don’t have the tagging, forecasting, or accountability needed to catch those costs before they get out of control.

“AI spend just behaves differently from a normal application workload. It spikes, it’s hard to pin on one team or feature, and you often don’t know the real cost per outcome until the invoice lands. Companies that already had solid FinOps habits before AI adoption took off are adjusting faster because visibility and ownership were already part of how they worked. Companies that treated FinOps as an annual cleanup are the ones getting caught out.”

Siarhei Sukhadolski, Chief Delivery Officer & Head of Competence Center at Innowise

Why visibility and shared spend ownership create an advantage

Flexera numbers on discount usage point to the same issue: fewer than 50% of organizations are using the most basic cost optimization tools. The adoption of tools like reserved instances or savings plans is slow, with only 48% of companies using Google Committed Use Discounts and 45% using AWS Reserved Instances. Too many others are leaving low-risk savings on the table.

In many cases, the real problem is a lack of ownership and visibility. If no team owns the cost of a workload, no one has enough reason or enough information to choose the right pricing model. That is where the competitive gap starts to open: some companies can explain and act on their spend quickly, while others cannot.

“A mistake we still see a lot is trying to optimize the bill instead of the system behind it. Deleting unused resources saves money once. Redesigning how workloads scale, how environments get spun up, and who’s on the hook for what keeps costs under control for good. That’s the difference that turns into a real competitive edge later.”

Siarhei Sukhadolski, Chief Delivery Officer & Head of Competence Center at Innowise

What visibility and shared ownership look like

FinOps operates well when at least three structures are in place.

  • Every workload or inference endpoint has a clear owner tied to its cost.
  • Cost and usage data are shared and available to everyone who needs them before they have to ask.
  • There is an ongoing review cadence designed around continuous optimization.

Teams that jump into dashboards before assigning ownership and establishing the data flow often end up with visibility but no accountability. Teams that start with ownership, even with basic tooling, get a different result. They tend to see savings stick instead of resetting every few months. That proper order is the single biggest predictor we have seen across cloud and AI cost engagements.

What business leaders need to know

There is a simple test a CEO or CFO should apply to FinOps. It’s whether the company can clearly state what a workload costs and whether it is worth that cost right now. Does it take finance three weeks to answer? Can the company provide an answer in real time? Are there live numbers and clear ownership behind every workload? That is what will enable leaders to make confident calls on where to invest and where to pull back.

FinOps is becoming a proxy for how well a company manages technology. With shared ownership comes high visibility. Add in continuous optimization and companies gain an advantage. These are not just cost-saving tactics. They represent operational discipline, separating those who can move quickly on AI from those who spend heavily only to find themselves still behind the rest of the pack.

Why Cisco is redefining its CIO role

The CIO job description is being rewritten in real time. As AI agents take over the interface layer and connect directly to any data source, the skills that once defined great IT leadership — UX fluency, applications integration, build-versus-buy judgment — are giving way to an entirely different set of questions surrounding not how a process works, but whether it needs to exist at all.

Thimaya Subaiya is living that shift firsthand. At Cisco, he oversees IT and says the ideal CIO candidate today might not have a traditional IT background. Here, he explains why he split the company’s AI leadership out as its own function and why he’ll merge back in, what he’s really looking for in a CIO candidate, and why the Cisco CIO job is such a good one.

How would you describe your role at Cisco?

I lead operations for one of the world’s largest supply chains, as well as security and trust, including product security, internal systems, and data center security. I also lead the CIO organization and have revenue operations, partnership management, and accountability for our AI strategy. Two and a half years ago, I consolidated AI from throughout the company and named a CAIO. I then split out the role to give us a boost in the AI space, but eventually, the CAIO role will merge into IT.

How did you conceptualize the CAIO role?

At first, it was a leader who could pull use cases from all our operations and execute. The role also included the ethical use of AI systems, and prioritized what to guardrail and push out to employees.

But it’s evolved. To take a step back, Cisco pioneered enterprise networking, then built Compute with Cisco, Storage with Cisco, Networking with Cisco, Security with Cisco, and Observability with Cisco. Today, the CAIO is moving up the stack with an AI framework for MCP connectors, which has really moved us forward.

This CAIO group can tell the Cisco-on-Cisco story for AI, because we have a testbed for new ideas. If we continue to rely on multiple vendors, as in the past, we won’t be able to integrate at scale. This is why we isolated the CAIO role, to focus exclusively on AI governance and execution.

You’re in the middle of a CIO search. What are you observing about the CIO talent market?

With AI, the CIO role has completely changed. It’s no longer about UX and applications integration because with MCP, we can connect to any data source at any time, and agents have replaced the interface. The CIO role is now more about rethinking a process and then deploying an agent to execute, rather than reworking a process.

So the ideal CIO is a traditional one who’s learned to think differently, or even someone without a CIO background, but who’s led in product management, innovation, or transformation. The role today requires someone who’s been disruptive, and has had to rethink how a company operates, not just how its applications work.

Our top criteria are strategy, speed of execution, and the ability to scale because we’re not investing in science projects. For example, when the sales team requests a better forecasting tool, a CIO traditionally would make a build or buy decision. But in today’s world, the right question should be if you need a solution to forecast at all, or can an agent do it. Or better yet, do we even need this process?

So what’s the right background for today’s CIO?

Product managers have a relevant background because they manage multiple aspects of how a product comes together: user needs, business outcomes, fit in the market, and getting it built. This understanding of product strategy, marketing, and adoption is extremely important right now because we treat our AI initiatives like products. So a great path for our CIO is data scientist foundations, product management, and transformation.

What about enterprise security?

I treat enterprise security as a separate organization, which every company should do. Testing and evaluating new cyber solutions for frontier models requires a lot of work like scanning everything, taking a neutral view of what’s broken, deciding which tools become standard within development frameworks, which cryptography tools to use, and then maintenance. Abstracting that into its own organization creates focus. It also lets us move at the speed of AI.

When AI attacks, you need AI to defend you, and if security is embedded within the CIO organization, it’s not top of mind for the business. Security has become its own board-level conversation. For today’s CIO, I’d keep AI in but take security out.

A year after the CIO is in place, what will success look like?

Our applications footprint has been reduced, we’ve seen pure productivity gains from accelerating the back, and the speed of new releases is increased. The team is becoming more effective with the same resources, and we can say that our CIO drove us to leverage everything new technologies offer without blowing up on tokens. We’re looking for a new way to operate IT.

Why is the CIO job at Cisco a great opportunity for the CIO you’re describing?

It’s possibly the coolest job out there. We have an entire AI stack end-to-end that nobody else can claim because we bring networking and security together, complemented by observability and collaboration. That combination means we can create net-new solutions that define what technology looks like in the future.

On the security side, we’re one of the very few companies truly integrating AI into defense in a way that can be leveraged across a much broader market. That’s exciting, because it means free access to an entire stack that lets you innovate in ways the industry hasn’t seen before.

I call AI today’s generational technology. Every generation gets a technology that redefines how it operates, including the internet, iPhone, and now AI. Cisco is about to become the first company to launch a personalized AI agent for every employee, reachable through Webex. Think of it this way: the average person has an IQ of around 100. Now every employee is paired with an AI agent that can exponentially increase human capacity, built entirely on the technology available today.

Getting to build things like that, with no proven methodologies or limitations, and nothing but the question of how we get to the future, is the most exciting thing there is if you’re an innovative leader.

AI agents need to learn when enough is enough

For the past few years, enterprise AI programs have focused on making models more useful, accurate, and autonomous. In that phase, a bad answer was still usually something a human could accept or reject before taking action. But once agents start invoking tools and acting inside business workflows, success should no longer be measured only by how much work they complete. A more important metric is how well an agent recognizes when it lacks the authority, context, or judgment to continue.

When helpful becomes risky

According to Allan Dabre, technology compliance and AI lead at PwC, a behavior that has to be deliberately designed into the system is, “I don’t know.” AI is built to be helpful, so an agent will generally try to do something useful unless it’s been configured not to.

“The fact that AI systems can hallucinate illustrates that tendency,” Dabre says. “When they lack enough information, they may still produce an answer. In an agentic workflow, that impulse can become more dangerous because the output may become an action, rather than remain a suggestion.”

He adds that many enterprises still test AI primarily for completeness and accuracy. That made sense when the central question was if the model could produce a reliable response. But as models improve and agents gain more operational authority, he argues that CIOs need to prioritize something else: restraint.

“Can it stop at the exact moment you want it to stop?” he asks. “Are you testing for that?”

Confidence is not authority

Dabre makes a simple but important distinction. An AI agent may be 99% confident a record should be updated, a refund should be approved, or a legacy database can be decommissioned. But that doesn’t mean the agent has the authority to act. Confidence is about the probability the system believes it’s right. Authority is about whether the organization has delegated that action to the system in the first place.

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Allan Dabre, technology compliance and AI lead, PwC

PwC

He gives the example of an agent asked to analyze legacy software and recommend what can be decommissioned. The agent may conclude, with high confidence, that several databases have little user impact and can be deleted. But even if the system is confident, most organizations wouldn’t want it to delete those databases on its own.

The same logic applies across business processes. An agent may be confident a customer record should be updated, an opportunity in a CRM system should be closed, or a transaction appears legitimate. But once that action flows into other systems, the potential consequences expand.

That’s why Dabre argues for what he calls an agent harness: a controls or orchestration layer outside the model that defines what the agent can and can’t do. In a refund workflow, for example, a company might let the agent approve small refunds, require human approval for larger ones, and stop the process entirely above a defined threshold. The agent may gather the relevant context, explain the request, and prepare the case for review, but the decision is governed by the authority boundary encoded into the system.

“It’s not a policy document and it’s not a prompt,” Dabre says. “It’s software or a configuration you can apply to an agent.”

The case for least agency

Matt Graney, chief product officer at Celigo, a business automation and integration platform provider, approaches the same problem through a principle he calls least agency. The idea is to give an agent the least amount of autonomy required to complete a job.

According to him, there’s a temptation to throw AI at broad, nebulous problems. But many business processes are still largely deterministic. They follow established rules and perform repeatable work. Within those workflows, AI may be useful at the point where rigid rules give way to interpretation. But that doesn’t mean the agent should own the entire workflow. “The smaller you make that surface area, the better,” he says.

Graney says the same logic applies to tools. An agent with too many tools can become confused, especially as context windows grow and the task becomes more complex. “Because Celigo is an integration platform,” Graney says, “the company’s approach is to expose agents to fewer, more powerful tools that reach enterprise systems through governed connections.”

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Matt Graney, chief product officer, Celigo

Celigo

That’s another form of restraint. Instead of letting an agent reach into enterprise systems ad hoc, the business gives it a narrow, governed toolset designed for the task at hand.

Graney also argues that guardrails should sit outside the model. If the same agent that makes a decision is also responsible for judging whether the decision is acceptable, the control is weaker. A separate guardrail can check the agent’s inputs and outputs before a downstream action occurs.

That same design discipline applies to escalation. “I don’t know” shouldn’t be treated as a chatbot phrase. In an enterprise workflow, it’s a handoff path that should be defined before the agent reaches it.

Make escalation part of the workflow

Turning uncertainty into a handoff is where Matt Quinn, CTO at CarGurus, an automotive marketplace, sees agentic AI becoming less a pure technology challenge and more a management challenge. At CarGurus, Quinn says agents are evaluated according to what they know, what they can do, and what data they operate on.

CarGurus receives a high volume of cases from dealers, and each one needs to be classified and routed. The company now uses an agent to review incoming cases, draw on account history, and route them to the appropriate next step. Quinn says the agent handles about 70% of those cases end to end without human involvement.

But when agents move toward consequential actions, he says the consensus is having a human approval step. The agent may return with a simple prompt like, I’m about to do this. Do you want me to proceed? That simplicity matters because a handoff shouldn’t bury the reviewer in complexity.

Quinn says the human remains ultimately accountable for the work. That principle is especially important in engineering, where agents may help write code or fix bugs. Quinn adds that CarGurus still expects engineers to follow the practices they’d use for any other production change, which includes running quality checks.

The company has adopted the phrase healthy speed to describe the balance it wants. The goal is to move faster without letting quality degrade. An agent can accelerate work, but if teams abandon the practices that make work safe, the speed becomes reckless.

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Matt Quinn, CTO, CarGurus

CarGurus

This is also where human judgment remains difficult to replace. Quinn describes it as high judgment people develop through experience. A human may look at an AI-generated output and sense something’s wrong, even before fully articulating why. “Agents are improving,” he says. “But humans still play a critical role in deciding when the system shouldn’t continue.”

That doesn’t mean every workflow needs the same level of review. Quinn says CarGurus doesn’t have a target percentage of work to automate. The right level depends on the job and the task. A simple bug fix may require a lighter review than a change to a sensitive backend service, and a personal summary may carry little risk. But a document sent under someone’s name still needs human review.

Make autonomy accountable

That kind of pragmatic approach may be the best lesson for CIOs, making the goal of agentic AI appropriate rather than maximum autonomy.

That also means ownership has to be clear. Dabre argues ownership should be divided before deployment. The business defines the outcome, technology builds and configures the agent, risk and compliance set the guardrails, and governance monitors whether the system still behaves as intended. The authority to pause, stop, or retire an agent should be defined before production, not negotiated during an incident.

Graney makes the same point with a simple analogy. If a company hires an untrained intern, gives that intern access to the crown jewels of a business process, and something goes wrong, the intern isn’t the real problem. The process is. The same applies to agents. Accountability belongs with the person who owns the workflow.

That may be the shift CIOs need to make as enterprises move from pilots to production. AI agents shouldn’t be treated as magical workers that absorb accountability. They’re components in business processes, and those processes need accountable owners.

As AI adoption increases, the next phase of enterprise maturity won’t be defined by agents that always answer or always complete the task. It’ll be agents that know when not to act.

Economic process modeling: Business cases beyond cost accounting

Cortney Pagel has learned to expect a particular question whenever she proposes changing how work gets done. Pagel, a senior business analyst and change manager at ENGIE Impact, often hears it from the digital side of the organization before she has finished explaining the change: How much money are we looking to save here? “I don’t always have an answer,” she says. “And so that is very frustrating for me.”

The frustration comes from a familiar structural problem. Most financial systems connect money to departments, accounts, cost centers and, in more mature implementations, activities; however, they rarely connect economics to the anatomy of the work itself. Pagel describes processes that are still too manual to be tracked cleanly, while in other cases the financial detail exists but was never tied to the process model. For a period, she resorted to putting cost estimates in comment bubbles on process diagrams because there was no systematic place for them. “There’s been no great system or method to do it,” she says. “It’s definitely not just me.”

The question she keeps being asked therefore exposes a broader lacuna in management accounting. Finance can usually explain what a process consumes, and it can often estimate what a proposed change might save. Yet it is far less equipped to show which process components contribute value, which destroy it, which absorb risk and which create information or options whose economic effects may surface much later. Consequently, a transformation business case can become highly precise about one side of the equation while leaving the other largely narrative.

A ledger on the ledge of usefulness

The general ledger reports what a department or cost center consumes. Activity-based costing (ABC), where organizations have implemented it with sufficient discipline, pushes that resolution further by assigning costs to activities. Both approaches remain useful; nonetheless, their analytical center of gravity is consumption rather than contribution. They can tell management where resources were spent with increasing granularity, while offering much less visibility into what an individual activity economically produced.

Double-entry accounting, dating back 500 years to Venetian merchants, earns its reputation for symmetry, although the symmetry belongs primarily to bookkeeping. The two sides of an entry describe the same financial event, while revenue generally appears when a transaction is recognized rather than carrying a lineage back through the many process components that helped create it. A renewal, expansion, avoided loss, faster decision or improved customer relationship may depend on dozens of steps, yet the contribution of any one step rarely has an account to which it can be posted.

This creates an analytical asymmetry that can influence investment decisions more than finance leaders may realize. When a CFO or operating executive evaluates a proposed process change, the cost side often arrives quantified while the value side arrives as prose, judgment or a collection of indirect metrics. The quantified side therefore tends to carry disproportionate weight because it is already denominated in the unit in which the decision is made: money. Indeed, acknowledged uncertainty may be safer than one-sided precision, because the latter can carry the authority of a number while obscuring what the model omitted.

One process, many economic artifacts

Consider the process of customer onboarding. Operationally, it is a sequence of tasks needed to establish a customer, configure services, obtain approvals and move the relationship into a steady state. Economically, however, those same steps may establish relationship patterns that influence retention, create the account depth that enables a later cross-sell, generate behavioral and preference data whose usefulness compounds over the customer lifecycle, and reduce churn risk through investments made well before the customer has a reason to leave.

Embedded in that same process may be approval controls whose original compliance rationale has waned, manual handoffs between systems that were never integrated, duplicate checks and wait times that gradually erode the loyalty the process was intended to build. Some components may therefore create value; others may protect it and still others may quietly consume it. Yet a conventional cost model compresses this heterogeneous mesh (or mess!) of economic activity into a single process cost, which is useful but incomplete.

Improving or transforming the process requires a more discriminating account of what each component is doing economically: which steps build value, which erode it, which create unnecessary friction, which absorb risk, which generate ancillary benefits and which perform economic work that becomes visible only after the step is removed. Without that component-level view, an efficiency initiative can readily eliminate something valuable simply because its cost was easier to suss than its contribution.

Economoic process model: sample customer onboarding.

Sample customer onboarding — economic process model.

LINQ.it

Putting economics on the process map

Economic process modeling (EPM) supplies that missing layer. As I described in a recent column, Business Transformation Needs a True Economic Approach Rather Than Guesswork, EPM decomposes a process into its constituent components—the information flows, human decisions, system actions and organizational touchpoints that make up the actual work—and then attributes economic effects to each across five dimensions: revenue contribution, cost and friction, risk exposure, option value and information value.

The component level matters because the economically significant finding often sits buried within the process as a whole. Two steps that look roughly equivalent on a process diagram may carry very different economic profiles once attribution is applied. A seemingly minor validation step, for example, may generate information that reduces downstream risk, while a more conspicuous approval step may be largely vestigial. A cost-only review can easily misread the two because it sees effort more readily than consequence.

Pagel describes the capability she wants in similarly practical terms: the ability to see processes at an organizational level, understand what they cost in aggregate and then break those economics down step by step. That level of resolution helps because process transformation decisions are rarely made at the level of an abstract end-to-end flow; they are made by automating, eliminating, combining, outsourcing or redesigning individual components. Consequently, finance needs an economic view at the same level where the design decision is actually being made.

The oft-ignored value of information itself

Information value is where this analytical oversight may be most acute, particularly because most processes today generate data as a byproduct of execution. For example, a credit review produces repayment-behavior signals, a claims intake creates fraud indicators, and a procurement approval accumulates supplier-performance evidence. Those outputs may have future utility well beyond the transaction or process that generated them, even though conventional cost accounting typically has no place to represent them.

Most organizations, however, still treat much of this data primarily as documentation, exhaust or a compliance burden rather than as a potentially monetizable asset. A process redesign can therefore appear efficient while externalizing, degrading or destroying information whose economic contribution was absent from the business case. Infonomics, the discipline of treating information as an economic asset with attributable value, provides the grounding for this dimension of EPM and helps expose value that can otherwise disappear during an ostensibly sensible transformation.

From cost review to capital allocation

EPM extends cost accounting by adding an economic perspective the ledger wasn’t designed for. Sure, cost remains an indispensable computation. However, the business case becomes materially more complete when the components proposed for automation or elimination are also evaluated for revenue contribution, risk absorption, optionality, information yield and the friction they create or remove.

This can change the quality of the capital-allocation discussion. A step that costs $500,000 annually certainly may be a strong automation candidate, yet the savings figure is incomplete if the same step prevents $2 million in avoidable losses, preserves a customer relationship, generates valuable information or creates an option the business may need later. Conversely, a relatively inexpensive step can still be economically destructive if it adds delay, rework or customer attrition. The point is not to manufacture spurious precision around every benefit; rather, it is to make the relevant sources of value visible, estimable and subject to the same scrutiny as cost.

Which brings us back to Pagel and the question she hears whenever she proposes a change: “How much money are we looking to save here?” Savings are only one side of the economic case. The more consequential question may be what each affected component contributes today, what value may disappear if it is changed, and what new value the redesigned process could create.

Indeed, a transformation can look compelling when the savings are visible and the value at risk remains invisible. Economic process modeling gives finance a way to juxtapose both in the analysis, so that a proposed change can be judged not merely by what the organization expects to spend less, but by what the work itself is actually worth.

Why we need technology economists

The advent of AI is precisely why organizations need technology economists, not just IT finance professionals.

IT finance is primarily concerned with budgeting, accounting, cost allocation, depreciation, chargebacks and financial reporting. These disciplines remain important, but they assume a relatively stable relationship between technology spending and business outcomes. AI breaks that assumption.

Technology economics asks a fundamentally different question: How do technology investments create, destroy, shift or delay economic value?

AI introduces a set of economic dynamics that traditional IT finance was never designed to evaluate.

AI creates non-linear economics

In traditional IT, spending $10 million typically produced a somewhat predictable capacity increase or operational improvement.

With AI, a $10 million investment might generate $100 million in value. It might generate no value at all. It could increase costs while appearing successful. It could also create strategic advantages that do not show up in financial statements for years.

A technology economist studies the relationship between technology inputs, organizational capability, productivity outcomes and economic value creation.

IT finance largely records the spending.

AI changes the economics of labor

AI is not merely another technology platform. It acts as a form of digital labor.

Organizations now face questions such as:

  • Should work be done by humans, AI, automation or a combination?
  • What is the marginal cost of an AI-generated transaction versus a human-generated one?
  • How does AI affect productivity elasticity?
  • When does AI create labor substitution versus labor augmentation?

These are economic questions, not accounting questions.

AI simultaneously creates technology inflation and deflation

A fascinating paradox is emerging: AI can reduce costs in some areas while dramatically increasing costs elsewhere.

For example, fewer coding hours. More GPU costs. Lower service desk costs. Higher cybersecurity costs. Reduced consulting expenses. Increased data management expenses.

Technology economists study entire economic systems and value chains.

IT finance often sees only line items.

AI requires measuring economic outcomes, not technology outputs

Historically, organizations measured projects delivered, systems implemented, budgets achieved and uptime percentages.

The AI era requires measuring:

  • Revenue generated
  • Margin improvement
  • Risk reduction
  • Productivity gains
  • Decision quality improvement
  • Time-to-market acceleration
  • Innovation capacity

Technology economists focus on these outcome measures.

This is one reason why AI performance measurement frameworks, including AI-focused balanced scorecard approaches, are becoming increasingly important.

AI introduces massive opportunity costs

One of the largest AI risks is not technological failure.

It is investing in the wrong AI initiatives.

A bank might spend $50 million building an AI solution that saves $5 million annually while ignoring another opportunity that could have generated $500 million in new revenue.

Technology economics focuses on capital allocation efficiency, opportunity cost, marginal returns and portfolio optimization.

These concepts sit outside traditional IT finance.

AI makes technology a strategic production function

Historically, technology supported the business.

Increasingly, technology is the business.

In many industries, AI determines customer experience, operating efficiency, innovation speed and competitive advantage.

Technology is becoming a primary production factor alongside labor, capital and natural resources.

Organizations therefore need experts who understand the economics of technology as a production asset.

AI creates new forms of technical and economic debt

Many organizations are deploying AI rapidly without understanding:

  • Long-term infrastructure costs
  • Model maintenance costs
  • Data quality costs
  • Governance costs
  • Security costs
  • Regulatory costs

A technology economist examines the total lifecycle economics.

The cheapest AI solution today may become the most expensive solution over the next decade.

Why this matters

The central challenge of the AI era is no longer “Can we build it?”

The challenge is, “Should we build it, where should we deploy it, what value will it create, what risks will it introduce and what is the optimal economic allocation of technology capital?”

Those are technology economics questions.

IT finance professionals are essential for controlling and reporting technology spending.

Technology economists are essential for determining whether that spending creates sustainable economic value.

As AI becomes embedded into every business process, the organizations that outperform will not necessarily be those with the biggest AI budgets. They will be those that best understand the economics of technology itself — how AI, data, infrastructure, labor, risk and innovation combine to create measurable business value. That is the domain of technology economics.

Why a cheaper model won’t lower your AI bill

Somebody in finance has already forwarded you the pricing comparison. A downloadable frontier-class model at a fraction of what you’re paying now, with the obvious question attached: why are we still on flagship rates? The honest answer comes in two halves. Open weights are the best thing to happen to enterprise AI buyers since the category existed, and switching to one still won’t lower your bill this year. Both of those are true, and the gap between them is where the useful work sits.

How we got here

In mid-July, a 2.8-trillion-parameter open-weight model shipped with performance close to the commercial frontier, and the full weights followed ten days later under a custom license. Markets moved before Washington did. Semiconductors were hit hardest that session and one widely held chip ETF finished the week almost 9% lower (CNBC). Washington began weighing restrictions on open-weight models soon after, and the industry answered inside a fortnight.

On July 24, twenty-five companies published a letter asking policymakers to leave downloadable model weights alone (Tom’s Hardware). Not a single founding signatory sold access to a closed-frontier model. Three major labs were absent at launch, two signed within 72 hours (TheNextWeb) and the roster passed 270 organizations inside ten days (Forbes). The sole holdout published its position days later, agreeing with much of the letter while disputing two safety claims.

What open weights hand you

Start with what genuinely changed, because it’s larger than the coverage suggests. A downloadable model at frontier-class capability puts a permanent public floor under what that capability can be sold for. But no supplier prices at whatever the market will bear once a comparable input is obtainable elsewhere, and that shift doesn’t reverse.

Stakeholders already know how to think about this.  They just haven’t been filing AI under the right heading, which is concentration risk. A single provider holding a load-bearing production input, controlling both pricing and release schedule, would sit on the risk register in any other procurement category. The only reason AI was able to bypass this was that there was no alternative worth naming. Now there is one.  The leverage shows up at renewal whether or not you ever deploy an open model, since the negotiating position changes the moment the alternative becomes credible.

It changes what you can responsibly commit to, as well. Until now, a multi-year AI investment has meant betting the program on one supplier’s pricing decisions and deprecation schedule, and that’s a hard paper to take into an investment committee. Commitments get easier when the input underneath them has a substitute. Workloads governed by data residency rules come back into scope too, and for some companies that means markets they’d written off.

Investors’ point of view is a little different in this scenario, and probably more accurate. Valuations built on sustained pricing power at the model layer assume something the capability data no longer supports. As models converge, the primary durable margin moves toward distribution, proprietary data and internal workflows that the customers cannot rip out. This happens to be the ground that the coalition’s founding signatories already hold.

None of that requires a single enterprise to switch models. So, the case against restrictions is a real one, whatever mix of principle and self-interest sits behind it. And note one of its own asks: public funding for shared evaluation frameworks, which the signatories evidently agree don’t exist yet.

The monopoly is breaking, just not on the scoreboard everyone watches

Stanford’s 2026 AI Index puts the leading closed model ahead of the leading open model by 3.3% as of March 2026, having been 0.5% ahead in August 2024 (Stanford HAI). The same chapter records six labs clustered inside 25 Arena Elo points at the top and reads that convergence as pushing competition toward cost and reliability. For most enterprise work, a 3.3% capability gap is not a reason to pay a multiple.

Market share tells a different story. Menlo Ventures, surveying 495 US enterprise AI decision-makers, puts three vendors at 88% of the enterprise LLM API market between them, on 40%, 27% and 21% (Menlo Ventures). The same research found enterprises tend to stay with whichever vendor they picked, upgrading within that provider even where switching costs are low.

Both are true and reconciling them is the point. Suppliers price differently when they know you can leave, and that holds whether or not you ever. The alternative never has to be used to change what you pay. So, the pricing monopoly is gone while market share sits exactly where it was. Pricing power was the monopoly that mattered to buyers, and open weights broke it.

The headline price is not the cost

This part is arithmetic. On published rates one recent open model looks roughly a third the price of a leading commercial system. Cost per completed task tells a different story, and the firm that measures it states the mechanism plainly: because cost tracks real token usage, models producing longer answers or more reasoning bill more per task even at identical per-token prices (Artificial Analysis). One current frontier model cost about twice its predecessor per task on that measure, driven entirely by token consumption and not by any price change.

It’s not possible to move a production workload to a budget-friendly mode without having a per-tasking definition of good enough, and literally no one has that handy. Ask what accuracy a workflow requires and you’ll hear crickets, or a number invented on the spot. Ask what it currently achieves and you’ll get the same silence. Same shrug, different meetings. Until both questions have answers, a pricing table is just somebody else’s workload dressed up as your business case.

I work on AI infrastructure, and the primary hurdle that I keep hitting isn’t a technical one. Writing down what good enough means requires somebody to put their name on a number they’ll be held to later. That’s an organizational decision, and not an engineering one, which is exactly why these documents don’t exist in most companies. Teams spend multiple quarters comparing models but barely spend a week agreeing what exactly they’re comparing them for. The related thing I’d say from that seat is that most groups believe they evaluated a model when what they did was try it. Someone ran twenty prompts, liked what came back, and the decision got made in the room. That’s a demo. Demos flatter every model about equally, which is why they can’t tell you whether the cheap one is costing you anything.

Self-hosting won’t rescue the math for most buyers either. A frontier-scale checkpoint runs well past a terabyte, so outside the regulated cases above, the win shows up as hosted providers competing for your workload.

What will actually move your bill

Two things will, and neither of them is a model release. First is the eval infrastructure, since it turns any price difference into a decision that you can defend. Scrape a few hundred real queries from the peak-load and freeze them as your golden test set. Get the people who own the business outcome to write down what a good response looks like, in specifics instead of adjectives. Test the existing model first and generate the scores. Most teams underestimate this step, and it’s the one that makes every future comparison possible.

The second is your own compliance position, which the August decision did nothing to simplify. Every proposal in circulation points at documentation and audit trails, and the holdout lab’s own position points the same direction from the opposite side of the debate. What reaches the buyer either way is a demand for evidence about what your systems can do and what they did. Most 2027 budgets don’t carry that line.

And on the stock market question

Expect volatility on release days and don’t mistake it for repricing. A capable open model lands and chip stocks sell off within hours, one such session costing a single chipmaker close to $600 billion (CNBC). They recover over the following weeks.

My read is that markets keep filing these as demand shocks when they are supply-side price events. Cheaper capability has driven adoption and compute consumption up together every time, which is the opposite of what a selloff assumes. So, keep the two conversations apart. A chip selloff tells you about supplier margins and nothing about your own AI spend, and boards that conflate them freeze budgets during a dip or wave them through during a rally. What would genuinely reprice this sector is a regulatory outcome raising the cost of shipping capability, or an adoption curve that flattens.

Shrewd IT hiring strategies have never been more critical

Major shortages of qualified professionals for key IT roles will lead to huge competitive challenges for organizations that fail to prioritize tech recruiting over the next couple of years, industry observers say.

Hiring the right staff has always been a prime concern for IT leaders, but the pressure to find the right candidates has never been higher, with qualified AI, cybersecurity, and data science professionals especially difficult to find.

Worse, those three domains, along with business/IT automation and risk management, make up the top five areas where CIOs are hiring today, according to CIO.com’s State of the CIO survey. Everyone appears to be hunting the same scarce resources — a market condition that’s already undercutting enterprise opportunities, around AI in particular.

As a result, IT hiring practices over the next 18 months to two years could make or break companies, with laggards risking a huge competitive disadvantage, experts suggest.

Organizations need to think both about hiring outside workers and retraining existing employees to cover gaps, says Adam Wachtel, CTO at employee onboarding platform provider Click Boarding. IT leaders should think wholistically about building capabilities in their teams, he suggests.

“The market for pure AI specialists is volatile and expensive and keeps shifting,” he notes. “What separates organizations right now is whether they’re building AI capability into the team they already have or waiting to buy it fully formed from outside; those building make progress while those waiting are falling behind, and it’s only becoming more expensive.”

There are major implications for organizations that fall behind, Wachtel adds.

“The most immediate risk is technical debt you can’t see accumulating until it’s expensive to fix,” he says. “I’ve lived through rebuilding a team and a platform from a thin, overstretched state, and the lesson that stuck with me is that understaffing or misaligned hiring fails quietly through slower development, more fragile systems, engineer burnout, and more time fixing versus building — it’s not fun for anyone.”

There are several implications for botched hiring efforts, notes Henry Vassal Jones, CIO at outsourcing provider Emapta.

“If you don’t have the people and capabilities to execute, transformation slows, product releases get pushed out, and existing teams carry more of the burden,” he says.

Risk of failure

Critical AI initiatives can fail without the right people in place, Jones notes. “Companies can invest heavily in AI platforms, but without people who understand the business processes, data, governance, and security behind those tools, much of that investment will never reach its potential,” he says.

Jones agrees that employee training, as well as strategic hiring practices, plays an important role in keeping organizations reaching their capacities.

Successful companies will broaden their approaches beyond constrained local or regional talent markets and develop their existing people, he says.

“Those that don’t risk seeing the gap between what the business needs and what their technology teams can deliver continue to widen,” Jones adds. “For CIOs and CTOs, this is no longer simply about filling open positions; it’s about building a talent model that gives the organization access to the right capabilities when needed.”

How to approach the talent challenge

When it comes to developing that talent model, Konstantinos Dolkas, CTO of cybersecurity upskilling and workforce development company Hack The Box, calls for IT leaders to broaden their geographic horizons. While talent is distributed, most hiring strategies still aren’t, he says.

He also advocates employee upskilling. “Recruiting externally can’t be the entire solution,” he says. “Build the majority, buy the scarcity. That could mean a few genuinely senior external hires to set patterns and mentor.”

Given shortages in AI and cybersecurity skills, hiring leaders should also focus more on demonstrated skills from outside hires than the titles they’ve held, Dolkas suggests.

“The best strategy is to hire for demonstrated ability, not credentials,” he says. “Put candidates in a hands-on environment and watch them work. It’s the only screen that survives contact with reality.”

Assessing AI security skills can be particularly difficult in a field that reinvents itself every quarter, Dolkas adds. Another challenge is separating genuine AI fluency from tool familiarity: “Prompting an assistant is not the same as securing an agentic system,” he says.

Anticipate the market and focus on future needs

In addition to building from within, smart IT leaders are focusing on the capabilities their organizations need one to three years from now, says Tom Ioele, CEO at recruiting firm TalentBridge.

IT leaders involved with hiring decisions should think about building talent communities before the demand exists, he says. Organizations should continuously identify and engage with people who have the skills they know they will need, instead of starting to search when a requisition opens, he advises.

“The biggest mistake companies make is treating hard-to-find technology talent like a traditional requisition,” Ioele says. “By the time an AI engineer, cybersecurity expert, or data scientist hits the open market, every company is competing for the same person.”

The companies that win won’t necessarily have the largest recruiting teams, but they will have the best talent intelligence and the ability to activate it faster than their competitors, he adds. Successful organizations will build talent capacity before they need it, he says.

“We’re entering a market where the skills companies need are changing faster than traditional workforce planning cycles,” Ioele says. “Organizations that continue operating through a simple post-a-job, screen resumes, fill-a-seat model will constantly be reacting to yesterday’s demand.”

See also:

Scaling enterprise AI without breaking the bank: A CIO’s guide to AI unit economics

Uber’s experience highlights a new enterprise AI challenge: adoption can scale faster than an organization’s ability to measure economic value. As companies move from AI pilots to widespread deployment, the question is no longer whether employees will use AI — it is whether every AI investment can justify its cost.

Generative AI is changing the economics of enterprise technology. Every inference request, AI agent execution and model interaction can create recurring costs, while cloud infrastructure, GPUs, data, security, integration and governance add to the total cost of delivering AI. The economics that made an AI pilot look compelling can look very different at enterprise scale.

The next phase of enterprise AI will not be defined by the number of models deployed or pilots launched. It will be defined by sustainable business value. For CIOs, CFOs and business leaders, success depends on maximizing business outcomes while controlling the cost of delivering AI.

AI success is an economics problem, not just a technology problem.

AI unit economics: The new measure of AI success

Manufacturers measure cost per unit produced. Banks track cost per transaction. Enterprise AI requires a similar discipline — not measuring how many models are deployed, but how much business value is generated for every dollar invested.

Traditional software investments typically involve predictable costs. AI introduces a dynamic cost structure where every interaction creates ongoing expenses, including compute, inference, storage, data retrieval, monitoring, integration and governance.

A simple framework for evaluating AI investments is:

AI unit economics = (business impact × adoption × reusability) ÷ total cost of delivering AI

Consider an illustrative AI-enabled invoice-processing workflow. If AI reduces processing time, increases straight-through processing and the same capability can be reused across accounts payable, procurement and supplier onboarding, its economics improve not simply because the model is cheaper — but because the value and reuse increase faster than the cost.

This equation reflects a simple principle: AI investments create the most value when they solve high-impact problems, achieve broad adoption and create reusable capabilities while keeping operating costs under control.

Business value may include productivity improvements, faster decisions, improved customer experiences, revenue growth, cost reduction or reduced operational risk.

The objective is not to minimize AI spending. It is to maximize the value generated from every AI dollar.

Understanding the cost drivers and metrics that matter

AI unit economics depends on understanding both consumption drivers and business value drivers. Infrastructure, GPU compute, inference usage, data management, security, compliance and governance all contribute to AI costs.

CIOs should move beyond tracking total AI spend and monitor metrics such as cost per inference, token consumption, GPU utilization, model usage, latency, adoption rates, productivity improvements, automation levels and business impact.

CIOs should treat AI consumption as a portfolio allocation problem — not simply an infrastructure problem.

The winners will not be the organizations that deploy the most AI. They will be the organizations that know where every AI dollar creates measurable business value.

Optimizing AI unit economics: Practical strategies for CIOs

Improving AI unit economics requires more than reducing costs. It demands thoughtful architectural and operational decisions that maximize business value while minimizing unnecessary AI expenditure. The following strategies can help CIOs achieve that balance.

1. Use the right technology for the right problem

Not every business problem requires a large language model. Many structured prediction challenges — such as demand forecasting, fraud detection, predictive maintenance, churn prediction and pricing optimization — are often better solved using traditional predictive machine learning models.

These models typically require fewer computational resources and can deliver comparable or superior performance for well-defined prediction problems.

Large language models create the greatest value for language-intensive tasks such as enterprise search, document analysis, conversational assistants, software development and content generation.

The right question is not, “Where can we use generative AI?” It is, “What is the simplest technology capable of delivering the required business outcome?”

2. Manage AI as a portfolio, not a collection of projects

Many enterprises still evaluate AI initiatives individually. Leading organizations manage AI as a strategic portfolio.

Every AI investment should have clear business objectives, success metrics, ownership and exit criteria. Experiments should either demonstrate measurable value and scale or be discontinued.

A portfolio approach helps eliminate duplicate investments, increase reuse of AI capabilities and shift funding toward initiatives with the strongest business impact.

3. Optimize AI architecture and model selection

AI infrastructure decisions are now financial decisions. Unlike traditional applications, AI workloads create continuous demand for compute resources, making inference costs a major operational expense as adoption grows.

Organizations are increasingly adopting hybrid AI architectures that combine public cloud flexibility with private infrastructure for high-volume, sensitive or regulated workloads. This approach can improve resource utilization, reduce data movement costs, strengthen data sovereignty and create more predictable operating expenses.

However, infrastructure optimization alone is not enough. Enterprises must also ensure that each workload runs on the right model. Not every interaction requires the most advanced — and most expensive — foundation model.

CIOs should adopt intelligent model routing strategies that match workloads with the right models based on complexity, performance and cost. Smaller language models, open-source models and domain-specific models can handle routine tasks such as classification, extraction and summarization at significantly lower cost.

Premium foundation models should be reserved for complex reasoning, advanced analysis and high-value decision support where their additional capabilities justify the expense.

The goal is not to maximize model size or infrastructure investment — it is to optimize AI consumption for measurable business outcomes.

4. Redesign business processes — Don’t just add AI

Adding AI to inefficient processes rarely creates transformational value. The biggest improvements come from redesigning workflows around AI capabilities.

For example:

Traditional workflow:
Employee → AI Assistant → Invoice

AI-enabled workflow:
Invoice → AI Agent → Human Exception Review

In this model, AI handles routine tasks while employees focus on complex decisions.

As organizations transition from basic copilot tools to autonomous agentic AI architectures capable of independent execution, the greatest value will come from designing workflows where AI agents handle multi-step operational tasks while humans focus on exception handling, complex judgment and strategic goals.

5. Measure outcomes and strengthen AI foundations

Providing employees with AI licenses does not automatically create productivity gains. Without clear use cases, adoption strategies and outcome measurement, organizations can increase AI spending without achieving proportional business value.

Leading enterprises focus on value realization by measuring outcomes such as hours saved, productivity improvements, automation rates, customer experience improvements, revenue impact and cost reductions.

However, productivity measurement alone is insufficient. Sustainable AI economics also depends on the foundations that make AI reliable, scalable and trusted. High-quality data and strong governance act as value multipliers by reducing errors, improving adoption and enabling responsible scaling.

Weak foundations can quickly erode AI economics. Poor data increases operational costs by creating inaccurate outputs, more human review, lower employee trust and repeated model execution.

Similarly, governance should not be viewed only as a compliance requirement. As IT leaders navigate the operational costs and requirements of AI governance, strong responsible AI practices — including security controls, explainability, regulatory oversight and human oversight — reduce operational risk while increasing confidence in AI-driven decisions.

Clean data improves model performance, while effective governance ensures AI systems are reliable, secure and scalable. Together, they improve AI unit economics by reducing waste, increasing adoption and maximizing the business value generated from every AI investment.

Measuring AI economics is only useful if organizations build the operating discipline to manage it continuously.

6. AI FinOps: Operationalizing AI unit economics

Cloud computing created FinOps to bring financial accountability to infrastructure consumption. As explored in CIO.com’s breakdown of FinOps expanding beyond traditional cloud costs, managing variable enterprise technology costs requires unified collaboration between engineering, finance and business leaders.  AI requires the same discipline, but with a more direct connection between technical consumption, financial accountability and measurable business outcomes.

The key is connecting technical consumption metrics with financial and business outcomes:

Consumption MetricsBusiness Impact Metrics
Inference cost per transactionRevenue impact
Token consumptionProductivity improvement
GPU utilizationHours saved
Model utilizationAutomation rate & cost savings achieved

AI spending should become as transparent, measurable and accountable as any other strategic operating expense.

Financial discipline is no longer optional; it is essential for scaling AI responsibly.

From AI adoption to AI advantage

The organizations that lead the next phase of enterprise AI won’t necessarily deploy the largest models or spend the biggest budgets. They will make better AI investment decisions.

They will choose the right technology instead of the newest technology. They will redesign business processes instead of simply automating existing ones. They will build reusable enterprise capabilities rather than isolated pilots.

Most importantly, they will manage AI as an economic asset — not merely a technological one.

The future of enterprise AI belongs to organizations that maximize AI unit economics: scaling adoption, reusing capabilities across functions and maintaining disciplined control over infrastructure, inference, operations and governance costs while delivering measurable outcomes.

The future winners will not be those who deploy AI everywhere. They will be those who know where AI creates economic leverage — and where it does not.

A spreadsheet is not a strategy

Picture the meeting. A slide goes up, a number goes down and somewhere in the room, someone claps.

The line item is a renegotiated managed services contract, a hardware order trimmed to “just enough,” or a headcount freeze that quietly became a headcount decrease. Whatever it is, it looks great in the deck. The CFO nods. The COO nods harder. Everyone agrees this was smart.

Three months later, something breaks — an incident nobody can escalate fast enough, a part that doesn’t arrive in time, a senior engineer who finally takes that recruiter’s call. Nobody connects it back to the slide. The slide was right. The spreadsheet said so.

This is the part where I’d like to gently suggest that a lot of very smart people are managing to the cell instead of managing to the outcome — with total confidence, because the cell is the only thing anyone asked them to optimize.

To be clear, this isn’t a jab at the leaders doing it. I’ve done it. I have a Six Sigma certification and a well-worn habit of measuring things, and measuring things is good — right up until the measurement becomes the mission. The problem was never the spreadsheet. It’s mistaking it for a map.

Outsourcing: The invoice goes down, and so does everything you can’t put a price on

I’ve watched this failure mode play out more times than I can count. Across nearly three decades in infrastructure — as chief technology architect at GE Medical Systems (now GE Healthcare), integrating roughly 300 acquired companies into a 420-location footprint — the pattern held: The moment a relationship with the people who actually knew a system got treated as a line item instead of an asset, the organization lost something the spreadsheet never had a row for.

Vendor and MSP contracts are the cleanest modern example: Savings are easy to show, losses are easy to miss. You cut the line item. What doesn’t show up anywhere is the on-call engineer who used to just know — the environment, the history, the thing that broke in 2019 — replaced by a support queue and an SLA that’s met on paper while your business is down in practice.

None of this is the vendor’s fault — they’re delivering exactly what the contract asked for. CIO.com’s own reporting on the hidden costs of outsourcing makes the same point from the other side: Ineffective knowledge transfer and high vendor-side attrition can permanently erode institutional knowledge the client never gets back. The contract took away flexibility. The person who used to just fix it — the one who wore six hats and closed the gap on a Tuesday afternoon — gets replaced by a role with a scope of work. Scopes of work don’t wear hats. A five-minute favor becomes a change request, routed through a ticketing system, against a rate card. You didn’t just outsource a function. You outsourced your ability to handle it — and bought back a slower, costlier version of the same fix, one billable hour at a time.

JIT procurement: A factory formula applied to a business that isn’t a factory

Just-in-time assumes something that doesn’t exist: A crystal ball good enough to see today’s need and whatever shows up next. I learned that the hard way at GE Medical Systems, when a new customer opening a facility wanted several hundred patient-critical bedside monitors customized to match a color scheme from their marketing department. Our processes were built entirely around clinical function — the thing that keeps a patient alive — and nothing accounted for a hospital wanting its equipment to match its brand. It came in from left field. We had no SKU for “must match burgundy.”

We ended up standing up a new department — Specials — because the existing process had nowhere to put a request like that. Building the flexibility after the fact was expensive. But it became a real differentiator: As far as I know, we became the first and only medical device manufacturer with a dedicated specials department. It told customers something that mattered more than paint color: They came first, and we’d find a way to say yes.

The lesson wasn’t “predict better.” It was “build systems with teeth” — flexibility designed in, not bolted on after reality shows up sideways. Dell and HP figured this out decades ago: You can order a PC built to your exact spec and have it shipped in days, because their systems were engineered for change. Most IT organizations still build for demand they can already see, then treat every surprise as an exception instead of the job itself.

This wasn’t a one-off: Supply chain analysts at SupplyChainBrain noted that during the 2021 chip shortage, many manufacturers found their lean JIT models weren’t built to flex under real disruption. “Just in time” only works until the time arrives and the thing isn’t there.

Headroom is savings, too — paid out in advance instead of on the back end, which is why it never gets credit. Nobody puts “the department we didn’t need to build in a panic” on a savings slide, because avoided cost doesn’t announce itself the way cut cost does. The expense of headroom is visible and immediate; the expense of its absence is invisible until it isn’t. It’s an incident report.

The hour that reads as free

I lived a version of this at GE Medical Systems. We built life-critical patient care products in a market crowded with giants — Philips, Siemens, HP — where nothing shipped until it cleared FDA review. On one release, scope crept weekly because sales kept promising new capability to close deals, and no one above us would draw a line around what “done” meant. What we got instead of a defined scope was a war room: Catering, a fridge stocked with Mountain Dew, enough M&Ms to open a candy counter — everything money could buy to keep engineers at their desks around the clock, except the one thing that would have actually helped: Someone willing to tell sales no.

We hit the deadline. The product cleared FDA review. When it shipped, there was no “great work,” no pat on the back — just the quiet message that this was expected of us. We won on the software. We failed on the people. That’s sunk-cost thinking in its purest form: Once a team’s extraordinary effort becomes the baseline, the extraordinary disappears the same way the ordinary already had.

Salaried time reads the same way on every spreadsheet I’ve seen since: Already paid for, so effectively free. Nothing stops it from being spent — on the meeting that could’ve been an email, on the ticket queue treated as bottomless, on “just have IT handle it” as the default answer. Burnout doesn’t have a line item either, until it shows up as attrition, and attrition finally does, at which point everyone acts surprised. It’s not small: Gallup estimates disengaged employees cost the global economy trillions a year — roughly 9% of global GDP, sitting outside any single department’s budget. The spreadsheet didn’t lie to you. It just never had a cell for the thing that mattered most.

The ledger nobody built

Zoom out from these three stories and the pattern is the same: Reporting structure decides which questions get asked. When technology reports through a CFO or a COO, the question every quarter is “what did this cost us today?” Almost never “what did this cost us to keep?” An organization that only asks the first will keep hiring smart people to answer it well — in exactly the wrong direction, forever.

None of these leaders are bad at arithmetic. Most are excellent at it. The tragedy isn’t the math — it’s the ledger: Precise, defensible calculations against books never built to hold the costs that matter most, and calling it leadership. I’ve seen this enough times to give it a name: The leader who runs a technology organization strictly by the numbers handed to them, gets good at it and never gets fired for it — not because they succeeded, but because the failure never had a cell to live in. The spreadsheet balanced. The building didn’t burn down that quarter. They got promoted.

That’s the actual scandal, worth saying to the room and not just the page: A CIO who has never once been wrong on a savings initiative hasn’t been managing technology. They’ve been managing a spreadsheet, and calling the absence of visible damage “success.”

Before the next savings initiative gets a round of applause, three questions worth asking honestly, out loud, in front of people:

  1. What does this cost that will never appear on an invoice — and am I certain, or just unbothered?
  2. Who inherits that cost, and will I still be in this seat when the bill comes due?
  3. If I can’t put a number on it, have I decided it’s zero — and whose job was it to notice first?

The savings will still show up in the deck. If nothing else shows up beside it, that isn’t restraint — it’s the tell.

Why AI TCO is so tricky — and how to start calculating it

Achieving return on investment is impossible without knowing the total cost of ownership (TCO) of an initiative — and when it comes to AI, CIOs are finding cost calculations anything but straightforward.

Subscription and token costs are a big part of the calculus, but several other factors go into the cost of AI projects, says Ben Schein, chief AI and analytics officer at AI data platform provider Domo. Chief among those are cloud infrastructure costs and the human time involved in guiding or correcting AI outputs, he notes.

In addition, many organizations have multiple divisions using different AI tools for vastly different purposes.

“There’s not like a single ledger,” Schein says. “Right now, and maybe for the foreseeable future, there’s sort of like a multiple ledger approach to how all this works.”

A shifting paradigm

While token costs have dropped significantly in the past two years, costs vary wildly between models and AI providers, and the price drops are often offset by increased usage. And AI providers have also explored other kinds of consumption-based pricing, including API calls, compute time, or documents processed.

All this makes it difficult to measure TCO, Schein says.

“You have sort of these subscriptions, you have the consumption and the tokenization, you have some of the infrastructure you might be paying for,” he says. “There’s also a human tax that introduces new time for verification and review, and if the AI is sloppy or creating slop, you might be inadvertently adding to your costs without knowing it.”

It’s difficult to measure TCO because AI doesn’t have a single cost center, agrees Shane Cronin, head of FinOps and ITAM services at systems integrator SHI.

“By the time you’re looking at the bill, you’re dealing with token consumption, cloud infrastructure, multiple AI models, governance tooling, integration work and, increasingly, autonomous agents making decisions across systems,” he says.

IT leaders at many organizations still define AI success through narrow technical metrics instead of prioritizing business outcomes, Cronin adds.

“Calculating token costs is relatively straightforward,” he adds. “Calculating whether those tokens actually created measurable business value is much harder. That’s where most CIOs are today.”

Unpredictability and hidden costs

Michael Moran, chief technology and information officer at contact center outsourcing provider NQX, sees several other factors leading to further unpredictability over AI costs.

For example, data center costs are rising, AI vendors are starting to shift from subsidized pricing to profitability, and organizations have increasingly complex AI use cases, he says.

“IT leaders should temper expectations that AI inherently reduces costs,” he adds. “Instead, it’s important to understand that full automation is likely to be prohibitively expensive for most enterprises, and that brands will need to balance AI investment with human engagement strategies that improve long-term value rather than cut costs in the short term.”

If AI implementations work exactly as expected right out of the grate, TCO should be relatively easy to calculate, he says. But agentic AI implementations often require much more human training and intervention than expected.

“These are the hidden costs that are often underestimated or ignored altogether when initially calculating TCO,” Moran adds.

Visibility is the first step

Chris Cagnazzi, chief innovation officer at IT solutions provider Presidio, is one IT leaders seeking to get a handle on the complexity of calculating AI TCO by applying playbooks from the cloud migration era.

Cagnazzi has adapted Presidio’s cloud cost optimization platform, PRISM, to track AI costs internally and to help customers do the same.

The first step toward tracking AI costs is visibility, he says. IT leaders should know every model running across their organizations, the cost per user per month, and what kinds of prompts each user is writing, he explains, adding that Presidio is using real telemetry to track internal AI use, as well as internal tools to direct prompts to cost-efficient AI models.

The second, more difficult, step is turning visibility into action, he adds. “You have to think about mapping the usage back to the owners, whether it’s users or groups,” he says. “Then you look at, what are some of the anomalies? And if you’re looking at those anomalies, do you have governance in place around overspend?”

What Presidio has found is that the bill for AI services represents only about 30% of the total cost, Cagnazzi notes.

“The other costs really lie in areas around the hidden AI stack,” he says. “Those things around orchestration or retrieval, observability of the guardrails, or the rereads and the redos. There’s a lot of cost that people are missing.”

While traditional IT costs can be fairly predictable, AI costs are driven by usage and can increase because employees are repeatedly using inefficient prompts, Cagnazzi says.

“The spend is hard to forecast; it’s hard to see the true hidden costs behind the bill,” he adds. “If that prompt is less efficient, it might produce a bill that’s 100% higher than what it should be.”

The good news, says Domo’s Schein, is that IT leaders have a lot of variables to play with to control AI spending. They can encourage users to use more cost-efficient AI models, they can track employee usage of AI, and they can test different prompts and other interactions for cost effectiveness, he says.

“The price spread on the different models is crazy,” he says. “You could say, ‘I have no ROI on this investment; if I could get the same outcome with a model that costs one-30th as much, I may have ROI.’”

Meta’s plans to replace workers with AI fell flat, report says

Earlier this year, Meta, one of the industry’s loudest AI advocates, was ready to slash up to 60% of the members of some teams and replace them with AI, as part of what it called Project OT (Organization Transformation), an initiative to make Meta “AI native.”

But it backed off at the last minute after internal data showed that the plan wasn’t working out, according to a Reuters investigation published Wednesday. For example, Reuters said, code changes made to the internal software platforms and infrastructure that employees used on the job were up 220% year-over-year, according to an early June post by Meta CTO Andrew Bosworth, yet changes that led to new or upgraded features reaching Meta users were only up 36%.

Meta executives also saw “’reliability warning signs’ caused by the AI coding surge,” according to an internal post, Reuters reported. “Another post, in April, said that unchecked AI agents were performing ‘large-scale, disruptive actions that humans are unlikely to execute.’ The result: Major technical and security incidents, such as service disruptions and possible data leaks, spiked 40% from the previous year, with the time staffers had to spend firefighting them up 70%.”

Reuters also noted that, in April, Meta had mandated that tracking software be installed on US employees’ devices to capture their keystrokes and mouse clicks to teach its AI agents to replicate how humans interact with computers. Believing they might be training their own AI replacements, employees rebelled.

To try and placate the workers, Meta promised to increase spending on travel and social events, and “to improve snack quality in office microkitchens.” Unsurprisingly, none of that seemed to help boost morale, and, Reuters reported, “Zuckerberg has stuck to the words ‘company-wide’ and ‘this year’ in discussing layoffs with employees, according to his internal communications. That has prompted some employees to speculate that he’ll continue trimming the ranks via team-specific cuts or performance-based dismissals – or delay company-wide headcount reductions until next year.” 

A cautionary tale

Consultants and analysts said enterprise IT executives should read the Meta story carefully, because it precisely illustrates what happens when AI marketing hype is not challenged aggressively.

Noted Sanchit Vir Gogia, chief analyst at Greyhound Research, “Meta trusted a forecast of AI capability before it existed in production, a different failure from trusting AI too much.”

He said, “Meta booked a forecast as capacity. Agents will improve, but the error was budgeting that improvement as production capacity before it arrived. Prove the action before widening the authority, and the authority before removing the human control. Only then is removing human capacity a decision, not a bet.”

Terra Higginson, a principal research director at Info-Tech Research Group, added that no experienced enterprise IT leader should be surprised by Meta’s experience. 

“Unchecked AI agents are a bad idea. Removing humans is a bad idea. That’s not the future anyone wants,” she said. “We want work to be reimagined so the human part still matters and technology makes it better. We are all still wrapping our heads around how AI and agentic AI will change the way we work.”

She noted, “what we are already seeing, though, is lots of output and action without always getting the outcome we actually want. We should not use AI output as a proxy for productivity. Humans bring judgment and friction before taking actions with significant consequences; agents can remove that friction. We don’t want easy outcomes, we want good outcomes.”

Tom Findling, CEO of Conifers.ai, also suggested that IT leaders should take the Meta report as “the best opportunity” to go to their board and argue that this is what happens with unchecked AI rollouts. 

“Tell them that we now have the opportunity to get it right. Say that you may not get a 500% productivity boost, but IT can show them a meaningful way to get 300%,” Findling said. “If you don’t want to end up like Meta, there is a way.”

Justin Greis, CEO of consulting firm Acceligence, added that he thinks that IT’s takeaway from the Meta situation is the disconnect between activity and actual value creation.

“AI can make an organization extraordinarily busy without necessarily making it more productive,” he said. “We have spent decades teaching technology leaders that lines of code, tickets closed, and projects launched are imperfect proxies for business value. AI makes that measurement problem much more acute because it can manufacture activity at machine speed.”

If an AI agent produces ten times as much code, analysis, or work product, that does not mean the enterprise created ten times as much value, Greis said. “It may simply mean the company created ten times as much material that somebody now has to validate, secure, integrate, maintain, or clean up. That is why I think the most important question for executives is not ‘How much work can AI produce?’ It is ‘What measurable business outcome improved because AI produced it?’”

This article originally appeared on Computerworld.

Google adds pay-as-you-go Gemini pricing as enterprises seek control over AI spending

AI agents could make software development and other enterprise tasks more productive, but they are also making technology spending harder to predict. Unlike traditional software licenses, the cost of running an agent can vary depending on the models it uses, the number of tokens it consumes, and how long it runs.

Google on Wednesday added new pricing options, discounts and cost-management tools to Gemini Enterprise that it says are aimed at helping enterprises reduce the cost of certain AI workloads while giving enterprises better visibility into where their AI budgets are going.

As part of the new pricing options, the hyperscaler introduced a pay-as-you-go model and Flexible Savings Plans (FSPs).

While the pay-as-you-go model allows enterprises to pay for the compute and tokens they consume instead of committing to a base subscription, which in turn avoids paying for empty seats or unused capacity, the FSPs offer discounts of 10% for one-year commitments and 20% for three-year commitments on Gemini Enterprise spending.

Flexible pricing lowers barriers, but adds new trade-offs

For enterprise teams and their CIOs, the pay-as-you-go model lowers the barrier to adoption and is better suited to experimentation, temporary projects, and agent workload bursts, said Stephanie Walter, practice lead of AI stack at HyperFRAME Research.

“While Per-seat pricing forces you to buy capacity before you know if an idea is worth it, the pay-as-you-go lets you spin up an agent experiment on a Friday afternoon and only pay for what it actually burns,” echoed Manoj Chandra Jha, principal analyst at Nord-IQ Research.

That means the newer pricing model also removes procurement friction from the experimentation loop, Paul Chada, cofounder of agentic AI startup Doozer AI, pointed out.

“When a pilot requires a license commitment, every experiment needs a business case. When it’s metered, an engineer can run the pilot on Tuesday and show finance a real bill on Friday. That shortens the distance between idea and evidence, which is where most enterprise agent programs die,” Chada said.

However, these advantages come with their own set of trade-offs, especially predictability.

“One user request can trigger an opaque chain of model calls, reasoning steps, and tool invocations, so consumption can grow much faster than employee headcount with the possibility of surprise bills at the end of a billing period,” Walter said.

That unpredictability also means the new pricing model does not automatically translate into cost savings, echoed Jha.

“It’s mostly a shift, not a discount. But matched to the right workload, it can save real money: bursty, unpredictable agent usage no longer subsidizes idle seats, while steady, high-volume usage may still be better suited to a committed plan. The savings come from matching each workload to the right pricing model, not from pay-as-you-go being cheaper by default,” Jha added.

However, the FSPs have their own caveats, especially the three-year plan.

While the FSPs can offer meaningful savings for enterprises with steady or growing AI usage, the three-year commitment is harder to justify in the wake of models, prices, and application architectures changing so quickly, Walter said, adding that the commitment is not just financial but also about choosing a platform as well.

Further, the analyst cautioned that the FSPs are less suitable for enterprises that have yet to establish a reliable baseline for consumption, as committing too early could turn an unpredictable operating expense into a predictable overcommitment.

Currently, FSPs are available for self-serve customers and customers already on enterprise agreements.

The pay-as-you-go model, though, remains only available to select customers with the hyperscaler planning a broader rollout “soon”.

Deferred execution trades speed for lower inference costs.

In addition, the hyperscaler is introducing a third cost-cutting option that is based less on how much an enterprise consumes than on how quickly it needs the result.

The option, named deferred execution pricing, will allow enterprises to mark eligible agent workloads for execution during off-peak capacity windows, with Google offering discounts of up to 50% on inference costs in return, the hyperscaler said in a statement.

Deferred execution pricing, it added, is aimed at workloads that can tolerate delays, potentially giving enterprises a way to lower costs for background tasks and other agent workloads where an immediate response is not essential.

That up to 50% discount in inference costs, Walter pointed out, can be material for CIOs at scale.

However, they must decide which work can safely wait and whether delayed tasks still meet the business requirement, Walter cautioned, adding that Deferred execution fits tasks such as evaluations, document processing, indexing, batch summarization, code analysis, and other background tasks.

Further, the analyst warned that CIOs also need to consider a “development tax” when considering deferred execution: “If agents need to be redesigned to accommodate real-time vs. asynchronous execution, the engineering effort required to build and maintain those different workflows can offset some of the savings.”

But even for workloads that can tolerate those trade-offs, the option will not be immediately available, with Google initially limiting deferred execution pricing to select workloads only. Details of which workloads are eligible were not immediately available.

New FinOps tools target AI spending visibility

Separately, Google is also adding an AI spend anomaly detection capability with root-cause analysis and pairing centralized billing reports with a FinOps agent that can generate natural-language summaries of where AI budgets are being spent.

While the anomaly detection capability is designed to flag projects where AI spending is trending higher than normal and identify the top three SKUs driving the increase, the FinOps agent is intended to make it easier for CIOs to understand where their AI spending is going.

Anomaly detection as a feature, according to Walter, can be valuable for agentic workloads because their consumption can increase through loops, retries, or unexpectedly long execution chains that users never see.

“Identifying what is driving a spike shortens the investigation for CIOs and enterprise teams,” Walter added.

The FinOps agent, meanwhile, Jha said, could help CIOs and other business leaders ask questions about AI spending that traditional dashboards were not designed to answer. However, its usefulness will depend on the quality of the underlying cost attribution, as the agent can only explain spending it can accurately associate with particular teams, projects, or workloads, Jha added.

10 steps to implement an effective AI training program

It’s no surprise that reaping the rewards from AI requires careful guidance, especially in helping staff use tools safely and productively. Yet evidence suggests some CIOs and their executive peers aren’t providing the level of guidance employees require.

While three-quarters of IT staff have access to AI tools, one in five technologists are expected to self-learn, and 23% are waiting for formal training, according to the recent Harvey Nash Tech Talent Salary Report, which surveyed over 3,600 technology professionals globally.

The research suggests AI explorations are commonplace, but tailored learning and development initiatives are not. Digital leaders who want to turn AI into a value-generating opportunity, though, must educate their staff. But what elements should AI training schemes include? Here, industry experts offer 10 steps to implement an effective program.

1. Take a comprehensive approach

Michael Cole, chief technology officer at the DP World Tour, the men’s professional golf tour that oversees 42 tournaments in 25 countries, says AI training is an organization-wide effort.

“I’ve asked the training coordinators in our HR department to help me deliver what I believe is going to be a fit-for-purpose training and development program for not only my IT team here at the European Tour, but equally across the business,” he says.

Cole says the crucial element to emphasize is that AI and the range of capabilities it brings is about much more than learning how to use technology. “Using AI effectively is about process, mindset, and culture,” he says. “So, when we start to think about the training and development needed to bring an organization like ours into this AI-enabled era of transformation, it’s a comprehensive program that must extend across the business.”

2. Educate the boss

In an organization-wide program, everyone needs AI education, including the boss. That’s why Emmanuel Frenehard, chief digital officer at biopharmaceutical giant Sanofi, says his firm takes a multi-layer approach to AI training.

The executives there completed Drive Digital, a program that Sanofi designed with the ESSEC business school in Paris. The initiative focused on core considerations, such as use cases and value generation. After 150 managers passed through the program, it was extended to more than 1,000 other professionals across the organization.

“Don’t just look for the solution; don’t just think about Claude or ChatGPT,” says Frenehard, referring to best-practice lessons. “Think about the challenge you’re trying to solve. In our case, that approach means focusing on what we’re doing, the value we’re looking to create, and the dependencies the project will create.”

He says training also needs to help AI doubters overcome their fears. “You have to make it fun and as risk-free as possible,” he says. “People shouldn’t feel they need to be super-technical to use AI productively.”

3. Build clarity and agency

Jo Bishenden, chief learning officer at tech training and talent provider QA, says AI education is often treated as a one‑off awareness session, a compliance requirement, or something reserved for technical specialists. 

The best programs get three things right. They provide a baseline for everyone across the organization, the courses focus on role-specific applications to show how AI impacts everyday activities, and they provide continuous learning to encourage a behavior change as new AI tools are introduced.

“When done well, organizations see better return on AI investment, improved productivity, and more confident decision‑making,” says Bishenden. “Employees gain clarity and agency, understanding how AI augments their expertise rather than replaces it. Ultimately, AI success isn’t determined by the technology alone, but by the capability of the workforce using it.” 

4. Put the human in the loop

Ankur Anand, group CIO at recruiter Harvey Nash, says AI training is often a work in progress, with his firm’s research suggesting one in five technologists are expected to self-learn. “There’s a rush to deliver the tools, but then organizations aren’t investing enough in enabling the capability of the people,” he says.

While technological skills like prompt engineering are an important part of AI learning and development, Anand said the best programs go beyond IT expertise to ensure humans in the loop have thorough understanding of their responsibilities.

“There are so many softer elements that need to be handled as part of AI training,” says Anand. “Good training is about using the tool as well as the governance and risk frameworks that need to be changed accordingly.”

5. Showcase individual successes

Louise Newbury-Smith, head of UK&I at Zoom, says it has AI enablement teams at the local and global level. And while the company provides courses and self-learning opportunities, Newbury-Smith says the enablement element brings AI training to life.

“Our approach is about showcasing individual successes, making it real, and repeating best practices,” she says. “We have what we call a Cook Along session with our AI evangelists. We’ll do those sessions together a lot as a group, and that makes the process fun. If you’ve got champions who can share incredible successes, then that goes a long way.”

She says the key to success is sharing knowledge. “We’re very much focused on the human,” she adds. “All the services, content, and direction of AI is about how we can give humans time back so they can have more valuable interactions with other staff to empower them with the information they need.”

6. Focus on the finer details

Dan Cherowbrier, CTO at Formula E, the motorsport championship for electric cars, is another digital leader whose business focuses on enablement. The company has a dedicated AI engineer who helps employees exploit emerging technology.

“We’ve got an innovative culture and we weren’t short of ideas of what we could do with AI,” he says. “What we needed were the resources to get people going, get the technology tested, and get it out there.”

The AI enablement engineer works with other tech specialists in the company to ensure tools are deployed safely and securely. “We’re beefing up our data and AI team so we can help users across the business plug in and understand APIs, get access to data, run security checks, and then put AI into production,” he says.

7. Develop reusable skills

Murali Swaminathan, CTO at technology firm Freshworks, says there’s so much information about AI models that people can easily take the wrong direction without guidance.

“We’re trying to give our staff structured learning,” he says. “We understand they’re not all on the same page. Some are ahead of others so you need to provide knowledge that applies to their specific job roles.”

Swaminathan says senior managers discuss how to train people effectively, as AI experiences and capabilities vary considerably across business units. However, the chosen pathway to AI learning and deployment must suit the individual and the company.

“I had this challenge with my engineers,” he says. “Initially, we gave them four different tools. Everybody was using AI, but it was so inconsistent, and everyone was trying to do the same thing in different ways. So we’re now trying to build reusable skills. And that approach must be replicated for every job function.”

8. Learn by doing

Luke Gebb, head of global innovation at American Express, says the financial services firm has various training programs. Having seen AI education in different forms, he advocates for learning by doing, or as a second-best strategy, watching someone else use the technology.

“Hearing or reading about AI, or being presented with something where you’re not actually seeing it happen is not nearly as helpful,” he says. “The best thing is to get a homework assignment and try something.”

Gebb says this approach plays out regularly across the people working in his 120-strong innovation group. The team runs one-hour show-and-tell sessions where an employee demonstrates how they use AI tools in their everyday activities.

“Then they get a bunch of questions, they post their best-practice lessons, and then others try the same thing. It’s an approach that works really well.”

9. Use pioneering techniques

Stephen Wood, COO at Rathbones Asset Management, says AI training in his organization is mandatory. “We want everyone to be versed in different types of AI,” he says. “We’re not expecting everyone to be a coding genius and an expert in all this stuff, but everyone needs to understand it.”

The firm takes a proactive approach to training, using education sessions and spreading best practices via digital champions. The company also embraces pioneering techniques, including running a hackathon to help identify in-house capabilities.

“The hackathon showed that with some searching on Google and YouTube, you could start to create agents that could do basic functions,” he says. “That process taught us, with the right training, and repeated sessions and continuous development, we wouldn’t necessarily need to hire people to create big productivity gains. That was quite an exciting moment.”

10. Evaluate new possibilities

Emerging technology can’t exist in a vacuum. Bernhard Seiser, VP of digital, data, and IT at AOP Health, says anyone using AI must be aware of potential consequences. “It’s your responsibility to validate whether what you’ve created is correct,” he says.

Operating in a regulation-heavy industry means AI training is linked to data governance. “We leverage it in areas where compliance isn’t an issue,” he says. “For example, writing text, creating images, and so on. Certain things can be done.”

As new AI tools emerge, AOP Health will consider its options and develop a training program. “That approach could mean bringing in specialized tools for specific tasks,” says Seiser. “It’s part of my job, and part of my team’s job, to evaluate AI for each use case.”

Nobody knows where their AI budget is going

The early stages of AI adoption focused primarily on getting organizations to adopt AI-enabled solutions such as copilots, agents, assistants and AI-powered workflows. The priority was demonstrating that AI had the potential to deliver value, not scrutinizing every dollar spent along the way.

At the time, this made sense. There were relatively few people using AI, budgets were often funded through innovation initiatives and the cost of experimentation was low compared to the potential upside.

Organizations treated AI spending as a learning experience, so finance teams had little reason to scrutinize every model call or workflow because the priority was learning what AI could do, not optimizing what it cost.

When AI moves from pilot to production

After AI goes from pilot programs to production environments, the financial implications of using AI become very different. Every time a user enters a prompt, calls a model, retries an action, takes an action with an AI agent or executes an AI workflow, the total amount of money spent on AI goes up.

The amount at stake is rising quickly. Gartner expects worldwide spending on AI to reach $2.59 trillion in 2026, an increase of 47 percent from 2025. As companies put AI into more products and daily tasks, even small inefficiencies will repeat across millions of requests and add up to substantial costs.

The movement toward autonomous AI agents will only accelerate this trend. An agent may take much longer to perform its task than a single chatbot interaction. It may also use several different models, interact with other systems and continue to operate autonomously until it completes its task. Since organizations are likely to deploy many agents, AI costs will increase due to two factors: more people using AI and AI performing more work.

These concerns are already affecting which projects survive. Gartner predicts that more than 40 percent of projects involving AI agents will be canceled by the end of 2027 because of rising costs, unclear value or weak controls. A company may approve an agent because it works during a pilot, then reconsider it once thousands of people begin using it and every task triggers several paid requests.

A small number of experimental interactions will become thousands or millions taking place throughout an organization each day. Organizations that were previously asking how to expand their use of AI will now be asking: Do the benefits derived from each AI initiative outweigh the long-term operational costs associated with it?

This change in perspective is a good thing. Organizations are starting to look at AI as an operational expense instead of a shiny new toy. They are now evaluating AI initiatives based upon whether the initiative’s benefits exceed its long-term operational costs, rather than evaluating them based upon enthusiasm for the new technology.

Organizations can see the total AI bill but not where it comes from

The biggest issue for organizations is understanding where those costs originate.

Model providers typically send invoices detailing total usage for a given period. Companies know how much they spent on AI, but they often cannot explain which workflows generated those costs, whether that spending created meaningful business value, or which teams should ultimately own it.

Research from the FinOps Foundation shows that practitioners rank controlling the cost and use of tokens in software delivered as a service as their top concern related to AI. The reasons include bills that reveal little about what caused the expense and systems that provide no built-in way to trace costs back to the people or work responsible for them. A total on an invoice cannot tell a company whether one useful product feature caused the expense or whether an agent repeatedly called a model without improving the result.

That lack of visibility represents a significant blind spot. Imagine receiving your monthly cloud infrastructure bill without knowing which applications consumed your compute capacity, or receiving your monthly utility bill without knowing which buildings consumed your electricity. Most organizations would never tolerate that level of uncertainty elsewhere in their technology stack, yet many are currently managing AI spending in precisely this way.

Without attribution, organizations cannot determine which AI systems are delivering measurable business value, which workflows are inefficient or where unnecessary costs are accumulating.

Optimizing AI spending requires visibility into workflows

Much of today’s discussion about optimizing AI spending revolves around selecting a lower-priced model or negotiating better pricing with model providers.

Those discussions are valid because pricing is one of the few variables organizations can easily measure. Lower model prices alone will not solve the problem. As AI becomes embedded in more workflows and autonomous agents perform more work, organizations often consume far more tokens than they save through lower pricing. The greater opportunity to reduce AI costs often lies in workflow design.

OpenAI provides a clear example. Developers often send the same instructions or previous conversation back to a model each time they make a request. OpenAI introduced Prompt Caching in 2024 so developers could reuse material the model had recently processed and receive a 50 percent discount on those input tokens. The company later increased the discount to 75 percent for repeated material sent to its GPT-4.1 models. The savings come from changing how an application sends information to the model, which means a company can lower its bill without choosing a less capable model or negotiating a new contract.

Optimizing AI spending therefore requires organizations to develop visibility into how work flows through their AI systems. Organizations need to understand how agents interact with one another, how workflows execute, where redundant processing occurs and which steps provide the greatest value relative to their cost.

They also need to know when a system retrieves information it never uses, repeats a failed request, sends the same material several times or calls an expensive model for work that a cheaper one can complete. Each decision may add only a fraction of a cent to one task, but the same mistake repeated across millions of tasks can erase the financial benefit the system was supposed to produce.

Once organizations gain this level of understanding, they can optimize intelligently instead of simply selecting the least expensive model.

AI spending will drive organizations toward greater selectivity

Over the last several years, the AI industry has focused on determining whether AI belongs everywhere. The next stage will focus on identifying where AI creates the greatest value.

Some workflows will produce enough business value to justify substantial AI investment, while others simply will not. Long-term success will depend on understanding where AI creates meaningful value, where the costs outweigh the benefits and which AI systems actually justify their ongoing expense.

Companies should begin by recording which team, product, customer and task caused each paid request. They should compare that expense with the result the system produced, set limits that stop agents from retrying work indefinitely, and alert the people responsible when the cost of a task rises unexpectedly. Engineers can then inspect the costly work, remove repeated steps, reduce the amount of information sent with each request or choose a less expensive model when the quality remains acceptable.

A monthly invoice arrives too late and says too little. Companies need to trace spending while the work takes place, assign responsibility for it and decide whether the result earned its cost. Those practices will help leaders determine where AI deserves more investment, where the system needs repair and where it should be turned off.

5 hard truths of change management

Mohan Sankararaman calls the old approach to technology-driven transformation — the kind of change that lands every few years and reshapes the organization in one push — a trap. As executive vice president and CIO of First Horizon, a regional bank headquartered in Memphis, he’s focused on driving digital transformation the way a bank funds risk: incrementally with room to pull back.

Every CIO is under similar pressure to rethink change management for the AI era. Wanda Wallace, managing partner at Leadership Forum, has advised CIOs on change management for years, and she thinks the job itself hasn’t changed much.

“The hardest and most critical aspect of making change happen and stick is convincing people to adopt a new approach,” she says. “AI doesn’t change that need or that process. It is a human-to-human dynamic.”

Talk to the practitioners and researchers closest to the work, and a version of her view emerges again and again. What has changed is how many things are competing for an organization’s limited capacity to absorb them — AI chief among them. Here are five hard truths IT leaders face about change management today.

1. There’s no finish line

Ashish Parmar, CIO of Standard Industries, a global industrial conglomerate with more than 20,000 employees across roughly 50 countries, has watched the nature of transformation shift beneath him. In the past, he says, change was treated like a project with a start date and an end date — whether the trigger was a new ERP system, a reorg, or a cost-cutting mandate. That model doesn’t hold anymore.

“Today, change is continuous,” Parmar says. “Our strategy is focused on building resilience and adaptability rather than getting to a single destination.”

AI is the clearest example of how the old model breaks down, says Fran Maxwell, who leads Protiviti’s people and change practice, though he’s quick to note it isn’t the only one. Unlike an ERP rollout, which lands as a discrete event, an AI transformation keeps moving.

“The technology evolves continuously, use cases emerge rapidly, and the impact on roles is often uncertain,” Maxwell says. The common misstep is treating any major shift, AI-driven or not, like a one-time project with a training curriculum and a communications plan, he says. The fix is building a permanent capability for adaptation rather than staffing up for a single push.

None of that continuous adaptation is possible if the underlying systems can’t support it, notes Manosiz Bhattacharyya, CTO of Nutanix.

“Technology is not the barrier to transformation; application modernization is,” he says. Years of accumulated dependencies, legacy integrations, and fragmented data are what actually slow an organization down. And layering new tools on top doesn’t make that debt disappear.

“Applying AI blindly does not remove technical debt,” Bhattacharyya says. “It amplifies it.”

2. Bandwidth isn’t just a network problem

A 2026 survey of roughly 3,000 HR leaders by talent firm LHH found that no single cause dominates why companies reshape their organizations: AI and automation, skills mismatches, M&A activity, and strategic shifts were each cited as drivers in the previous year by about a fifth of respondents. In other words, most organizations are contending with several forms of change at once, not just AI.

All that change at once runs into a hard limit: An organization can absorb only so much at a time.

“Every organization’s capacity for change is finite, so leaders cannot endlessly stack new initiatives on top of existing workloads,” Parmar of Standard Industries says.

Rather than treat that ceiling as a constraint, he argues CIOs should use it to force discipline. IT leaders should determine their non-negotiables and point the team’s energy there instead of spreading it thin across AI pilots, reorganizations, and everything else competing for attention.

Sankararaman arrived at nearly the same conclusion at First Horizon. Banking used to reward slow, occasional overhauls, the kind that could take years to prove out, he says. But that approach has become untenable.

“It’s tempting to treat transformation as one big initiative, but with technology evolving this fast, that’s a trap,” he says. Instead, Sankararaman releases funding in stages, each tied to a measurable result before the next is approved. Then, the organization can adapt and build confidence as it goes rather than betting everything on a single multi-year plan. “We reward progress, not perfection,” he says.

Kevin Martin, chief research officer at the Institute for Corporate Productivity (i4cp), has data that supports the value of incremental improvements. When leaders want to move faster, the reflex is to restructure: delayer, widen control, redraw reporting lines. i4cp’s research found no statistical relationship between those structural moves and organizational agility or market performance. What separates agile organizations are routines: scenario planning, faster resource reallocation, clear decision rights, continuous workforce planning, targeted reskilling, and disciplined execution.

“You don’t reorganize your way to agility,” Martin says. “You build it into how the organization operates.”

3. Shadow IT doesn’t belong in the shadows

Employees finding their own tools to get work done isn’t new. Shadow IT has taken on various forms over the years, from personal file-sharing accounts to unsanctioned SaaS subscriptions.

Today, it’s shadow AI, and Sankararaman argues most CIOs still treat it as a security or compliance issue rather than what it really is: information about what the organization needs and isn’t getting.

“Shadow AI is already happening in every organization. If you’re not addressing it through your change management strategy, you’re addressing it too late,” Sankararaman says.

Sankararaman’s approach starts with curiosity rather than restriction: understanding what employees are trying to accomplish with the tools they’ve found on their own, which makes it easier to agree on how the business should govern those tools.

“That’s a change management conversation, not just a policy conversation,” he says.

4. Trust has to be designed in, not repaired later

Every new system that changes how decisions get made must earn trust before it achieves adoption. Agentic AI raises the stakes because it doesn’t just inform decisions; it takes actions on its own within workflows.

That’s a fundamentally different dynamic from anything change leaders have managed before, Sankararaman says. The resistance it produces is often quieter, showing up in questions about how a system reached a conclusion, who’s accountable when it’s wrong, and whether it’s replacing what someone does. “Those questions deserve real answers, not reassurance,” he says.

At First Horizon, IT is building trust into the foundation with permissioned access, centralized guardrails, human oversight, and outputs that are consistent, reviewable, and explainable. “If people can’t understand how the technology reached a conclusion, you haven’t earned their trust,” Sankararaman says. “And without trust, adoption doesn’t hold.”

Employees today worry less about learning a new tool than about what it means for their role, their skills, and how their performance will be judged once a machine does part of the job.

“They worry about career relevance, accountability, job security, and how performance will be evaluated,” Protiviti’s Maxwell says. Closing that gap, in his view, takes more than a rollout plan. It demands transparency about what’s changing, what isn’t, and how people add value once the tool is in place.

5. Tired isn’t the same as unwilling

Ask IT leaders about change fatigue, and they frame it as a capacity problem rather than resistance. Late 2025 saw a wave of five-day return-to-office mandates that landed on top of continued layoffs. Tech companies alone cut more than 66,000 jobs between May and November, according to Newsweek and TechCrunch, exactly the kind of concurrent disruption that erodes an organization’s capacity for change.

Among what i4cp calls “coasting incumbents” — companies that still perform well despite low organizational agility — 51% of employees report finding change fatiguing, and just 8% say change management is an organizational strength. At “agile pacesetters” — the highest-agility, highest-performing organizations in i4cp’s research — only 12% report high fatigue.

“AI is an accelerant,” Martin says. “But organizational friction is the fuel.”

Protiviti’s Maxwell argues that change fatigue is often less about resistance and more about capacity. “Employees are far more likely to embrace change when leaders are clear about what matters most, what success looks like, and just as importantly, what is not a priority right now,” Maxwell says. His advice to CIOs: Empathize and prioritize before accelerating.

One structural fix most CIOs underuse is shared ownership. “Don’t go at it alone,” advises First Horizon’s Sankararaman. “Partner with others across the business, your CHRO, CFO, COO, and make them co-champions of the change, not just stakeholders who get updates.” A message that arrives from multiple leaders, he’s found, carries more weight and lasts longer than one delivered by IT alone.

The absence of fatigue, in Wallace’s view, is its own warning sign. “If your organization isn’t change-fatigued,” she says, “then I am worried about what you have been doing.”

Inside TIAA’s massive IT transformation to fuel business growth

When Sastry Durvasula joined TIAA in early 2022, he saw an organization fighting against outdated legacy technologies and in need of a major IT refresh.

Since then, the financial services organization has completed two phases of a comprehensive transformation initiative called Technology Ecosystem Transformation, or TETRIS, leading to a huge reduction in tech debt and a major expansion of functionality for customers.

The ongoing project, anchored in cloud and AI technologies, started in 2023 with phase one that modernized the core technology stack with 10 new enterprise platforms. Phase two, launched in late 2024, went further by enabling 87 use cases across all major lines of the business.

The project, for example, allowed TIAA to launch its MyChoice Multi-Year Guaranteed Annuity product, and helped create the TIAA Gateway portal, an API-based suite that integrates with partners in retirement and wealth planning using industry standards.

TIAA Gateway took home a CIO 100 Award in 2025, and phase two received a CIO 100 Award in 2026.

Durvasula, TIAA’s chief operating officer, pitched the multimillion-dollar TETRIS project to the board as a three-pronged strategy, with empowering business growth, fueling innovation, and transforming the IT core as its key goals.

Not only did TETRIS need to modernize the company’s IT systems, decommission legacy processes, and automate other processes, but Durvasula pitched it as the way to expand the reach of TIAA’s products and move the company into the future.

“As you expect in a company of our size, we have problems of yesterday, today, and tomorrow being solved at the same time,” he says.

Focus on business use cases

As TETRIS moved into phase two, project leaders shifted their goals from pure technology modernization to business outcome-driven prioritization. So once phase one delivered needed IT platforms like a data cloud and design studio, TIAA pivoted toward enabling business use cases.

This business-first approach ensured continuing executive support and clear ROI at every key milestone, TIAA says.

In 2022, just before the project launched, more than 80% of TIAA’s IT workloads resided in fragmented, end-of-life platforms, which created operational risk, compromised security and resilience, and constrained its ability to innovate. Through TETRIS phase two, however, the organization has cut that tech debt nearly in half.

And consolidating 17 design systems also led to digital products looking and behaving differently, depending on the team that designed them, and accelerated product launches by 35%, enabled multi-lingual capabilities, and increased accessibility to more than 185,000 customers who don’t speak English.

In addition, TETRIS allowed TIAA to combine multiple middleware systems and data lakes, Durvasula says, and the organization moved mainframe applications and data center infrastructure to the cloud.

A giant leap forward

TETRIS has been a huge project, with the company saying it empowered TIAA to have one of the largest leapfrog moments in company history in its submission for the 2026 CIO 100 Award.

Despite the reported failure rates of large transformation projects — some estimates suggest up to 95% fail to meet their goals — TETRIS was essential to keep TIAA competitive and move it forward in the market, Durvasula says.

A big part of the project has been workflow modernization, he says, because TIAA were using some technologies and workflows that were decades old.

“There’s your classical platform and application rationalization, and then there’s your end-of-support, end-of-life stuff that should’ve been remediated long ago,” he says. “Some of the processes we have, because we’re such a large, old company, were designed when the internet just came along.”

Stick to the metrics

Two keys to pulling off such a large project are establishing metrics for success and transparency with leadership, Durvasula says. Project leaders set milestones to indicate when things went well, and they planned for bumps in the road so the TIAA board knew when setbacks happened.

“Not everything is as pretty as it sounds in an awards application, but the success measures we established with our board were based on both phases,” he says. “For the first one, we said we’d deliver enterprise-grade platforms and accomplish migration objectives, but not tied to any specific business objectives.”

Phase two metrics focused more on business objectives, and the project team kept the TIAA board updated as TETRIS moved forward. Setting realistic goals was important, he says, with the team determined not to overpromise results.

“Large programs have a range of objectives, and if you publish the outcomes you’re looking for, people start looking for them, especially stakeholders, the C-suite, and board,” he says. “You have to be honest about which metrics or KPIs you can deliver in the first and second year, and when you’ll start seeing real business scale and impact, which definitely won’t be that soon in a large program like this.”

Goals also need to be flexible, Durvasula says, so transparency with leadership sometimes means telling them the project needs to reset. “If something doesn’t go well, what’s the level of fungibility you have?” he says. “We pick this tool, but what if it doesn’t work? You need to have a plan B.”

So TIAA’s IT team is heavily focused on flexible systems, and what was contemporary three years ago is probably legacy now, especially thanks to AI.

The power of change management

Another big lesson from a project of this size is the need to focus on change management. Retiring old IT systems requires the organization to bring employees along on the journey and convince them the changes are for the better.

TIAA established a multi-disciplinary team to implement a change management program focusing on breaking down silos and setting common adoption goals across the organization and its lines of business. Stakeholder forms and a huge focus on continuous collaboration helped employees understand the need for the changes.

“It’s a big organizational change,” Durvasula says. “If you’re working on a legacy system, and you think at some point it’s going to be modernized, then you become a legacy talent, and won’t have a job.” But the right change management program can convince these employees they can upskill and bring value to the new systems.

“You can bring your functional knowledge of the business and learn new technical skills,” he says. “It’s a massive culture- and people-change initiative as much as tech initiative.”

TIAA’s change management efforts were also made easier because TETRIS happened at the same time as the recent AI boom and involved AI elements. So it wasn’t hard to convince employees they needed to improve their AI skills.

“Because of AI, everybody woke up to this new reality,” he says. “We rode that wave when transformation drove from a cultural and organizational change management point.”

The GPU bill is the new AWS bill

The call usually opens with praise. The AI feature shipped on time, users love it and engagement charts are pointing the right way. Then finance closes the quarter, and the feature everyone celebrates loses money on every single request. That is the part the CTO called about. I get some version of this call every week. I work in developer relations at a GPU cloud provider in Silicon Valley, putting me in the room, or at least on the video call, when engineering teams decide how to buy and run AI infrastructure. The longer I do this work, the more familiar the pattern becomes. I watched companies learn cloud-cost discipline in the 2010s, usually after an end-of-month bill delivered a nasty surprise. GPU spending is the same lesson with two important changes: the hardware costs roughly ten times more per hour, and mistakes pile up faster. We’ve seen this movie.

We know the ending

Before moving into AI infrastructure, I spent years in data analytics at an automotive software company. One part of that job was cleaning up a decade of accumulated cloud enthusiasm, which sounds harmless until you inherit the bill. After we consolidated three overlapping analytics platforms into one, we cut about 220,000 dollars a year while keeping every capability intact. That money accumulated through reasonable-sounding subscriptions, one after another, because nobody owned the basic question: what did it cost to produce those numbers? The industry still hasn’t solved it. Flexera’s annual State of the Cloud research has for years found that organizations estimate more than a quarter of their cloud spend is wasted. An entire discipline, backed by the FinOps Foundation, grew around squeezing that waste back into a manageable shape. It took most companies years to learn those habits.

What bothers me is simpler: I keep seeing solid engineering teams drop that discipline the moment the purchase order says GPU. AI spend gets treated like a bold bet instead of an operating cost, and then the ordinary scrutiny disappears. That is where the trouble starts. The waste patterns of 2015 come back wearing 2026 pricing. The invoice tells you what you paid, separate from what you earned.

The invoice tells you what you spent, separate from what you got

GPU capacity is priced by the hour, so teams naturally budget and report by the hour. It feels neat. It lines up with the bill. And it hides the problem that really crushes margins. The number that decides whether an AI feature survives is cost per request: everything you spend on inference infrastructure divided by the requests you serve. Those two metrics line up only when your hardware stays busy. For user-facing AI, that stays rare. Traffic moves with human attention, so it flares for a few hours and then drops off a cliff.

One team I worked with had reserved a cluster built for a peak that showed up for about two hours a day. On the invoice, the hourly rate looked almost cheap. Once we divided it by served requests, it was ugly, and the team had honestly seen it for the first time when we ran the numbers together on a call.

That division is the most useful exercise I can offer a reader of this column. Take last month’s total inference spend. Divide it by the number of requests you served. If the answer makes someone in the room go quiet, you have found money and you found it with arithmetic a spreadsheet has been waiting to do for you.

Workload shape, rather than vendor choice, decides the right pricing model

When the number looks ugly, the instinct is to push for a better rate or go hunting for another provider. I sell GPU capacity for a living, so I’ll say it plainly: the rate is rarely the issue. The unit price of AI compute keeps falling; Stanford’s AI Index has documented inference prices dropping by orders of magnitude in just a few years. That still leaves a team paying for capacity it barely touches. Waste eats the discount whole.

The fix lasts longer when you match the buying model to the workload itself, which is a point Andreessen Horowitz made well in its guide to the cost of AI compute: access to compute matters less than the shape of the commitment you sign for it. AI workloads usually split into two very different cases, and they want opposite deals. Sustained work, such as training runs, fine-tuning and batch processing, keeps hardware busy around the clock. This is what reserved or dedicated capacity is for. The economics are simply better when the machines stay hot. Reserved or dedicated capacity is made for that, and the per-unit economics pay you back for the commitment. Spiky work, which covers almost everything with a person on the other end, is the reverse. Usage-based pricing earns its markup there, because you only pay when you serve. The per-unit price goes up and the total bill drops. Finance teams resist that sentence until the numbers hit their own sheet.

The best production setups I see are hybrids. A team keeps a modest baseline, sized to the floor of traffic, the level demand almost never sinks below, and lets usage-based capacity soak up the rest. Teams under roughly ten million tokens a month often skip infrastructure entirely and stay on a model-as-a-service API until volume justifies the switch. The reserved slice stays busy. The bursts stay covered. The architecture quietly records a choice the team meant to make, which is rarer than it should be.

3 questions are cheaper than a contract

When a team asks me to review a GPU commitment, I keep coming back to the same three questions, and I would rather they ask them before the signature than after it.

  1. What does our measured utilization curve look like? Skip the neat projection in the deck. Instrument a week of production traffic before signing anything at all. Teams almost never predict their own curve correctly, and that surprise costs nothing before the contract, then plenty afterward.
  2. What is our cost per request at ten times today’s volume? Scale can change the answer, sometimes in our favor. Spiky demand may smooth as users spread across time zones, shifting the calculation toward reserved capacity later. If nobody in the room can answer, the organization is buying a snapshot, not a strategy.
  3. What would switching cost us? Open-weight models let us rerun the analysis with any provider, then act on what the numbers say. Proprietary endpoints tie your costs to another company’s pricing whims. Either route can make sense, yet flexibility has a dollar value and deserves space beside the hourly rate.

The discipline is the differentiator

Outside my day job, I’ve judged more than eight AI hackathons this past year, at Microsoft offices in Chicago and Mountain View, plus events with OpenAI and Google Developers Group. Even there, surrounded by teams building through a weekend, I can see the production problem waiting ahead: brilliant models, minimal thought about what serving them will cost. Nobody wins a hackathon with a unit economics slide. Plenty of companies quietly fail without one.

Years in data analytics left me with a conviction I repeat to every team willing to listen. A dashboard nobody costs out is a liability; an AI feature carries the same risk. The companies that survive the next pricing cycle will be the teams able to name their cost per request from memory and explain their infrastructure in one sentence, with a week of traffic data behind it, rather than those squeezing the lowest hourly rate from a vendor. Ten years ago, cloud bills taught engineering leaders to ask what their systems cost. Now the GPU bill is asking again, at ten times the stakes. The lesson lands harsher now: guessing survives only until the next ugly bill arrives at the worst time. The teams that move first will claim the margin everyone else is still chasing.

❌