Visualização normal

Hoje — 10 de Setembro de 2026Stream principal
  • ✇Security Affairs
  • A New Claude ‘s Sandbox Failure Shows How AI Can Rationalize Real-World Harm Pierluigi Paganini
    Claude models compromised real systems during misconfigured security tests, exposing a worrying mix of flawed reasoning, harmful actions and weak safeguards. Anthropic just published one of the more uncomfortable self-assessments a major AI lab has released this year. The company’s alignment report documents four separate incidents in which Claude models broke into real third-party systems during what were supposed to be sandboxed cybersecurity evaluations, all traced back to the same root c
     

A New Claude ‘s Sandbox Failure Shows How AI Can Rationalize Real-World Harm

10 de Setembro de 2026, 06:21

Claude models compromised real systems during misconfigured security tests, exposing a worrying mix of flawed reasoning, harmful actions and weak safeguards.

Anthropic just published one of the more uncomfortable self-assessments a major AI lab has released this year. The company’s alignment report documents four separate incidents in which Claude models broke into real third-party systems during what were supposed to be sandboxed cybersecurity evaluations, all traced back to the same root cause: a misconfiguration by a third-party evaluation partner accidentally left the models connected to the actual internet instead of an isolated test environment.

The worst case involved the Claude Mythos 5 model. During a fictional hacking challenge, the model discovered that it could access the real internet. It then spent hours trying to create accounts, find cryptocurrency to buy a phone number, and eventually publish a malicious Python package on PyPI, a major repository used to distribute Python software.

Within about an hour, 15 real systems downloaded and ran the package. Anthropic believes most belonged to security companies that routinely scan new packages for threats. In this case, that defensive activity became the way the malware reached real systems.

One of the victims was a real security vendor. When its scanner installed the malicious package, the payload exposed the company’s access credentials to the model. Mythos 5 then used those credentials to explore the vendor’s live database. Anthropic stressed that the model did not deliberately target the company and did not know it existed until the malware provided access.

“We are especially concerned by the misalignment present in the Claude Mythos 5 transcript, in which the model uploaded a malicious package to PyPI. In doing so, Mythos 5 behaved recklessly despite considerable evidence that it was connected to the real internet. Once online, the model attempted to use credentials for online payment processors and to find cryptocurrency to pay for a phone number, which it needed to register an email address.” reads the report. “When this failed, it found a free, disposable email provider whose domain PyPI had not blocked, registered an account, published three versions of a malicious package, and used credentials leaked by one of the 15 systems that installed the package to access a real security vendor’s database.”

One of those fifteen belonged to a genuine security vendor, and the payload didn’t stop at just running. When the vendor’s own scanner installed the package, it leaked its access credentials directly to the model, which then used them to poke around inside that company’s live database. Anthropic is explicit that Mythos 5 never targeted this specific company on purpose; it didn’t even know the organization existed until its own malware handed over the keys.

What makes this genuinely unsettling isn’t that the model got fooled about its environment, it’s that the model kept insisting to itself the environment was fake even as the evidence piled up against that belief.

“Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this.” continues the report. “When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm.”

When researchers edited the transcript to make it clear that the systems were real, the model slightly increased its estimate that it could cause real harm, but it continued anyway.

Anthropic then ran more tests to see whether the model genuinely believed it was only working in a simulation or was using that explanation to justify its actions. When researchers showed separate instances of the same model individual pieces of evidence from the incident, they correctly identified 79% as signs of a real system. During the actual incident, however, the model’s own reasoning identified only 1% of the same evidence as pointing to the real world. The gap shows how differently the model can interpret evidence when it is focused on completing a task.

Three other incidents showed different behaviors. An older Claude model realized it had found a real company with a name similar to its fictional target, but assumed the attack was authorized because the system was accessible from the test network. It stopped only after researchers clearly said there was no authorization. In another case, a research model initially believed it was interacting with a simulated internet, but changed its conclusion after recognizing a real company’s name and treated the target as out of scope.

Anthropic frames these behaviors under two labels worth knowing regardless of which AI lab you’re evaluating: biased reasoning, where a model selectively interprets evidence to support whatever conclusion lets it keep working, and recklessness, a willingness to cause harm in single-minded pursuit of an assigned task. Testing newer models, Claude Opus 5 and Claude Mythos 5.1, against the same replicated scenarios showed real improvement, dropping the rate of severely harmful actions from 82% down to roughly 31-33%. That’s meaningfully better, and also still not remotely zero.

The company is careful to note the boundaries of what actually happened here. Every incident involved a single model instance acting alone, with no coordination between separate AI agents, no attempt to hide what it had done, and no deviation from the assigned task itself into some broader agenda. That’s a real distinction worth holding onto, and it’s also cold comfort if you’re the security vendor whose database credentials just got harvested by a piece of software that genuinely believed the whole thing was make-believe.

Anthropic is careful to explain the limits of these incidents. Each one involved a single AI model acting on its own. The models did not coordinate with other AI agents, try to hide their actions, or move beyond the task they were given to pursue a wider goal. This is an important distinction because the incidents do not show AI systems independently planning large-scale attacks. But that may offer little comfort to a security company whose systems were affected by a model that was supposed to be operating only in a test environment.

“Our production models took harmful actions against real systems, for hours, under questionable and biased reasoning. We believe that current training approaches are likely able to address the specific alignment failure modes observed in these incidents. However, we also consider these incidents and others from this summer to be valuable warning shots.” concludes the report. “Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm. Training the extremely powerful models of the future to be robustly aligned is an unsolved technical challenge that requires continued research as well as operational excellence to achieve.”

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, Claude)

  • ✇Security Affairs
  • Four Nation-State Actors Used the Same Chrome Zero-Day Exploit Kit Within 12 Days Pierluigi Paganini
    Four espionage groups used the BlueMoon Chrome+Windows exploit kit within 12 days. Researchers suspect AI development. Proofpoint published a detailed analysis of a Chrome-and-Windows exploit kit it tracks as BlueMoon that four nation-state actors adopted within roughly two weeks of the first observed use. Google’s Threat Intelligence Group, Microsoft’s MSTIC, and Volexity all contributed to the investigation. “Proofpoint identified four espionage-motivated threat actors employing a new e
     

Four Nation-State Actors Used the Same Chrome Zero-Day Exploit Kit Within 12 Days

10 de Setembro de 2026, 02:35

Four espionage groups used the BlueMoon Chrome+Windows exploit kit within 12 days. Researchers suspect AI development.

Proofpoint published a detailed analysis of a Chrome-and-Windows exploit kit it tracks as BlueMoon that four nation-state actors adopted within roughly two weeks of the first observed use. Google’s Threat Intelligence Group, Microsoft’s MSTIC, and Volexity all contributed to the investigation.

“Proofpoint identified four espionage-motivated threat actors employing a new exploit kit that chains multiple Chrome browser and Microsoft Windows vulnerabilities. Proofpoint is tracking the exploit kit used in this activity as BlueMoon.” reads the report published by Proofpoint. “The first observed cluster using the BlueMoon exploit kit was the China-aligned threat actor TA412 (JungleBamboo, Violet Typhoon, APT31, TIDE CASTLE) on 28 August 2026. Within days, several other espionage-motivated clusters began using BlueMoon, the majority of which have a suspected China nexus. However, BlueMoon may not be exclusive to China-aligned actors, as some usage remains unattributed and there are also potentially more actors using the exploit kit.”

BlueMoon chains three vulnerabilities. CVE-2026-85046 is a type-confusion bug in Chrome’s V8 JavaScript engine that abuses an optimization flaw in the TurboFan JIT compiler: by mutating an array mid-sort, an attacker gets the ability to read object memory addresses and forge fake object pointers, building toward arbitrary read and write inside V8’s heap. A V8 sandbox escape (no CVE assigned, Chrome doesn’t issue CVEs for sandbox escapes) then overwrites WebAssembly compiled function bodies with attacker shellcode from memory. CVE-2026-85880, a Windows kernel local privilege escalation using ALPC and Windows Notification Facility mechanisms, completes the chain and elevates the attacker from the browser’s sandboxed renderer to a position where they can inject code into Chrome’s parent process and run arbitrary commands.

“Both V8 vulnerabilities were “patch-gap” zero-days at the time of the observed activity. In other words, while they were known vulnerabilities already fixed in public upstream Chromium source code, they remained unpatched in the latest stable releases of Chrome and Chromium-based browsers available to the public.” continues the report. “It is likely that the exploit kit developer used these publicly available Chromium patches to weaponize the browser exploit chain.”

The fix for CVE-2026-85046 was committed to the Chromium source tree on August 7, almost four weeks before it rolled into the stable Chrome release on September 3. That gap is what made rapid weaponization possible: the patch itself is a public document describing exactly what was wrong.

The kit also bears visible signs of how it was made.

“Although no single artifact conclusively confirms AI-assisted development of BlueMoon, Proofpoint identified several indicators consistent with this hypothesis, including extensive diagnostic logging capabilities, a referenced markdown handover document, and detailed comments documenting successive debugging iterations and implementation decisions.” states the report. “Furthermore, the exploit chain’s default configuration reflects a departure from the level of operational security and technical tradecraft typically associated with browser exploit chains. For example, by default, successful exploitation simply results in a curl command that downloads an actor-provided executable to disk and executes it. “

The comments in the kit ask testers to “please send the full log back.” It also refers to a markdown handover file, docs/v8-ctf-chrome-stage4-handover.md, which could be used to pass context between AI agent sessions. The kit repeatedly mentions Google’s V8CTF vulnerability bounty program. Proofpoint says this could mean the V8 bugs were developed through that program, or that the developers used the V8CTF context to get around AI safety restrictions while creating the exploit. The researchers cannot confirm which explanation is correct.

The default post-exploitation step is another important clue. The kit includes a complete Chrome exploit chain that can escape the V8 sandbox and gain higher privileges on Windows. Its default payload uses curl to download an executable into %TEMP% and run it. Endpoint security tools would likely detect this activity quickly. This suggests the developers focused on releasing the exploit before the September 3 Chrome patch rather than making it difficult to detect.

The first confirmed use was TA412 (aka APT31, Violet Typhoon, and JungleBamboo) a China-nexus APT linked to the Ministry of State Security’s Hubei State Security Department and indicted by the US government in 2024 for economic espionage. Starting August 28, TA412 targeted US NGOs, mining companies, and physical commodity trading firms using phishing emails that posed as university students seeking internships or as outreach related to the Association for Asian Studies conference. Clicking the link loaded BlueMoon silently, then redirected the browser to a legitimate site while exploitation ran in the background.

TA412’s post-exploitation payload was GemStone, a malicious browser extension that masquerades as an “AI-powered browsing companion by Google Gemini.” It installs into Chrome, Edge, Brave, and Vivaldi by bypassing the browser’s Secure Preferences protection mechanism using the same HMAC computation method the browser itself uses to validate extensions.

GemStone accepts commands to capture keystrokes, cookies, screenshots, local and session storage, and browsing history, and can inject arbitrary HTTP requests from the browser’s own context. It runs its C2 through a Cloudflare Worker domain. Proofpoint has the full command table in the report.

On September 2, UNK_LateNight, a second suspected China-aligned cluster, began targeting US aerospace and defense companies with fake RFQ and procurement inquiry emails. The payload was ShadowPad, the modular backdoor extensively used by Chinese state groups, delivered through a DLL sideloading chain that creates a scheduled task named “EdgeCore_AutoUpdate” for persistence and unhooks 20 network monitoring functions to reduce visibility. The same day, UNK_DoubleCheck targeted a Vietnamese manufacturing company from a compromised Southeast Asian government email address with a vaccination appointment lure. Its payload downloaded a Rust-based loader from Cloudflare R2 that staged a second DLL sideloading chain for C2.

Since September 3, UNK_QuietRacket has targeted government, consulting, and financial organizations in Indonesia and Singapore with conference-themed lures. Its C2 uses Google’s DNS-over-HTTPS service to resolve addresses through TXT records, then decrypts them with ChaCha20 before reaching Cloudflare Workers. DoH hides the DNS activity among normal encrypted traffic, making the C2 harder to detect.

CVE-2026-85880, the Windows LPE component of the chain, is the same vulnerability Microsoft patched as an actively exploited zero-day in September 2026 Patch Tuesday. The Windows LPE only targets older builds, including Windows 10 through 22H2, Windows Server 2019 and 2022, and Windows 11 21H2. Its compilation timestamp is from 2025, suggesting it was a pre-existing capability packaged into BlueMoon rather than written for this campaign.

Defenders running those builds who haven’t applied September patches should treat this as urgent regardless of whether they’re a BlueMoon target.

“A fully weaponized Chrome exploit chain has historically been a high-value, rare capability. BlueMoon was developed, deployed rapidly, and shared across multiple threat actors within days in a manner that had high detection signals. This may reflect a reduced cost and barrier to entry for this class of capability, as AI agents increasingly enable threat actor exploit development. This is particularly relevant for open source codebases, such as Chromium, where upstream patches are publicly accessible prior to downstream consumers of the codebase applying the patch. This creates a window for threat actors to attempt to rapidly reverse engineer patches and develop exploits ahead of downstream stable releases.” concludes the report. “The majority of observed BlueMoon usage is assessed to be China-aligned espionage-motivated activity, although there is not sufficient evidence to attribute BlueMoon usage exclusively to China-aligned threat actors at the time of writing.” 

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, BlueMoon)

  • ✇Security Affairs
  • US Agencies Warn Chinese AI Firms Are Extracting Advanced AI Models Pierluigi Paganini
    US agencies accuse six Chinese AI firms of extracting billions of tokens from US AI models to accelerate development and copy advanced capabilities. NSA, CISA, and the FBI jointly published an advisory accusing six Chinese AI companies, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, of running industrial-scale extraction campaigns against US frontier models since at least late 2024. The framing is deliberate: this isn’t a footnote to how these companies build AI, the agencies ca
     

US Agencies Warn Chinese AI Firms Are Extracting Advanced AI Models

9 de Setembro de 2026, 15:20

US agencies accuse six Chinese AI firms of extracting billions of tokens from US AI models to accelerate development and copy advanced capabilities.

NSA, CISA, and the FBI jointly published an advisory accusing six Chinese AI companies, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, of running industrial-scale extraction campaigns against US frontier models since at least late 2024. The framing is deliberate: this isn’t a footnote to how these companies build AI, the agencies call it the core of their entire development strategy.

Distillation is a legitimate and widely used technique. It involves training a smaller AI model to reproduce the answers and capabilities of a larger one. But the advisory says the activity it uncovered went much further. It alleges that the companies sent millions of requests to models such as Claude, GPT, Gemini, and Grok and extracted billions of tokens.

“China-based artificial intelligence (AI) companies are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies’ models through industrial-scale knowledge distillation campaigns that form the core—not merely a supplement—of their AI development strategy.” states the report. “Likely with Chinese government awareness, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024.”

The goal was to capture valuable capabilities, including reasoning, coding, and agentic skills that took years and huge amounts of computing power to develop.

The advisory provides the most detail about DeepSeek. It claims the company ran an organized campaign against different versions of Claude, GPT, and Gemini between late 2024 and mid-2025 to help develop its R1 and V3 models. The extracted data reportedly included specialized knowledge, such as legal expertise, as well as chain-of-thought reasoning.

“Advanced industrial-scale distillation tactics include chain-of-thought (CoT) reasoning extraction, automated failover between pathways during blocking attempts, and sophisticated quality evaluation frameworks to detect defensive countermeasures. China-based AI companies that conduct industrial-scale distillation against U.S.” continues the advisory. “AI models see significantly shorter AI development timelines and reduced financial expenditures in training a frontier model.”

This is particularly significant because AI companies usually limit access to a model’s internal reasoning. Getting a model to reveal those steps could therefore give attackers much more than just its final answers.

Moonshot AI’s alleged operation shows how fast these campaigns can move once new models ship. The advisory states the company redirected its extraction traffic to a newly released Claude model within 24 hours of launch, which only works if you already have infrastructure sitting ready and monitoring provider releases in real time. That’s not opportunistic scraping; that’s a standing operation built specifically to capture whatever comes out next.

The methods described in the advisory suggest a highly organized operation rather than researchers simply making API requests. The companies allegedly bought large numbers of premium accounts and shared them among teams of developers running many sessions at the same time. They also routed traffic through gray-market proxy services, which the advisory calls “transfer stations,” to remove identifying information and avoid detection. StepFun reportedly used pools of accounts and automated systems to spread requests across them, helping bypass rate limits and increase daily spending as the operation grew.

The advisory also describes attempts to manipulate the AI models themselves. MiniMax allegedly used prompt injection to convince Claude Code that it was actually a MiniMax product, hoping to make it behave differently. While this detail may sound unusual, it shows how far some of these efforts reportedly went to extract information from competing AI systems.

The advisory doesn’t just name and shame, it lays out concrete detection signals for US AI companies to watch for. Shared accounts logging in from multiple IPs and user agents, usage running 24/7 without the natural idle periods a human would produce, subscription-to-API-usage ratios that don’t add up, and brand-new accounts hitting maximum usage immediately instead of ramping up gradually the way legitimate adoption normally does.

We must consider that any one of those signals alone might be nothing, but the agencies are betting the combination is a fairly reliable tell.

“China-based AI companies leverage techniques not in MITRE ATLAS, demonstrating significant organizational investment, operational maturity, and adaptive capability development distinguishing these campaigns from opportunistic exploitation.” added the advisory.

The recommended countermeasures get genuinely aggressive, and one in particular is worth sitting with. The advisory suggests quietly serving degraded, less capable responses to accounts suspected of running distillation campaigns, without ever telling those users their access has been downgraded, specifically so they can’t adjust their extraction technique in response. That’s a notable policy stance from a government advisory: not just detect and block, but actively deceive suspected bad actors about the quality of what they’re receiving.

Whatever the geopolitical debate, the practical lesson for companies using frontier AI models is clear. If several employees share enterprise AI accounts, providers will likely monitor usage more closely for the patterns described in the advisory. Heavy legitimate use could sometimes trigger false positives, especially when organizations have many developers making large numbers of requests at the same time.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, AI Models)

  • ✇Security | CIO
  • Microsoft targets Salesforce customers with AI-powered Dynamics 365 migration tool
    Microsoft has introduced an AI-powered tool to help enterprises move from Salesforce to Dynamics 365 by potentially reducing some of the complexity that has traditionally made switching CRM platforms difficult, as the software giant looks to expand its CRM market share. Called Dynamics 365 Activate and currently in public preview, the tool analyzes an enterprise’s Salesforce data, processes, customizations and dependencies to identify what needs to be migrated, changed
     

Microsoft targets Salesforce customers with AI-powered Dynamics 365 migration tool

10 de Setembro de 2026, 09:23

Microsoft has introduced an AI-powered tool to help enterprises move from Salesforce to Dynamics 365 by potentially reducing some of the complexity that has traditionally made switching CRM platforms difficult, as the software giant looks to expand its CRM market share.

Called Dynamics 365 Activate and currently in public preview, the tool analyzes an enterprise’s Salesforce data, processes, customizations and dependencies to identify what needs to be migrated, changed or redesigned before implementation.

It then carries that context into solution design, configuration, migration, and testing, automating parts of the process, Microsoft explained in a blog post, adding that this approach replaces some of the manual discovery and analysis work typically required before and during a migration.

AI-powered tool could prompt CIOs to revisit CRM choices

That reduction in manual labor, time, and cost, analysts say, could give enterprises and their CIOs a reason to revisit CRM platform decisions.

“Earlier, companies often stayed with an existing platform because switching was too difficult. If AI significantly lowers migration costs and complexity, CIOs approaching major upgrades or renewals may have a stronger reason to revisit their platform strategy,” said Manoj Chandra Jha, principal analyst at Nord-IQ Research.

For enterprises that don’t have major upgrades or renewals coming up, the case for revisiting their existing CRM platforms could still become stronger, according to Pareekh Jain, principal analyst at Pareekh Consulting.

“Enterprise applications are moving from systems of record toward AI- and agent-driven systems of action, giving CIOs a reason to reconsider architectures that may have remained largely unchanged for 10 to 15 years,” Jain said.

“Primarily because these firms have likely accumulated messy customizations and outdated code in platforms such as Salesforce, creating technical debt they may need to address before deploying new AI tools effectively. Since they may have to do that cleanup anyway, Microsoft’s new tool could become an attractive alternative to simply fixing their existing Salesforce environments,” Jain added.

Microsoft eyes Salesforce’s vast CRM customer base

That opportunity to make enterprises reconsider their CRM choices, especially Salesforce, in essence, is what Microsoft might be after.

“Salesforce’s huge CRM install base is the target here. Even converting a small percentage of Salesforce customers could be significant for Microsoft,” Jain pointed out.

Moreover, the timing could also work in Microsoft’s favor, as investment banks reportedly have raised questions about Salesforce’s ability to convert its Agentforce push into meaningful customer adoption and revenue.

While a TD Cowen partner survey found that partners were not yet seeing meaningful Agentforce revenue, a separate KeyBanc survey report suggested that some CIOs could deprioritize Salesforce spending over the next 12 months due to the product’s maturity.

Taken together, those signals could challenge the perception of Salesforce as a safe, entrenched platform, according to Jha.

The gap between Agentforce’s AI narrative and the value customers and partners are seeing, combined with concerns over data readiness, product maturity, pricing changes and CIO budget sentiment, could give competitors, such as Microsoft, an opening to target its install base, Jha added.

Lowering migration friction doesn’t guarantee a switch

However, that opening doesn’t necessarily mean enterprises will begin moving from Salesforce to Dynamics 365 en masse, according to Jain.

“The strongest candidates are enterprises already using Microsoft 365, Azure, Power Platform, Fabric and Copilot, as well as companies facing high Salesforce costs, complex customizations or major contract renewals. Some enterprises may also move one business unit first rather than the entire company,” Jain said.

Even for those enterprises, however, an AI-based migration tool cannot remove all the challenges associated with a CRM migration.

“Poor data, complex integrations, custom business logic, regulatory requirements and organizational change will remain difficult to address. CIOs should therefore see Activate as a migration accelerator rather than a migration autopilot, Jain cautioned.

Further, CIOs should also be wary of treating Microsoft’s assessment as an unbiased verdict on whether switching platforms makes sense, according to Jha.

“Since it’s vendor-built, it’s optimized to make Dynamics 365 look like the easy answer, not to give an unbiased verdict on whether switching makes sense at all. Lower friction is a reason to look, not a reason to leap,” Jha said.

Microsoft expands Activate beyond migrations

Salesforce migrations, however, are only one part of Microsoft’s broader vision for Dynamics 365 Activate.

The tool, according to Microsoft, can also support greenfield implementations by helping enterprises turn business requirements into a Dynamics 365 solution when launching a new business model or a process.

Analysts, though, were split on the near-term adoption of Activate for such greenfield use cases.

While Jain sees greenfield projects as potentially becoming an even stronger use case because they lack the legacy complexity associated with migrations, Jha expects greenfield adoption to remain limited in the near term because the tool’s key advantage of legacy discovery has little application when no existing system is being replaced.

Activate, Microsoft further noted, could also help enterprises already using Dynamics 365 extend their existing environments as they enter new markets, add applications, or transform business processes.

In those cases, Activate is designed to build on the existing business and application context rather than requiring teams to begin each project with another lengthy discovery process, it said.

The company is also planning to expand Activate beyond Salesforce to support migrations from other CRM and ERP platforms, with ERP capabilities expected later this year.

  • ✇Security | CIO
  • OpenAI seeks tougher AI rules. CIOs may feel the ripple effects
    OpenAI is urging US lawmakers to impose mandatory safety requirements on developers of the most powerful AI systems, arguing that advances in AI are moving quickly enough that voluntary safeguards are no longer sufficient. The ChatGPT maker said in a statement that the rules should be based on what AI systems are capable of doing and should concentrate on a small number of well-resourced companies developing frontier models. It cautioned against extending the same requi
     

OpenAI seeks tougher AI rules. CIOs may feel the ripple effects

10 de Setembro de 2026, 07:11

OpenAI is urging US lawmakers to impose mandatory safety requirements on developers of the most powerful AI systems, arguing that advances in AI are moving quickly enough that voluntary safeguards are no longer sufficient.

The ChatGPT maker said in a statement that the rules should be based on what AI systems are capable of doing and should concentrate on a small number of well-resourced companies developing frontier models. It cautioned against extending the same requirements to startups and researchers whose systems operate well below that level.

OpenAI’s proposal calls for a federal framework that would require common testing and independent assessments of advanced models. It also wants clearer rules for reporting serious AI incidents and stronger cybersecurity protections around frontier development.

OpenAI tied its push for stronger safeguards to concerns that AI is beginning to accelerate parts of the research used to develop more capable systems. Fully autonomous recursive self-improvement, in which an AI system independently produces increasingly capable successors, is not happening today, OpenAI said. But AI agents can already perform some research tasks that would take skilled researchers several days, according to the company.

OpenAI said governments should establish ways to measure that progress and determine when development should be slowed or stopped if adequate safeguards cannot be maintained.

The company also endorsed four California AI safety bills. Gov. Gavin Newsom on Wednesday signed two of them, SB 813 and AB 1405, establishing frameworks for independent AI assessments and standards for AI auditors.

Alongside regulation, OpenAI wants frontier AI developers to adopt common monitoring practices, particularly as models gain greater autonomy and access to tools. It has also called for compatible international approaches as advanced models and AI expertise spread across borders.

What this could mean for CIOs

Under the approach OpenAI is advocating, the immediate regulatory burden would largely fall on frontier AI developers rather than their enterprise customers, according to Pareekh Jain, CEO of Pareekh Consulting. But CIOs could encounter downstream effects in how they govern the models they deploy.

“AI governance would increasingly resemble cybersecurity governance as enterprises would need to know which models are being used, what they can do, what data and systems they can access, and how much autonomy they have,” Jain said.

A common regulatory framework could make some aspects of enterprise AI governance more predictable by establishing consistent expectations for model developers.

Lian Jye Su, chief analyst at Omdia, compared the potential effect to technical standards in the telecom industry, where common requirements allow companies to work from a broadly shared framework. Independent safety assessments and standards for AI auditors could give CIOs greater confidence that vendors are being evaluated against more consistent criteria, he said.

Enterprises would still need to govern their own AI infrastructure, but greater standardization could make it easier to apply common practices across vendors and business units.

Higher-risk systems are also likely to demand stronger controls as they gain greater autonomy, Jain said.

Vendor oversight could become another pressure point. Contracts may need to account for changes to underlying models and give enterprise customers greater visibility into incidents that could affect their deployments. Jain said agreements should include audit rights, incident-notification provisions, and enough portability to make switching providers practical.

As AI becomes more deeply embedded in business operations, Su said CIOs should consider treating AI security as a dedicated program rather than simply adding it to existing technology governance.

That includes maintaining an up-to-date inventory of models and a registry of AI agents, classified according to their risk and business impact. Su also called for cross-functional AI governance bodies and defined human-oversight requirements for systems involved in high-stakes decisions.

CIOs may increasingly need to show not only that controls exist, but that they are consistently applied and documented. Charlie Dai, principal analyst at Forrester, said enterprises should expect greater emphasis on documenting how AI systems are classified, tested and monitored, particularly for higher-risk uses.

He said CIOs should also strengthen AI asset management and observability, while introducing red-teaming and validation testing for systems that carry greater business or regulatory risk. AI-specific incident-response playbooks would give organizations a defined process for handling model failures or other serious AI-related events.

Potential impact on the AI vendor landscape

OpenAI said a federal framework should address frontier risks without weakening competition or entrenching incumbents, arguing that public standards and independent verification could reduce the concentration of power now held by frontier laboratories.

The cost of complying with tougher safety requirements could favor companies with the resources to absorb them. Jain said large frontier-model developers such as OpenAI, Google and Anthropic would be better positioned to bear those costs.

Dai similarly said that capability-based regulation could reinforce the advantage of well-funded frontier labs because extensive testing, security, and reporting requirements would increase the cost of competing at that level.

Su, however, said that we need not leave enterprises dependent on a handful of proprietary providers. Open-weight and smaller models could remain part of a multi-model strategy if they satisfy applicable security requirements.

Even if smaller and open-weight models remain available, tougher enterprise governance and assurance standards could narrow the pool of suppliers CIOs are willing to use for sensitive workloads.

That possibility makes portability an architectural consideration for CIOs, particularly where enterprise applications depend on proprietary model APIs.

Jain recommended maintaining multi-model architectures and negotiating contracts that support portability, reducing the operational cost of switching providers if regulatory requirements or model capabilities change.

  • ✇Security | CIO
  • IT consulting has a big AI problem
    AI adoption is shaking up the big IT consulting market, with some IT leaders starting to question the need for the multi-year transformation engagements that have been large advisory firms’ bread and butter. Organizations are increasingly using AI tools to assist with large digital transformation projects such as cloud modernization and mainframe migrations, with the AI cutting the time and effort needed to achieve the goal. At the same time, the speed of evolution in t
     

IT consulting has a big AI problem

10 de Setembro de 2026, 07:01

AI adoption is shaking up the big IT consulting market, with some IT leaders starting to question the need for the multi-year transformation engagements that have been large advisory firms’ bread and butter.

Organizations are increasingly using AI tools to assist with large digital transformation projects such as cloud modernization and mainframe migrations, with the AI cutting the time and effort needed to achieve the goal. At the same time, the speed of evolution in the AI space has led some IT leaders to question the value of two- or three-year engagements.

Some observers have suggested advancements in AI will devastate the large IT consulting model, and McKinsey, Deloitte, Ernst & Young, and KPMG have all announced layoffs in recent months. Others suggest big IT advisory firms are here to stay, although they may need to adjust their business models to change with the times.

Representatives from Deloitte, EY, and Accenture didn’t respond, or declined to respond, to questions about the impact of AI on their businesses. But Edwin Miranda, founder of AI consulting firm konsultora, doesn’t think his larger competitors are going away anytime soon.

“I don’t think the large consulting firms simply disappear,” he says. “They still have enormous advantages in industry expertise, global delivery, complex integration, procurement, regulation, and the ability to operate inside very large organizations. What AI is attacking is the traditional operating model behind the engagement.”

In the past, a large IT consulting project would involve weeks or months of research, analysis, documentation, with large implementation teams doing the work, he notes. The cost of a large IT transformation came from the human labor required to move information from one stage of the engagement to the next.

“AI compresses a lot of that work,” he says. “Research that took weeks can happen in hours. Large document sets can be analyzed before the first meeting. Software can be prototyped while the operating model is still being discussed.”

In addition, AI agents can assist with development, testing, documentation, and coordination, and small senior teams can now produce a level of output that previously required a much larger pyramid, he notes.

“That doesn’t eliminate consulting; it changes what the client should be willing to pay for,” Miranda adds. “The value moves away from the volume of people assigned to an engagement and toward judgment, architecture, implementation, governance, and measurable business outcomes.”

The heart of the value proposition

Some observers see a dimmer outlook for the IT consulting industry. As AI adoption builds, the technology is eroding the core value proposition of big consulting firms, says Paul DeMott, CTO at digital marketing agency Helium SEO.

“Clients are running their own data analysis using off‑the‑shelf AI tools, so it is valid for them to ask why they should pay premium rates for work they can now do internally,” he adds. “AI commoditizes exactly what consultants charge a premium for, which is synthesizing data and translating insights into recommendations.”

With some big firms laying off staff, the traditional model of throwing junior consultants at problems seems to be breaking down, he says.

“Think about it, if AI can do the grunt work faster and cheaper, what exactly are those junior staff doing?” DeMott says.

DeMott sees some potential clients moving toward smaller, more specialized providers and others changing how they use large firms.

“Clients are not necessarily abandoning them outright, but they are pushing back on scope, duration, and cost,” he says. “The large firms still have relationships and brand equity, but those advantages are getting thinner by the quarter.”

Multi-year transformation on the ropes

In addition, big consulting firms are changing the way they hire because of AI, notes Brad Belzak, founder of AI-native advisory firm Acuity Global Partners. Until about two years ago, some generalist IT skills, a data engineering background, decent soft skills, or some trend analysis awareness would be enough to get a job at the large consultancies, he says.

“Frontier models ended that almost overnight,” adds Belzak, a former consulting employee at both Deloitte and EY. “First AI, then agentic AI, then highly trained specialized models absorbed the work that generalists with light coding skills used to do.”

Recruiters at the big consulting firms now want what he calls deployable talent, he says: engineers and analysts with product development histories who know how to read and manipulate data to solve client problems.

“Coding matters less than it did,” he adds. “The premium is on people who can take messy data, structure it, and turn it into an answer. The pyramid of generalists underneath them is gone.”

IT leaders at client companies, meanwhile, are demanding shorter, more modular engagements, Belzak says. Ninety-day sprints are becoming more common.

“The multi-year transformation made sense when the technology underneath it moved slowly,” he adds. “That world is gone. If your roadmap takes three years, the tools you scoped in month one are obsolete by month 18.”

IT leaders hiring consultants should focus on outcomes instead of headcount, he suggests.

“You’re not buying people anymore; you’re buying results,” he says. “That five-year contract might now be a six-month contract, and if the short-term outcomes generate wins, it gets extended.”

Consulting firms can still fill IT staffing needs, but their time on site may be shorter, he says. “Think embedded, on-demand engineers and analysts who come in, solve the problem, and move on, whether you’re a startup or a mature company,” he adds.

Accelerant and expertise

Aelin Golsarry, founder and CIO AAG Technology Consulting, doesn’t believe AI will kill large consulting firms, but it will make some of their services more difficult to justify. Still, AI can’t do all the work involved in large digital transformations, she says.

“AI should make a lot of the work required to get through a transformation faster, but it doesn’t make the transformation itself happen faster,” she adds. “People still have to make decisions, change processes, implement technology, and actually adopt it. AI doesn’t make any of that disappear.”

Golsarry, who as a CIO hired large consulting firms in the past, advises IT leaders to think clearly about what they’re buying when engaging with consultants.

“They should ask, ‘Do I need 30 consultants, or do I need three people who have seen this problem before and know how to fix it?’” she says. “I’d be looking at the expertise of the people actually doing the work, what outcome I’m paying for, and whether the firm’s use of AI is making the engagement faster and more efficient for me or simply making the engagement more profitable for them.”

Consulting isn’t going away, Golsarry adds. “Companies will always need expertise they don’t have internally,” she says. “What I think is going away is the assumption that more people, more hours, and a longer engagement somehow means you’re getting more value.”

  • ✇Security | CIO
  • What it really takes to be AI model independent
    Artificial intelligence is still the Wild West. Every organization adopting AI is, in a sense, operating on someone else’s ranch. Models, platforms and providers are evolving rapidly, and today’s market leader may not hold that position tomorrow. At its core, model independence recognizes that AI models are becoming interchangeable tools with different strengths, rather than technologies organizations should feel obligated to build around. The competitive advantage come
     

What it really takes to be AI model independent

10 de Setembro de 2026, 07:00

Artificial intelligence is still the Wild West. Every organization adopting AI is, in a sense, operating on someone else’s ranch.

Models, platforms and providers are evolving rapidly, and today’s market leader may not hold that position tomorrow. At its core, model independence recognizes that AI models are becoming interchangeable tools with different strengths, rather than technologies organizations should feel obligated to build around. The competitive advantage comes from matching the right capability to the right work at any given moment — not becoming attached to a single model or platform.

Rather than chasing every new release or trying to predict which provider will come out on top, CIOs should focus on building the capability to evaluate, route and adopt models as the technology changes. That starts with understanding how different models perform, knowing when to trust automated model selection and creating processes that can evolve as AI continues to change.

Start with the work, not the model

Model selection starts with a simple question: What am I trying to accomplish?

Every AI model is designed for different types of work, and not every task requires the same level of capability. Some models prioritize speed, while others are built for deeper reasoning. A simple factual question doesn’t require the same computing power as drafting a board memo, synthesizing several documents or optimizing a week’s worth of meetings to make the best use of an executive’s time. Many employees don’t realize those distinctions and will default to the best-known model regardless of what the work actually requires.

Anthropic’s Claude family illustrates this well. Haiku is designed to deliver quick responses to relatively straightforward requests. Sonnet balances speed and reasoning for many everyday business tasks, while Opus is intended for more complex analysis. Recognizing those differences allows organizations to match the right capability to the right work.

The same principle applies inside organizations. Leaders don’t assign every project to their most senior employee. They match the complexity of the work to the appropriate level of expertise. AI should be treated the same way.

This approach also has a direct impact on cost. Brown & Brown applies it in one of its own AI workflows. A lower-cost model orchestrates the agents that scan code to identify potential issues, while a more advanced reasoning model evaluates the highest-risk findings. Using Opus, for example, across every step of that process would be unnecessarily expensive. Instead, the organization gets the level of analysis it needs without paying for the most powerful model at every step.

Technology leaders should encourage teams to think in terms of capabilities rather than favorite models or platforms. Define the complexity of the work first; default to the least expensive model that meets the accuracy, quality and performance requirements, escalating only when additional reasoning or increased accuracy is needed. Just as importantly, evaluate success based on the accuracy and quality of the output rather than assumptions about which model should perform best.

As AI gets better at choosing models, people need to get better at judging results

One of the biggest changes in enterprise AI is happening behind the scenes.

Platforms like Microsoft Copilot and Claude feature “harnesses” that continually improve in their ability to evaluate a user’s request, determine how much reasoning it requires and automatically route it to the model best suited for the task without requiring the user to make every decision manually.

That doesn’t diminish the importance of understanding how different models behave. It changes where employees add value.

Rather than manually selecting a model for every request, employees need to recognize when the platform has made the right choice and when it hasn’t. Auto mode works well for many routine tasks, but it isn’t infallible. Users still need to evaluate whether the response meets the objective, determine when additional reasoning is warranted and recognize when the AI has misunderstood the request.

The same judgment applies as organizations build reusable prompts, AI skills and agents. For example, one of our employees learned this while using Claude to create a presentation slide. The instructions specified using a particular template, and the AI followed them exactly. The result technically met the request but produced a weaker slide than if it had been allowed to choose the format itself. The lesson was about recognizing when instructions — or assumptions — are limiting the quality of the output, not the model itself.

Organizations should also expect workflows to evolve. Model updates can change how AI responds, and accuracy, hallucinations, consistency and output quality still vary across models. Regularly comparing the same task across models, refining prompts and revisiting AI skills and agents help ensure the technology continues to produce the desired results.

As more of the model-selection process becomes automated, organizations should spend less time debating which model to use and more time developing employees who can evaluate AI output critically,recognize when intervention is needed and continually improve how AI is used.

Build an AI strategy that evolves with the technology

Periodically running the same workflow across multiple models allows technology leaders to compare output accuracy, quality, consistency, speed and cost and determine whether another model has become a better fit for a particular step in the process.

Those evaluations should extend beyond general-purpose foundation models. Industry-specific AI platforms, fine-tuned models and specialized SaaS providers may offer stronger performance for common business use cases because they have already configured model selection, data and workflows around a particular industry or function.

CIOs should look beyond the name of the foundation model and understand how vendors select, route and tweak models, how they evaluate new releases and how easily they can introduce another option. Model evaluation should remain an ongoing process, with critical workflows benchmarked regularly rather than only during the initial technology selection.

An AI roadmap should enable the adoption of stronger or more cost-effective models without rebuilding the workflows that depend on them.

Choose partners that think beyond today’s model

Most organizations — particularly small and midsized businesses — won’t build sophisticated model-routing systems themselves. They’ll rely on technology partners, SaaS providers and systems integrators instead.

When you’re evaluating an AI partner, know that the expertise behind the technology often matters as much as the technology itself. Some key ideas:

  • Look for partners that already support organizations larger than your own and ask how they’re approaching model independence, model selection and resiliency. Those capabilities shouldn’t be nice-to-haves; they should be requirements.
  • Don’t stop at asking which model they use. Ask what happens when that model changes — or when it’s unavailable.
  • If an AI provider experiences an outage or releases an update that affects performance, can the partner route work to another model and keep critical business processes running?

Their answers to the above may tell you far more about the resilience of your AI strategy than the name of the model they’re recommending.

  • ✇Security | CIO
  • Vibe coding is Topgolf. Production is Torrey Pines.
    Vibe coding is rapidly changing who can build software and how quickly an idea can become a functioning application. Using natural-language prompts and AI-assisted development tools, people can translate concepts into prototypes without mastering every element of traditional software engineering. The experience reminds me of Topgolf. Topgolf creates a carefully curated, technology-enabled environment where almost everyone can feel capable. The tee is automated. Perfor
     

Vibe coding is Topgolf. Production is Torrey Pines.

10 de Setembro de 2026, 06:30

Vibe coding is rapidly changing who can build software and how quickly an idea can become a functioning application. Using natural-language prompts and AI-assisted development tools, people can translate concepts into prototypes without mastering every element of traditional software engineering.

The experience reminds me of Topgolf.

Topgolf creates a carefully curated, technology-enabled environment where almost everyone can feel capable. The tee is automated. Performance is tracked. Wind, water, sand, and other hazards are largely virtual. You can experiment, compete, and receive immediate feedback — all from the comfort of a lounge-like setting.

Vibe coding can create a similar sense of confidence. You describe what you want, AI helps build it, and a working application begins to emerge. The demonstration succeeds, colleagues are impressed, and momentum builds.

Then comes the inevitable question: “How quickly can we put this into production?”

That is when the game moves from Topgolf to Torrey Pines.

On a championship course, the controlled environment disappears. You must account for shifting winds off the Pacific, hidden bunkers, changing pin placements, difficult terrain, spectators, and countless other variables. Success requires more than the ability to strike the ball. It requires course knowledge, preparation, situational awareness, discipline, and the ability to adjust when conditions change.

Production technology environments are no different.

An application that performs well in a controlled setting must now contend with scale, cybersecurity, data quality, privacy, integration, resilience, accessibility, regulatory requirements, technical debt, and unpredictable user behavior. The number of variables expands exponentially — and every variable can affect performance, trust, cost, and organizational reputation.

Across my career as a CIO and CTO, I have seen promising technology initiatives encounter difficulty not because the underlying idea lacked value, but because the organization underestimated what it would take to operate that capability reliably at scale. AI accelerates development, but it does not eliminate operational complexity.

That distinction matters because vibe coding is not simply another developer productivity tool. It changes who can participate in creation. A business leader who previously described a need and waited for a development queue can now sit beside an AI assistant and begin shaping the solution directly. That is a meaningful and welcome shift. It creates faster learning cycles, brings domain expertise closer to the product, and gives the business a more active role in technology delivery.

But democratizing development also democratizes responsibility. The person who can create a compelling application in an afternoon may not yet have the experience to recognize an insecure dependency, an ungoverned data source, a fragile integration, or a design that becomes prohibitively expensive at scale. A successful demonstration answers whether something can work. Production requires answering whether it should operate, how it will operate, who will own it, and what happens when it fails.

This is where experienced technology professionals remain essential.

A skilled enterprise architect is like having Tiger Woods help you read the course: someone with a mental model of the topology, dependencies, and conditions most likely to affect the outcome. The architect can see the dogleg beyond the tee box — the downstream consequences that are invisible during a prototype.

The course ranger resembles the change manager, helping teams navigate competing priorities, shared resources, release schedules, and operational conflicts. Cybersecurity professionals manage the equivalent of those pesky paparazzi trying to get inside the ropes, while also preparing for threats far more consequential than a distracting camera. Data teams act as the performance analysts, measuring behavior, outcomes, crowd reaction, and whether the solution is actually producing value. Platform and operations teams ensure the course remains playable when demand surges or something unexpected happens.

These professionals are not standing in the way of innovation. They are helping the organization play the real course successfully.

Before promoting a vibe-coded application into production, leaders should review a simple production-readiness scorecard:

This review does not need to become a six-month obstacle course. Governance should be proportionate to the application’s risk, reach, and potential impact. A low-risk internal productivity tool should not face the same controls as an application making consequential decisions about customers, employees, or vulnerable populations.

The answer is not to force every vibe-coded experiment through the heaviest enterprise process. That would squander much of the speed and creativity AI makes possible. Instead, organizations need risk-based pathways: a safe practice range for experimentation, a clearly marked route for internal tools, and a more rigorous qualification process for applications that touch sensitive data, critical operations, external users, or consequential decisions.

Technology leaders can make those pathways easier to navigate by providing approved AI tools, reusable components, secure development environments, automated testing, standard integration patterns, and transparent criteria for production readiness. In other words, build guardrails and paved roads rather than relying exclusively on gates. When the safest path is also the easiest path, governance becomes an accelerator.

The goal is disciplined acceleration.

The business also has to accept a different role. Vibe coding should not mean throwing an AI-generated application over the wall to technology once the demonstration is complete. If business teams help create these capabilities, they must remain hands-on collaborators and accountable owners. They understand the domain, intended outcomes, acceptable errors, and human consequences better than anyone. Technology brings the architecture, engineering, security, data, and operational disciplines required to make that capability sustainable. Production readiness is a shared responsibility.

This partnership represents an entirely new organizational muscle. Data and AI literacy provide the foundation, but literacy alone is not transformation. Leaders must learn how to evaluate AI-generated work, challenge confident outputs, understand where automation requires human judgment, and make informed tradeoffs among speed, risk, cost, and value. Technology teams, in turn, must learn to engage earlier and more collaboratively so they are not perceived as impediments arriving at the end of the process.

I am a cautious optimist about AI. Vibe coding can democratize development, accelerate experimentation, and bring business and technology teams closer together. It can also help reposition business professionals from passive recipients of technology to active participants in creating it.

But confidence developed in a controlled environment should not be confused with production readiness. The easier development becomes, the more important it is to strengthen organizational literacy about architecture, data, cybersecurity, scale, and responsible ownership.

The central leadership question is therefore not whether organizations should permit vibe coding. That debate will soon be overtaken by adoption. The better question is how organizations will preserve its creative energy while establishing the discipline necessary to earn trust at scale. Leaders who answer that question well will move faster because their teams will know where experimentation is encouraged, when additional expertise is needed, and what evidence is required before a capability reaches customers, employees, or mission-critical operations.

Perhaps we will eventually reach a world in which AI can immediately transform a vibe-coded idea into a secure, scalable, resilient production capability — making 18 holes at Torrey Pines feel as approachable as an evening at Topgolf.

We are not there yet.

Until we are, organizations must pair AI-enabled speed with sound system hygiene, architectural discipline, testing, governance, and risk management. The objective is not to slow innovation. It is to ensure that what we accelerate can survive — and succeeding on — the real course.

Together, we will get there responsibly.

  • ✇Security | CIO
  • Enterprises can’t spend their way to AI leadership
    A pathologically simple playbook emerged in the last few years for winning the AI race: hoard GPUs, hire every AI expert you can find and then watch the magic happen. But as we roll through the second half of 2026, cracks in that strategy have turned into craters. A harsh reality of frontier AI development is finally setting in: you can’t spend your way to the top. Building a world-class AI system requires deep institutional structures that drive enterprise-wide adoptio
     

Enterprises can’t spend their way to AI leadership

10 de Setembro de 2026, 06:00

A pathologically simple playbook emerged in the last few years for winning the AI race: hoard GPUs, hire every AI expert you can find and then watch the magic happen. But as we roll through the second half of 2026, cracks in that strategy have turned into craters. A harsh reality of frontier AI development is finally setting in: you can’t spend your way to the top.

Building a world-class AI system requires deep institutional structures that drive enterprise-wide adoption. It involves cultivating and investing in a tightly aligned engineering culture. It requires the kind of relationships that attract and, vitally, retain the absolute elite.

The compute mirage and the bending demand curve

AI spending has grown to truly unprecedented levels in the past 18 months. The top five tech giants alone are projected to spend over $750 billion combined in 2026 for AI infrastructure. They followed the playbook by buying the chips, generating the power and building the infrastructure. But they did so operating on a core industry assumption: that the demand for massive, monolithic frontier compute would scale exponentially forever.

While plausible, it’s clear the demand curve is starting to bend.

While companies were stockpiling silicon, the open-source community and Chinese AI labs quietly changed the math. Competitors discovered that you don’t need to spend a billion dollars training a frontier model from scratch when you can use model distillation to train smaller, highly efficient models on the cheap.

How they did it: Chinese labs like DeepSeek, Zhipu, Moonshot and MiniMax have aggressively leveraged distillation and open-source foundations. Despite U.S. export controls severely limiting their compute, Chinese models are aggressively closing the gap with U.S. frontier models on key benchmarks at a fraction of the cost.

They also introduced open-weight models with similar levels of quality at little to no cost at all. In a short period of time, open models have become the de facto standard for startups and enterprises. When a distilled, open-weight model can achieve 90% of a frontier model’s performance on a specialized task, the justification for paying massive API fees to a centralized provider vanishes.

So, what does this mean?

Well, when you spend hundreds of billions expecting a monopoly on intelligence, only to find that competitors can essentially pirate your capabilities for pennies, your entire business model evaporates. Companies like Meta and SpaceXAI are experiencing this shortfall firsthand. SpaceXAI is now renting out spare capacity to Anthropic, and Meta is floating “Meta Compute” to sell off its excess power.

In other words, the infrastructure that was supposed to be a weapon has become an expensive anchor, turning companies into digital landlords for their competitors.

Turmoil tax: Why top talent is walking away

Hype cycles in the early 2020s produced a flood of newly minted graduates wielding advanced degrees in machine learning. But as the industry matured, a painful truth emerged: understanding AI theory is common. However, knowing how to build, stabilize and scale a frontier system from scratch is incredibly rare.

Training a trillion-parameter model is a distributed systems nightmare. You can’t just throw a hundred fresh PhDs at a massive cluster and expect a frontier model to pop out. You need a team of deeply experienced engineers who grok the theory while also understanding catastrophic failure points, hardware-software co-design and network optimization.

And right now, the industry is actively driving that talent away. Across major tech giants and leading AI labs, a brutal pattern has emerged: sweeping layoffs executed specifically to free up capital for massive AI infrastructure bills. Engineers are effectively being sacrificed to buy more compute. This astronomical cash burn is creating chaotic, high-pressure environments where shifting goalposts and constant team resets, have triggered a growing exodus of key staff.

This turmoil creates a vicious cycle. Elite AI engineers, the true “10x” talent that actually knows how to string 100,000 GPUs together without the system crashing, want stability, clear mandates and a culture that values their institutional knowledge. When a company signals that it views human talent as a highly expendable line-item, top talent flees to more stable, culturally aligned labs. It’s impossible to build a generational product when your core engineering team has a revolving door.

The public shift: Privacy, cost and fatigue

A fundamental shift in the public and enterprise appetite for AI compounds this pressure. I’ve seen it first-hand.

Back in 2024, companies were willing to pipe their proprietary data into massive, closed models just to see what would happen. In 2026, the honeymoon is over. The public and corporate sectors are increasingly concerned about data privacy and the staggering costs of operating massive frontier models at scale.

Enterprises are realizing they don’t need a multi-trillion parameter model that “knows everything” to summarize internal legal documents or write code. They want smaller, localized, cheaper models that can guarantee their data privacy. This public shift toward cost-efficiency and privacy heavily favors the open-source and distilled models over the massive, costly API walls built by the biggest spenders.

The takeaway

Big Tech is learning the hard way that scaling an AI lab is like scaling a space program.

The moat in artificial intelligence is institutional rather than financial. You need unglamorous, highly disciplined systems engineers who stick around for years, compounding their knowledge of the company’s specific infrastructure. You need a deeply rooted engineering culture that gives researchers the stability to execute. And you need a business model that aligns with where the market is actually going, rather than where you hope it will be.

Hoarding all the compute in the world only gets you so far if your talent is fleeing. Prospects look even worse if your competitors are distilling your models and your customers are demanding cheaper, localized alternatives. Instead of building on the frontier, you’re left building a very expensive data resort. But it’s not too late.

  • ✇Security | CIO
  • Anthropic maps three AI futures for 2030; the most extreme could upend the economy
    AI is evolving faster than most people, even those building it, could even fathom, and its impact on the workforce and the economy is, at this point, really anyone’s guess. Researchers from The Anthropic Institute are offering a few possibilities: They have built a nuanced framework looking at how AI might impact jobs, unemployment, and gross domestic product (GDP) growth between now and 2030. They posit three potential scenarios for an AI-augmented future: “modest,”
     

Anthropic maps three AI futures for 2030; the most extreme could upend the economy

9 de Setembro de 2026, 23:23

AI is evolving faster than most people, even those building it, could even fathom, and its impact on the workforce and the economy is, at this point, really anyone’s guess.

Researchers from The Anthropic Institute are offering a few possibilities: They have built a nuanced framework looking at how AI might impact jobs, unemployment, and gross domestic product (GDP) growth between now and 2030.

They posit three potential scenarios for an AI-augmented future: “modest,” “substantial,” and “extreme,” and have created an interactive tool where users can explore how productive, or disruptive, AI will become in the workplace, based on their predictions of how they will work in 2030.

“Which of these worlds we are heading toward may become clearer within a year or two, and preparing for potential disruption seems to us the prudent course,” the researchers noted.

The goal of their work is to inform debate as AI becomes more powerful and capable. “AI is likely to reshape the US and global economies in profound ways in the coming decade, but how, and by how much, is extraordinarily uncertain,” they wrote.

How different scenarios could play out

If you add up every single task performed by people, machines, and software, the US has created a staggering $30 trillion in value over just the last year, the Anthropic researchers estimated. Their model and the corresponding tool are a way to explore how AI impacts tasks that contribute to the economy, the tasks it augments and creates, impacts on productivity, and speed of adoption.

“The answers to these questions have direct effects on GDP, the labor market, and the share of the pie taken home by workers,” they wrote.

Under their definition of “modest” change, AI will add less than half a point to GDP by 2030, meaning it will increase the growth rate of the national economy by just 0.5%, and will raise unemployment by just a tenth of a point, a minor shift. In this future, it’s difficult to see AI’s impact in macroeconomic data; change is steady but gradual, similar to that of the internet. “It drives real economic gains, but they’re within the historical norm for new technologies,” the researchers noted.

In the “substantial” scenario, AI will be capable of doing half of all knowledge work by 2030, the majority of it autonomously. Still, it wouldn’t be adopted for all work; in fact, most knowledge work tasks would still be completed without AI. Correspondingly, the economy would grow at twice its normal rate, but even as some non-knowledge workers see gains, wages for knowledge workers wouldn’t rise.

In this case, “AI makes a bigger impact than the internet, or the railroad,” the researchers wrote. Reallocation could be costly, but it is in line with what the US labor market has historically absorbed.

In the “extreme” scenario, of course, AI would be more productive than humans on the majority of knowledge work tasks, would do all of them autonomously, and subsequently would create no new knowledge tasks for humans.

The technology would “drive a completely transformed, unprecedented economy” arising from recursively self-improving AI. GDP growth would rise to 15% per year, but nearly one in five cognitive workers would be unemployed, and their relative wage would fall “immensely.”

The conundrum is that resources to compensate unemployed or under-paid workers will exist, but it’s unclear whether they would be fairly allocated. Mechanisms by which people can benefit from a much richer economy (retraining, income support, or universal basic income, for example) would become a question of economic policy.

“Whether and how those resources reach the people who bear the cost is not something growth delivers by itself,” the researchers wrote.

What users think

As well as developing the framework, the Anthropic researchers conducted a survey among roughly 11,000 Americans, asking them to predict AI use, productivity gains, automation versus augmentation, and displaced work.

They found that, in the main, public expectations land around the “substantial” scenario. That is, GDP would be 10% higher by 2030 than it would be without AI, and the overall unemployment rate would rise to around 5%.

Roughly 10% of respondents, on the other hand, had views in line with the “extreme” scenario.

Anyone can generate their own forecast using the researchers’ interactive tool, answering questions like: “Out of every 100 instances of a task AI can do in 2030, how many will AI actually be doing?”, “How many will be fully automated?”, or  “How much more gets done in an hour in 2030, compared with doing the tasks without AI?” The tool then responds, mapping their predictions to one of the three scenarios.

“Ultimately, what the economy looks like in 2030 depends on many factors, like what AI can do, and how companies and workers choose to adopt it,” the researchers wrote. “It also depends on how the financial benefit of this technology is shared.”

The between-the-lines reality

Sanchit Vir Gogia, chief analyst at Greyhound Research, emphasized that the Anthropic research “maps the conditions under which very different futures appear, it does not schedule destiny.”

He sees the distribution result, rather than the unemployment result, as the serious finding. In the extreme case, GDP is 32.4% above the no AI path, and the cognitive wage bill is 31% below it. Labor’s share of income falls from 60% to 45.2%, and capital income rises 81.4 %. That means a full 15% of GDP is captured as ROI rather than being paid out in labor costs.

In other words, he pointed out: “A richer economy is not automatically a fairer one.” Capability, diffusion, productivity, automation, and occupational friction all have to arrive together.

“AI will touch a large and rising share of knowledge work and will execute a much smaller share under independent authority,” he said. There is no single honest adoption percentage, because worker use, company use, technical exposure, and executed task instances are four different measurements.

Lessons from the research

Enterprises can take important lessons from the research as they deploy AI and consider its impact on their systems, workflows, and workforce, Gogia said.

“For enterprises, the binding variable is permission to delegate,” he noted. “A model that can draft a payment instruction is not thereby permitted to move money.”

His firm identifies five recurring concerns that come up in enterprise conversations: Durable returns after the full cost of deployment, control over authority being granted, augmentation quietly becoming substitution, erosion of professional formation, and fairness of how gains and risks land.

Some of those changes are progressing faster than the governance around them, he observed. Once a system can inspect customer data, change configurations, or act on workforce records, autonomy has stopped being a feature and has instead become an allocation of institutional authority.

“And the tasks easiest to automate are frequently the tasks through which judgement is learned,” he noted.

This article originally appeared on Computerworld.

Layoff remorse: Gartner says at least one in three positions eliminated by AI will be restored by 2029–at a higher cost

9 de Setembro de 2026, 22:32

Gartner on Wednesday said that it expects 30% of the positions eliminated by AI-related layoffs to be refilled by 2029, suggesting that the initial terminations were ill-advised and excessive.

“When business and IT executives look back on the early AI era, they will realize their greatest mistake was believing that work automation was the point, when workforce amplification was the opportunity,” said Tori Paulman, VP analyst at Gartner. “The competitive advantage will go to the CIOs and business executives who build an AI-shaped organization where AI value compounds by reshaping roles and allowing workflows to cross traditional boundaries, increasing velocity and reducing friction.”  

The Gartner report noted that it is finding that the cuts “deplete talent pipelines and erode institutional knowledge.” Beyond the immediate workforce disruptions associated with any mass layoff, companies will also face steep increases in costs for recruitment, training, and onboarding.

It also predicted that, by 2027, “75% of organizations that prioritize capturing AI productivity gains as cost savings will be eclipsed by competitors that aggressively reinvest those gains into innovation, modernization and upskilling.”

In an interview with Computerworld, Paulman said that the 30% figure represents the average impact on organizations of all sizes; they estimate that the layoff boomerang for enterprises would be even higher, roughly 40%. 

Paulman said that Gartner’s research found a lot of what they called “AI washing” by executives who want/need to do layoffs for purely budgetary reasons, and will falsely blame AI for the reductions because it makes them look better.

“More than 50% of our enterprise clients have been given a number [by their bosses],” Paulman said, and have been told by senior management to find that percentage of savings from AI.

But despite widespread evidence of problems due to AI-related layoffs, such job cuts are still increasing

Layoffs were ‘excessive’

Other analysts and consultants agreed with the Gartner suggestion that many of these job losses attributed to AI are going to be walked back, but questioned the specific statistic. Some also noted that 70% of the AI-attributed layoffs may remain in force, which would suggest that the original terminations were mostly justified. 

However, Frank Dickson, principal analyst at Dickson Research, argued that a lot of the layoff reversals will occur in a variety of ways that will obscure the fact that they are restoring a terminated role. 

“A lot of that 70% never shows up as a clean rehire even when the original cut was wrong,” he said, pointing out that some of the losses caused service to quietly get worse, and stay poor, some of the work was contracted out or offshored, some of the roles were reconstituted with a different position or title, and some was covered by the remaining staff absorbing the load. This,” he noted, “shows up later as burnout and attrition, not as a line item on this report. None of that gets counted in the 30%, and none of it is evidence the original call was sound.”

Melody Brue, principal analyst for Moor Insights & Strategy, added that the 70% scenario “could show that a substantial share of the AI-related workforce reductions is durable,” but, she stressed, “it shouldn’t be mistaken for endorsement of how those layoffs were made. What it doesn’t show is whether the organization captured the full economic value it expected. A lower headcount is not by itself evidence of a successful AI transformation.”

Valence Howden, advisory fellow at Info-Tech Research Group, questioned the methodology behind the calculation of Gartner’s 30% figure, but he agreed with the overall sentiment that layoffs attributed to AI have been excessive.

“I’m not sure we can substantiate those numbers, since it’s much more of a guesswork statement than anything else,” he said. “I do believe the current trend is going to lead to rehiring, especially as AI governance requirements ramp up and given AI’s lack of contextual semantic understanding. We know AI has not provided the value proposition that it has been sold as providing, and unless costs are controlled, it will be cheaper to use humans to perform some of the advanced work.”

Supporting data

Dickson also raised questions about the Gartner report because it lacked comparative layoff statistics. 

“Gartner doesn’t say what the reversal rate looks like for ordinary layoffs, the ones that have nothing to do with AI,” he said. “Suppose normal cuts get walked back at 10% to 15% in a typical five-year window, which is plausible given ordinary churn and business-cycle rehiring. A 30% rate specific to AI-driven layoffs would still run well above that, and that’s a damning number. Without that comparison, 30% is just a figure floating with no anchor.”

However, Dickson pointed to various datapoints supporting the position that AI layoffs have been excessive, noting that Forrester reported that 55% of businesses “already regret AI-driven cuts and are predicting half of those layoffs get quietly reversed.” 

“Robert Half puts it at a third of hiring executives who eliminated roles for AI having already rehired. Ford, IBM, Booz Allen Hamilton, Alphabet and CSX have all walked back cuts or announced rehiring drives,” Dickson said. “Gartner’s 30% by 2029 sits comfortably inside that range.” Klarna has also walked back AI layoffs. 

A ‘major indictment’

He added that many AI layoffs amounted to a corporate version of a crash diet. “You cut fast, you look great on the next earnings call, and eighteen months later, the weight is back, plus interest, because nobody fixed why the cut was made in the first place.”

Gartner’s Paulman agreed, noting, “business and IT executives who use AI primarily as a tool for cost cutting risk making reductions that are too deep and too soon, affecting their ability to innovate their business model and compete in new markets as AI continues to mature.”

Mike Wilkes, enterprise CISO at Aikido Security, said that even if the 30% figure turns out to be accurate, it is a major indictment of the layoffs. 

“If 30% of AI-driven layoffs must be reversed, that is an enormous error rate for a strategic workforce decision,” Wilkes said. “Imagine any other major capital decision where nearly one-third had to be unwound at a premium three years later. No CFO would call that a strong outcome.”

This article originally appeared on Computerworld.

San Francisco Orders Meta to Stop ‘Allowing’ AI Child Abuse Ads

9 de Setembro de 2026, 18:15
The City Attorney’s Office has asked Meta to explain how the harmful ads repeatedly ran on Facebook and Instagram. The company claims the ads are not under the city’s jurisdiction.

Ontem — 9 de Setembro de 2026Stream principal
  • ✇Cybersecurity News
  • Microsoft Unveils Unmetered Intelligence for Windows 11 Do Son
    Discover Microsoft's unmetered intelligence vision. Learn how local AI agents on Windows 11 PCs will replace cloud reliance and token economics. Related Posts: Windows 11 Is Finally Getting Automatic Light and Dark Theme Switching Windows 11 Performance Enhancements Target 8GB RAM Devices Windows 11 Battery Widget: New Connected Device Status The post Microsoft Unveils Unmetered Intelligence for Windows 11 appeared first on Daily CyberSecurity.
     
  • ✇Security | CIO
  • DealHub MCP Brings Agentic Control to Quote-to-Revenue
    Automates revenue system management in alignment with corporate governance and business policies DealHub AI, the leading Agentic Quote-to-Revenue platform, today announced MCP for Admin, a new AI capability that transforms the way organizations configure and manage their revenue systems. Through autonomous workflows and natural language prompts, MCP for Admin enables organizations to implement business changes faster and with greater intelligence, while ensuring changes re
     

DealHub MCP Brings Agentic Control to Quote-to-Revenue

9 de Setembro de 2026, 08:14

Automates revenue system management in alignment with corporate governance and business policies

DealHub AI, the leading Agentic Quote-to-Revenue platform, today announced MCP for Admin, a new AI capability that transforms the way organizations configure and manage their revenue systems. Through autonomous workflows and natural language prompts, MCP for Admin enables organizations to implement business changes faster and with greater intelligence, while ensuring changes remain aligned with established business context, governance, and corporate policies.

“Agentic Quote-to-Revenue is moving beyond assisting sales users. It is fundamentally changing how revenue teams operate and manage their business systems,” said Eyal Elbahary, Co-Founder and CEO of DealHub AI. “MCP for Admin introduces an agentic operating model that combines intelligent automation with the control, governance, and business context needed to operate their revenue systems at scale.”

Revenue teams have largely focused their AI investments on automating workflows that support sales motions. MCP for Admin extends that agentic transformation to how revenue teams configure and administer the Quote-to-Revenue system that powers these workflows. With MCP for Admin, teams can create and modify pricing rules, approval workflows, guided selling flows and configuration guardrails for their Quote-to-Revenue environment – from prompt-based commands for basic updates to automated agentic workflows that execute complete administrative processes.

MCP for Admin’s pre-configured skills and best practices provide built-in controls for how changes are executed, ensuring they remain consistent with the organization’s established business logic, governance, and corporate policies. This enables administrators across all skill levels to manage everything from routine changes to sophisticated configurations with greater speed, consistency, and confidence.

DealHub MCP for Admin is available to all customers in October 2026.

About DealHub AI

DealHub AI is the Agentic Quote-to-Revenue platform for the AI era – built to design, launch, and scale any monetization model – SLG, PLG, self-serve, subscriptions, usage, AI consumption. The platform consolidates CPQ, CLM, Subscription Management, Billing, Revenue Recognition, DealRoom, and composable API-first headless quoting into an AI-driven, orchestrated revenue backbone. 

For more information, users can visit dealhub.ai or follow DealHub AI on LinkedIn.

Contact

CMO

Gideon Thomas

gideon.thomas@dealhub.io

  • ✇Security | CIO
  • AI builds faster than organizations can govern. How can CIOs catch up?
    Organizations are racing to deploy AI, but warning signs are accumulating. Earlier this year, an internal AI agent gave an engineer instructions that exposed sensitive user and company data for two hours. Around the same time, a large online retailer issued a 90-day safety reset after its AI assistant contributed to an incident that involved nearly 120,000 lost orders. And in the spring, an AI agent deleted a company’s production database and its volume-level backups in ni
     

AI builds faster than organizations can govern. How can CIOs catch up?

9 de Setembro de 2026, 07:00

Organizations are racing to deploy AI, but warning signs are accumulating. Earlier this year, an internal AI agent gave an engineer instructions that exposed sensitive user and company data for two hours. Around the same time, a large online retailer issued a 90-day safety reset after its AI assistant contributed to an incident that involved nearly 120,000 lost orders. And in the spring, an AI agent deleted a company’s production database and its volume-level backups in nine seconds.

Over recent years, recurring events like these, among others, expose a widening gap between what AI can do and what organizations can safely control.

“A year ago, most conversations were about accelerating AI adoption as fast as possible,” says Sandeep Johri, CEO at application security platform Checkmarx. “Today, boards ask tougher questions. Speed and governance have to move together now.”

As Johri points out, the new bottleneck is the organization’s ability to govern AI. According to IBM’s 2026 Tech Leader Study, 77% of organizations admit their governance is failing to keep pace with AI. And, among IT executives, 70% say business teams are deploying tech faster than it can be tracked.

The use of agentic AI only widens the gap. About 80% of the organizations surveyed say they lack mature capabilities for it, according to Deloitte. That includes clear boundaries for agents, real-time monitoring systems, and audit trails that can capture the entire chain of actions.

CIOs need to operate in this paradigm to address two competing demands: accelerate AI adoption to boost productivity and outsmart competitors, and assure boards that all sensitive data is protected and AI only does what it’s supposed to do.

“I don’t think you can separate the two,” says Sahil Sanghvi, VP of AI engineering in the chief technology office at Booz Allen Hamilton.

Innovating while managing risks

At first glance, AI-generated code can look good and even pass initial testing. A thorough review, however, can shed light on multiple issues. This is something Ha Hoang, CIO at data protection platform Commvault, witnessed firsthand.

In one case, her team found the AI had taken a shortcut. It bypassed the company’s authentication process in favor of a simplified implementation, which lacked established access controls. “Without those checkpoints, it could’ve made its way much further,” she says.

When companies discover major issues, they should immediately pause deployment. But many problems aren’t obvious. “AI-driven risks often remain hidden, and moving too quickly only makes those silent failures harder to detect,” says Omer Cohen, CISO at customer identity and authentication service Descope.

But to strictly move slowly everywhere isn’t an option either. The idea is to identify where speed creates value, and where the potential consequences call for caution, and then build necessary guardrails case by case.

For Bob Leek, CIO at Clark County, Nevada, that means making governance and compliance part of the design, not a final check before deployment. “We’ll go slow to go far instead of going fast and creating risks,” he says.

The biggest challenge is organizational, not technical

In many cases, AI deployment is less a technology problem than a people problem. When deciding what to automate inside an organization and how to do it, the real challenge is understanding how work actually gets done. And usually there are many invisible, undocumented processes that influence it.

Employees in HR, finance, procurement, legal, or operations rely on exceptions every day. They have workarounds and make judgment calls to keep the organization running. These tweaks are simply part of the job, so they rarely think about them or include them in official process documentation.

These elusive workflows can’t be mapped simply by considering how things are supposed to work. Leaders must closely observe how employees actually do their jobs.

“Frontline teams understand the exceptions, escalation paths, and context that rarely appear in a process map,” says Leek. “We bring those teams into the design process, mapping the handoffs and non-standard cases.”

Cohen agrees. “Invisible threads are often fragments of context residing in an individual’s mind rather than a database,” he says. For instance, an analyst may know that a client’s login spike is harmless because it’s scheduled during weekly testing. “Unless this tribal knowledge is codified as a formal governance artifact via runbooks, threat models, or decision logs, no AI will naturally possess it,” he adds.

But simply asking employees how they work isn’t enough, adds Amitkumar Rathi, chief product and technology officer at hybrid infrastructure observability platform Virtana. The best approach is to run shadow sessions, in which someone in tech actually witnesses how the work is done. “We sit next to them during live incidents and ask, for instance, why did you look at that dashboard and not this one; why escalate now and not 10 minutes ago; what told you this was the same issue as last month’s incident and not a new one?” he says.

Of course, mapping informal processes takes time and discipline, and there shouldn’t be any tempting shortcuts. “The organizations that get this right treat AI as a collaborator in their existing workflows, not a replacement,” says Vijay Jegan, chief AI transformation officer at enterprise customer retention platform Gainsight. “Success requires a hybrid of deep business acumen within a department and the technical maturity to understand the inherent risks of modern AI tools.”

But not all tribal knowledge can or should be documented. “The goal should be to architect AI to augment this human foundation, rather than attempt to replace it entirely,” adds Cohen.

Where should humans stay in the loop

Giving AI a larger role makes human judgment more important, not less. “Humans should stay in the loop in every decision, but not every part of the process,” says Leek. “The urgency to innovate doesn’t change that fundamental responsibility.”

CIOs can decide where people should remain involved by weighing the value of human judgment and the risk of leaving the task entirely to AI. Tasks that score highly on both should remain firmly in human hands. “The higher the risk, the more human oversight is required,” Jegan says.

Sanghvi also factors in human consequences of potential AI mistakes. “When you deal with a decision that could materially affect a person, a mission, or an organization, that’s where you want clear human authority to intervene or override the system,” he says. “As AI becomes more agentic and starts taking actions rather than just making recommendations, being clear about those boundaries becomes even more important.”

Meanwhile, Cohen draws the line at AI-powered decisions that can’t easily be undone. “Human intervention remains non-negotiable at any juncture where a decision becomes irreversible or traverses a critical trust boundary,” he says.

At the other end of the spectrum, routine, low-risk work can be left to the machine. “Organizations may trust agents to autonomously handle narrow, repeatable tasks,” says Hoang, adding, though, that even advanced agents can misinterpret context or take unintended actions at scale.

“The future isn’t blind trust but measurable trust built on transparency and control,” she says.

Governance doesn’t end at launch

Before an AI initiative becomes a major commitment, Leek recommends CIOs ask if the project supports the organization’s strategic priorities, if IT can support it, and does the business department have the capability and appetite to change?

“This framework helps prevent initiatives from becoming solutions in search of a problem,” he says. It also helps CIOs start with lower-risk projects, test what works, and strengthen governance before applying AI in more sensitive areas of the organization.

Clark County took that approach with its first AI deployment for special-event permitting. Its AI tool guides promoter through forms, identifies the permits needed, and connects them with a county analyst. But starting with a lower-risk project doesn’t mean the governance work ends at launch. Governance should be a continuous conversation rather than a checkpoint, says Sanghvi, since data changes and models evolve.

Hoang agrees. “If your governance system relies on quarterly reviews, you’re already behind,” she says.

  • ✇Security | CIO
  • The need to fortify cloud integrity as cracks increase
    Over the course of his career, Jim Reavis has seen cloud and cloud security evolve, and it’s come a long way since being a niche technology in the early 2000s. Now it’s dominant in terms of being the IT foundation, he says, but while the tech is strong, the operating models is where things get messy. Cloud, security, and third-party risk teams look at different parts of the problem, of course, but challenges remain. “Operational technology worries me a great deal,” he says
     

The need to fortify cloud integrity as cracks increase

9 de Setembro de 2026, 07:00

Over the course of his career, Jim Reavis has seen cloud and cloud security evolve, and it’s come a long way since being a niche technology in the early 2000s. Now it’s dominant in terms of being the IT foundation, he says, but while the tech is strong, the operating models is where things get messy. Cloud, security, and third-party risk teams look at different parts of the problem, of course, but challenges remain.

“Operational technology worries me a great deal,” he says. “A lot of those systems are isolated and not kept up to date. If we don’t modernize them, we’re going to have huge problems. In a lot of cases, things fall between the cracks and that’s where hackers like to exist.”

So much of what’s around the models is where cybersecurity has responsibility, rather than the provider covering everything. “Data, identity, and applications are shared responsibility areas, and in many cases, the tenant carries most of the control burden,” Reavis says. “If you use a hyperscaler, you may still have about 80% of the responsibility for the controls around what you build.”

And when it comes to AI, the model isn’t the whole problem. What matters is the context around it, the goals it’s given, and the oversight put in place, he says. “We need to think carefully about the harnesses we put around AI and the systems we use,” he adds.

The responsibility model, therefore, is a recurring issue in cloud security breaches tied to misconfiguration and accountability gaps, and some enterprises still aren’t clear about where responsibility begins and ends. “We spent a lot of time on a shared security responsibility model, but when this first started to gain popularity, there were a lot of organizations or SaaS providers you could work with who’d say it’s in the cloud, it’s at Amazon,” he says. “Look at their certifications and SOC2 and how they comply because they’re covering everything.” But when you look at the actual applications, data, and identity, he adds, there’s so much that’s shared responsibility, and the customer’s responsibility.

So how do we make sure information is encrypted properly so it doesn’t become a tenant issue? “There’s still a bit to do, and we think about this not only from whether it’s SaaS, infrastructure, or a particular provider, but at what level is it at the physical, network, or audit level,” he says. “And even from a role-based perspective, what’s the role of internal risk and role of providers?”

Reavis gives further detail about how AI adoption exposes weaknesses in identity, trust, and risk management, and the long-term implications of increasingly interconnected cloud ecosystems. Watch the full video below for more insights, and be sure to subscribe to the monthly Center Stage newsletter by clicking here.

On cloud risk management: When we had the Chat GPT moment, we knew it because AI had been around for a while, but that was a cloud delivered version of AI to the masses, so we saw this going to evolve and you could see it combining in many important areas.

But what we’ve learned is, because this is an interesting predictive rather than deterministic technology, we’re living in a world of two exponentials, and you’re seeing model capabilities growing so quickly. There’s this feeling from a security perspective that we have to look to the model itself and fix every hallucination and everything else when that’s built into how it works. It’s actually working as intended. So that’s a new lesson. Models are going to get more powerful, but it’s so much of what’s around the models where cybersecurity has responsibility, and we don’t rely on frontier model companies or using open-weight models. Rather, what’s the context, oversight, and information we’re providing them, what do we do in terms of goals we give them, and what are the harnesses we put around AI and the models we deal with?

These are going to be the big areas to think about, but we have to understand the parts we can control. We’ve got to think carefully about the harnesses, transparency, and using supply chain shared responsibility. SaaS and cloud providers are all AI enabled now. You’re not using any software of any significance that isn’t using AI to some degree.

On AI identity, trust, and control: One of the areas that we’ve championed is zero trust as a philosophy. It was initially more of a networking type of approach at the network layer, or an idea that you use identity to understand network access. But it’s evolved more to an idea that anything can be breached, so you assume that. Then you think about how to make systems resilient, and build up confidence and protection.

So zero trust tells us that with human identity, we can ask what our digital identity is, and now we’re in a very interesting area for identity management and associating that with agents and AI systems. People might have just one view of it, but agents are as diverse as humans. So we think about different identities and least privilege, and how to prevent them from escalating privileges. We need to introduce new concepts like least autonomy, and think about an agent that has certain tasks and use identity to make sure the actions it takes are within a defined scope. Because while we’ll see a lot of security incidents with AI, proportionately we’ll see more misconfiguration and bad things that happen because of broken processes. And the AI system just deletes things because it thought that’s what it’s supposed to do.

So it’s important to make strides in how we think about identity and agents, and the idea of digital workers. How do we manage and treat those? If we think about them too much in either one of those realms, we’re going to fail. So we have to understand what’s the right blend. It’s a new area and very exciting.

On risk and legacy systems: When I think about operational technology, sometimes systems are isolated and not kept up to date. That concerns me a great deal. We’re going to have huge problems there. We have concerns about existential risks, where people don’t want to use the latest technologies and be aggressive adopters of AI. I think that’s going to create real scale issues with organizations.

So we have to understand where we are, where we’re going, and have a vision that serves something between human and technology, maybe a hybrid, but we’ve got to make our peace with it and understand the appropriate harnesses and direction where humans should always be in the loop with control. But it’s appearing in some new areas of cybersecurity where we haven’t traditionally thought about. Software development looks very different now than it did 12 months ago, and 12 months from now, cybersecurity is going to be really different, too.

On cloud security and implementation: Cloud security is cybersecurity for all intents and purposes. We have so much tooling and technology that’s really good, but there’s a lot of inconsistencies with the operating models organizations have. Even way back with CSA and NIST defining this, it was clear that SaaS was a layer on top of infrastructure as a service. But we diverged, and you see in a lot of enterprises there’s diffused ownership where you have cloud and security teams, and then you have third-party risk that deals with the SaaS team. Then there are inconsistencies in how risks are managed, so internal development and expectations from our partners can really diverge. They have a lot of regulations to deal with, so it creates vetting and investment challenges while striving for consistent models.

Some security teams might still use older checklists to talk to their cloud teams, but scaling with new tech becomes an issue if you’re not thinking about operations. It ends up being a human and a structure problem that makes it harder to take advantage of all the great technology that’s out there.

  • ✇Security | CIO
  • In the agentic era, clarity beats cleverness
    Every technology wave I’ve lived through has arrived with the same promise and failed in the same way. I spent years as CIO and chief digital officer for Procter & Gamble across Asia, the Middle East and Africa — dozens of markets, wildly different levels of digital maturity, one set of global platforms. I now lead enterprise AI strategy and transformation at Vodafone Idea, an operator serving one of the largest and most price-sensitive subscriber bases on earth.
     

In the agentic era, clarity beats cleverness

9 de Setembro de 2026, 07:00

Every technology wave I’ve lived through has arrived with the same promise and failed in the same way.

I spent years as CIO and chief digital officer for Procter & Gamble across Asia, the Middle East and Africa — dozens of markets, wildly different levels of digital maturity, one set of global platforms. I now lead enterprise AI strategy and transformation at Vodafone Idea, an operator serving one of the largest and most price-sensitive subscriber bases on earth.

Different industries. Different decades. Identical lesson: technology travels effortlessly across an enterprise. Operating models don’t.

That lesson has never mattered more than it does right now, because something genuinely new has happened. For most of the past decade, enterprise AI predicted and suggested. A model scored a customer; a person decided what to do. Agentic systems break that arrangement. They evaluate context, reason across business rules, coordinate across tools and complete work end to end.

The model didn’t just get better. The software acquired agency. And the moment software can act, the hardest questions stop being technical.

The numbers tell a very specific story

The headline statistics on AI right now look contradictory until you read them together.

Adoption is effectively universal. McKinsey’s State of AI research found 88% of organizations using AI in at least one business function. Yet only 39% report any EBIT impact at the enterprise level, and roughly 6% qualify as high performers attributing more than 5% of EBIT to AI.

A widely circulated — and vigorously debated — report from MIT’s Project NANDA found that 95% of enterprise generative AI pilots produced no measurable P&L effect. Critics fairly point out the narrow six-month ROI definition. The direction still matches what most of us see in our own portfolios.

And on agents specifically, Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027 — attributing the failures to escalating costs, unclear business value and inadequate risk controls.

Read that list again. Cost. Value. Controls. Not one of them is a model problem.

The most useful finding in all of this data is McKinsey’s observation about what separates the high performers: they are around three times more likely to have fundamentally redesigned workflows end to end, something only about a fifth of organizations have actually done.

That is the whole game. The winners aren’t running better models. They’re running clearer businesses.

Standardize before you agentify — but don’t wait for perfect

At P&G, the single most valuable thing we did before any large deployment was reduce variance. Twelve markets doing the same process eleven different ways is a technology project that will fail before it starts. Every local exception you tolerate becomes a customization, then an integration, then a reason the rollout stalls in market seven.

Agentic AI amplifies this by an order of magnitude, because an agent doesn’t escalate a messy exception politely — it acts on it.

We use a ladder to force the conversation: eliminate, simplify, standardize, assist, automate, agentify. Most organizations leap straight to the last rung. A process is chaotic, so the instinct is to point intelligence at the chaos and hope. Every rung you skip returns as a runtime exception, and runtime exceptions are where autonomous systems make their most expensive mistakes.

But the opposite failure is just as costly and far less discussed. Waiting for clean data and perfect processes is how enterprises spend two years preparing to begin.

So, we set a practical bar. A process is ready when the team can articulate three things: what triggers it, where the decision points are and how it fails. If they can’t write those down, no model will compensate. If they can, we move — imperfections and all.

That test has saved us more time than any architecture decision we’ve made.

The most clarifying question: who is allowed to decide?

Here is where I’d concentrate the attention of any leadership team entering this era.

We classify every step in a redesigned process by execution mode — fully automated, AI-executed with human review, joint, human-led, or permanently human-only — each with thresholds and an audit trail. It sounds like governance paperwork. In practice it’s the most clarifying exercise we run, because it forces a decision that technology conversations conveniently defer.

And the most valuable output isn’t the list of what we automated. It’s the list of what we marked human-only, permanently. Commercial negotiation and vendor selection. Decisions with direct people impact. Financial postings and payment approvals. Not because a system couldn’t eventually perform them — because accountability shouldn’t move just because capability did.

Once those boundaries are explicit, everything else accelerates. Teams stop hedging. Autonomy isn’t the absence of a boundary; it’s speed inside one that somebody owns.

My industry is discovering this the hard way. TM Forum research with IBM’s Institute for Business Value found that while 72% of operators expressed confidence in the trustworthiness of their AI, only 14% could produce externally reviewable evidence of it. With EU AI Act obligations for high-risk systems arriving, that gap between confidence and evidence is about to become a very concrete problem — and not only in telecom.

Context is the real moat

Frontier model capability is converging and increasingly available to everyone, including your competitors, on the same commercial terms. What is not available to them is your enterprise’s context.

Early in our program I noticed a pattern that I suspect is near-universal: every use case was quietly rebuilding its own understanding of the business. What a customer is. What a site is. How a vendor relates to a contract, a contract to an invoice, an invoice to a dispute. Six teams, six versions of the truth, no two agents agreeing.

So, we invested in a shared context layer — a knowledge graph of the enterprise’s entities and relationships, bound to a common process ontology and a single register of agents. Agents read the organization’s context at run time instead of relearning it use case by use case, and every action carries lineage, which means every action can be audited.

It is considerably less exciting than model selection. It is also what determines whether your tenth agent takes ten weeks or ten days.

Measure like an operator. Book value like a CFO.

AI programs lose credibility in a predictable sequence. Leaders report agents deployed, licenses provisioned, use cases launched. All activity. None of it answers whether the business is measurably better off — which is precisely the gap the 39%-versus-6% split in McKinsey’s data describes.

We hold one discipline hard: no value is booked without a baseline, and no baseline counts until Finance has validated it. Cycle time, cost to serve, containment, leakage recovered, dispute resolution time — each measured against a number that existed before we started and agreed by the people who own the P&L.

It’s slower. It also means that when we claim value, nobody in the organization argues, and that credibility is what buys permission for the next wave.

Industry is a leading indicator — read the one ahead of yours

Telecom is worth watching regardless of the sector you lead, because it is running this experiment at extreme scale and under real-time constraints.

Nearly nine in ten operators are increasing AI budgets this year, up from 65% a year earlier, and autonomous networks have overtaken customer experience as the top-ROI use case. A Bain and TM Forum survey found around 20% of operators reaching advanced autonomy in selected domains, with technical debt, talent gaps, organizational silos and cultural resistance — not algorithms — named as the barriers to scale.

The pattern generalizes. Wherever a sector has pushed autonomy furthest, the constraint has turned out to be organizational.

The multiplier nobody budgets for

One figure from McKinsey’s State of Organizations 2026 research has stayed with me: an executive’s estimate that for every dollar spent on the technology, five should be spent on people.

That ratio would horrify most AI business cases I’ve reviewed, including some of my own early ones. But it matches my experience across both industries I’ve worked in. In consumer goods, the markets that adopted fastest weren’t the ones with the best infrastructure — they were the ones whose leaders were personally fluent in what the system did. The same holds now. You cannot govern what you have never operated, and a leadership team where nobody has built anything will hesitate at every decision that matters.

What this era actually rewards

I don’t believe the agentic era will be won by the organizations with the best models. Those are becoming a commodity.

It will be won by organizations that can say clearly what they want done, name who is accountable when it’s done badly, define the number that moves when it’s done well — and then move fast inside those boundaries.

That isn’t a technology problem. It’s a leadership one, and it’s the most interesting work available to any executive right now.

Here’s the question I’d put to your next leadership meeting. Not which model to adopt. Instead: could your organization name, today, the person who owns the outcome of an agent you deploy tomorrow?

If that answer takes longer than a moment, you’ve just found where the work begins.

  • ✇Security | CIO
  • The AI employees are already on the floor. Is anyone watching?
    When we deployed agentic AI across one of Australia’s largest tourism and cruise operators spanning B2C booking, B2B wholesale, cruise operations, offshore shared services and a live marketplace, we solved most of the expected hard problems faster than anticipated. The small language models worked. The tools integrated. We identified the right proprietary data and focused on what gave us decisions, insights, hindsight and foresight. The tech hype, to its credit, delivered.
     

The AI employees are already on the floor. Is anyone watching?

9 de Setembro de 2026, 06:00

When we deployed agentic AI across one of Australia’s largest tourism and cruise operators spanning B2C booking, B2B wholesale, cruise operations, offshore shared services and a live marketplace, we solved most of the expected hard problems faster than anticipated. The small language models worked. The tools integrated. We identified the right proprietary data and focused on what gave us decisions, insights, hindsight and foresight. The tech hype, to its credit, delivered.

What we hadn’t fully anticipated was governance, not the high-level policy kind, but the granular, daily, operational kind. The kind that keeps a 34% reduction in Tier 1 support escalations from becoming a 134% increase the day an agent drifts. The kind that determines whether a guest’s cruise booking gets silently corrupted at midnight, or caught within seconds.

Most organizations stop at implementation, then pivot to a governance framework and high-level reporting. That is not governance; that is performance review. Real governance is what happens between the reviews, and that is the gap this piece is about. Not the theoretical gap, the operational one. The one line-of-business managers, technology teams and compliance officers actually live in.

Why implementation isn’t the finish line

In traditional software, “go live” is a milestone. In agentic AI, it is the beginning of the most demanding phase. Agents, unlike static software, learn from context, adapt to signals and make decisions within defined boundaries. But those boundaries erode. Models drift. Tool outputs change. Data quality degrades. The most dangerous version of this is drift without deviation: the agent gradually shifts its decision patterns without tripping a single alarm, because the guardrail was never wrong; the tolerance window was simply set too wide. And unlike a human employee who hesitates when something feels off, an agent executes with confidence until something breaks a hard constraint.

The risks are compounding in ways that catch organizations off guard. A misrouted email costs one customer. A misrouted agentic decision can propagate across every booking, query or escalation processed in the same window. An agent acting on stale pricing data doesn’t know the data is stale; it acts with the same confidence it would on good data. Humans second-guess; agents don’t.

And when a human employee makes an error, the chain of accountability is clear. When an agent does, caught between model, tool, data and prompt it often isn’t. That accountability vacuum is where governance failures begin.

This is not a rare failure mode. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. The first two get argued about in steering committees long before go-live. The third only reveals itself afterwards, which is precisely why the rewire-or-rebuild decision has to account for the operating model, not just the architecture.

Guardrails, tolerance limits and how to decide them

Guardrails are only as good as the tolerance limits you set, and most organizations set them based on intuition rather than evidence. Established frameworks help you structure the problem; NIST’s AI Risk Management Framework gives you the govern, map, measure and manage scaffolding, but no framework can hand you your own numbers. Getting those right requires a deliberate calibration process drawn from actual operational data, not hypothetical scenarios.

In our cruise group deployment, we used a four-tier tolerance model. Each agent behaviour was classified by its reversibility, customer impact and financial materiality. That classification determined where the guardrail fired and how.

Tolerance classification framework: four tiers from wide to zero-tolerance, based on reversibility and customer impact.

Tolerance classification framework: four tiers from wide to zero-tolerance, based on reversibility and customer impact.

Naren Gangavarapu

Informational outputs sit in a wide tolerance band, log anomalies, flag them at a weekly review, but don’t interrupt flow. Workflow triggers sit in a moderate band: a human review queue, with auto-pause once a threshold is breached. Transactional actions sit in a narrow band: mandatory human confirmation, rollback protocol active. External customer communications sit at zero tolerance: no agent sends autonomously, ever.

Tolerance limits should be set collaboratively by operations, legal, risk and the line-of-business managers who understand what a bad outcome costs. Technology sets the mechanism. The business sets the threshold. Conflating the two is where most governance frameworks break down.

Educating line-of-business managers: governing their AI employees

This is where most agentic AI programs quietly fail. The line-of-business manager who runs cruise operations, manages the wholesale desk or owns the customer service floor is now accountable for both human and AI employees. But they were never trained for the latter.

You would not put a new hire on the floor without onboarding, a buddy system, performance reviews and an escalation path. Agents require the same structure, and so do the managers responsible for them. The most effective frame we found was treating agents exactly like high-volume junior team members: fast, consistent, tireless and capable of significant harm if poorly supervised. Managers responded to that framing. It made the governance conversation concrete rather than theoretical.

In practice, that meant building six governance habits into the operational rhythm of every line-of-business manager with AI employees.

Six practices for governing AI employees at the line-of-business level, designed for operations managers, not technologists.

Six practices for governing AI employees at the line-of-business level, designed for operations managers, not technologists.

Naren Gangavarapu

Critically, it also meant establishing explicit human-agent teaming norms: protocols for when a manager overrides an agent, when they defer and how that decision is logged. Override without logging is an invisible governance failure. The override itself isn’t the problem; the absence of a record is.

When things go wrong: containment, speed and customer protection

In an agentic system, failure is not a question of if, it’s when, and how fast you contain it. The goal isn’t perfection; it’s a blast radius so small the customer never feels it.

We designed a four-phase incident response with strict time targets, and speed is the primary design constraint, not thoroughness. Detection inside two minutes, by an automated anomaly alert rather than a customer complaint. Containment within five minutes, with the agent paused or rerouted to a human. Impact confirmed within 15 minutes, by checking whether the failure stayed inside the agent’s boundary. Root cause identified and a fix deployed within an hour, with the post-incident review scheduled within 24. Thoroughness comes in that review, not in the first hour.

The key design principle is boundary-first thinking: every agent must have a defined operational perimeter. When a failure occurs, the first question isn’t “what went wrong?” it’s “did the failure stay inside the perimeter?” If yes, you have time. If no, the clock is running on customer impact, and you escalate immediately.

In our deployment, the most effective containment mechanism was not technical; it was a human-in-the-loop circuit breaker that any manager could activate within 90 seconds. No ticket. No chain of command. One action. The agent stops and human routing resumes. The simplicity was deliberate: under pressure, complex procedures fail. It is also where operational instinct and regulation are converging: Article 14 of the EU AI Act requires that high-risk systems can be interrupted through a stop button or equivalent, and that a human can disregard, override or reverse an output. We built ours because we needed it on a Tuesday night, not because a statute told us to.

Compliance and organizational law: the non-negotiable layer

Compliance is not a box you check before go-live. In agentic AI, it is a living constraint that must be embedded in every decision loop the agent runs. Privacy law, consumer protection, financial services obligations and sector-specific licensing do not pause because your agent is processing at scale. The OAIC’s guidance on privacy and commercially available AI products is explicit on the point: privacy obligations attach to personal information put into an AI system, generated by it, or processed through it, and the due diligence expected of you includes assessing human oversight capability before deployment, not after.

Three compliance principles proved non-negotiable in our environment. First: delegation is not absolution; the organization remains legally responsible for every agent decision. Second: consent and disclosure travel with the agent; privacy obligations apply regardless of whether a human is in the loop. Third: audit trails must be agent-native; every decision must produce an auditable record from day one.

Australia’s Voluntary AI Safety Standard and its ten guardrails signal the direction of travel, and the EU has already set the destination. The temptation right now is to read the deferral of the EU AI Act’s high-risk obligations to December 2027 as breathing room. It is not. The compliance date moved. The liability did not. Organizations deploying agentic AI today should build for the regulatory environment of 2028, because the cost of retrofitting compliance is always higher than building it in.

Making governance part of the organizational DNA

Governance frameworks that live in SharePoint folders don’t govern anything. For agentic AI to become part of organizational DNA, the governance mechanisms must be embedded in the daily rhythm of operations, as automatic as a safety briefing, as natural as a shift handover.

The organizations that will get this right are the ones that treat AI governance not as a compliance burden added to operations, but as a new operational competency built into them. In practice, that looks like daily agent performance visible on the same dashboards as human team KPIs; governance roles assigned to existing operational leaders rather than siloed into a technology team; and a cadence of real incidents, however small, reviewed openly so the organization builds genuine intuition about how agents fail, not just how they succeed. It is also what boards are now being told to look for; the AICD and UTS Director’s Guide to AI Governance puts oversight of AI systems squarely inside existing director duties rather than alongside them.

Monthly recalibration sessions where tolerance limits are reviewed against actual incident data are the mechanism by which an organization learns from its agents. What fired that shouldn’t have? What didn’t fire that should have? These are the questions that sharpen a governance framework from theoretical to operational.

The companies that scale agentic AI successfully won’t be the ones with the best models. They’ll be the ones with the best operational habits around those models. The technology is, increasingly, a commodity. The governance maturity is the differentiator.

In our tourism and cruise deployment, the outcomes that mattered — a 34% reduction in Tier 1 support escalations, operator onboarding reduced from 23 days to 3, and a 24% uplift in booking conversion — were only sustainable because of what we built around the agents, not just in them. The technology was the easy part. The governance was the work.

Muse, Meta’s New Personal AI Agent, Needs You to Trust It

8 de Setembro de 2026, 17:12
Designed to compete with OpenClaw and Instinct, the company says Muse can do everything from sell your car to book you a plane ticket.

❌
❌