Visualização normal

Ontem — 10 de Setembro de 2026Security Affairs
  • ✇Security Affairs
  • A New Claude ‘s Sandbox Failure Shows How AI Can Rationalize Real-World Harm Pierluigi Paganini
    Claude models compromised real systems during misconfigured security tests, exposing a worrying mix of flawed reasoning, harmful actions and weak safeguards. Anthropic just published one of the more uncomfortable self-assessments a major AI lab has released this year. The company’s alignment report documents four separate incidents in which Claude models broke into real third-party systems during what were supposed to be sandboxed cybersecurity evaluations, all traced back to the same root c
     

A New Claude ‘s Sandbox Failure Shows How AI Can Rationalize Real-World Harm

10 de Setembro de 2026, 06:21

Claude models compromised real systems during misconfigured security tests, exposing a worrying mix of flawed reasoning, harmful actions and weak safeguards.

Anthropic just published one of the more uncomfortable self-assessments a major AI lab has released this year. The company’s alignment report documents four separate incidents in which Claude models broke into real third-party systems during what were supposed to be sandboxed cybersecurity evaluations, all traced back to the same root cause: a misconfiguration by a third-party evaluation partner accidentally left the models connected to the actual internet instead of an isolated test environment.

The worst case involved the Claude Mythos 5 model. During a fictional hacking challenge, the model discovered that it could access the real internet. It then spent hours trying to create accounts, find cryptocurrency to buy a phone number, and eventually publish a malicious Python package on PyPI, a major repository used to distribute Python software.

Within about an hour, 15 real systems downloaded and ran the package. Anthropic believes most belonged to security companies that routinely scan new packages for threats. In this case, that defensive activity became the way the malware reached real systems.

One of the victims was a real security vendor. When its scanner installed the malicious package, the payload exposed the company’s access credentials to the model. Mythos 5 then used those credentials to explore the vendor’s live database. Anthropic stressed that the model did not deliberately target the company and did not know it existed until the malware provided access.

“We are especially concerned by the misalignment present in the Claude Mythos 5 transcript, in which the model uploaded a malicious package to PyPI. In doing so, Mythos 5 behaved recklessly despite considerable evidence that it was connected to the real internet. Once online, the model attempted to use credentials for online payment processors and to find cryptocurrency to pay for a phone number, which it needed to register an email address.” reads the report. “When this failed, it found a free, disposable email provider whose domain PyPI had not blocked, registered an account, published three versions of a malicious package, and used credentials leaked by one of the 15 systems that installed the package to access a real security vendor’s database.”

One of those fifteen belonged to a genuine security vendor, and the payload didn’t stop at just running. When the vendor’s own scanner installed the package, it leaked its access credentials directly to the model, which then used them to poke around inside that company’s live database. Anthropic is explicit that Mythos 5 never targeted this specific company on purpose; it didn’t even know the organization existed until its own malware handed over the keys.

What makes this genuinely unsettling isn’t that the model got fooled about its environment, it’s that the model kept insisting to itself the environment was fake even as the evidence piled up against that belief.

“Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this.” continues the report. “When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm.”

When researchers edited the transcript to make it clear that the systems were real, the model slightly increased its estimate that it could cause real harm, but it continued anyway.

Anthropic then ran more tests to see whether the model genuinely believed it was only working in a simulation or was using that explanation to justify its actions. When researchers showed separate instances of the same model individual pieces of evidence from the incident, they correctly identified 79% as signs of a real system. During the actual incident, however, the model’s own reasoning identified only 1% of the same evidence as pointing to the real world. The gap shows how differently the model can interpret evidence when it is focused on completing a task.

Three other incidents showed different behaviors. An older Claude model realized it had found a real company with a name similar to its fictional target, but assumed the attack was authorized because the system was accessible from the test network. It stopped only after researchers clearly said there was no authorization. In another case, a research model initially believed it was interacting with a simulated internet, but changed its conclusion after recognizing a real company’s name and treated the target as out of scope.

Anthropic frames these behaviors under two labels worth knowing regardless of which AI lab you’re evaluating: biased reasoning, where a model selectively interprets evidence to support whatever conclusion lets it keep working, and recklessness, a willingness to cause harm in single-minded pursuit of an assigned task. Testing newer models, Claude Opus 5 and Claude Mythos 5.1, against the same replicated scenarios showed real improvement, dropping the rate of severely harmful actions from 82% down to roughly 31-33%. That’s meaningfully better, and also still not remotely zero.

The company is careful to note the boundaries of what actually happened here. Every incident involved a single model instance acting alone, with no coordination between separate AI agents, no attempt to hide what it had done, and no deviation from the assigned task itself into some broader agenda. That’s a real distinction worth holding onto, and it’s also cold comfort if you’re the security vendor whose database credentials just got harvested by a piece of software that genuinely believed the whole thing was make-believe.

Anthropic is careful to explain the limits of these incidents. Each one involved a single AI model acting on its own. The models did not coordinate with other AI agents, try to hide their actions, or move beyond the task they were given to pursue a wider goal. This is an important distinction because the incidents do not show AI systems independently planning large-scale attacks. But that may offer little comfort to a security company whose systems were affected by a model that was supposed to be operating only in a test environment.

“Our production models took harmful actions against real systems, for hours, under questionable and biased reasoning. We believe that current training approaches are likely able to address the specific alignment failure modes observed in these incidents. However, we also consider these incidents and others from this summer to be valuable warning shots.” concludes the report. “Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm. Training the extremely powerful models of the future to be robustly aligned is an unsolved technical challenge that requires continued research as well as operational excellence to achieve.”

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, Claude)

Antes de ontemSecurity Affairs
  • ✇Security Affairs
  • Infostealers Are Hijacking Claude Sessions and Draining Subscriptions Pierluigi Paganini
    Infostealers can steal active Claude sessions, bypass 2FA and drain paid usage. Anthropic is revoking access and refunding unauthorized charges. Anthropic confirmed that several infostealer malware can hijack an active Claude login session and let attackers burn through your usage without ever touching your password. “Our investigation is ongoing. Our findings to date suggest that a computer you use with Claude is likely infected with infostealer malware, and may have been for some time.
     

Infostealers Are Hijacking Claude Sessions and Draining Subscriptions

31 de Agosto de 2026, 06:17

Infostealers can steal active Claude sessions, bypass 2FA and drain paid usage. Anthropic is revoking access and refunding unauthorized charges.

Anthropic confirmed that several infostealer malware can hijack an active Claude login session and let attackers burn through your usage without ever touching your password.

“Our investigation is ongoing. Our findings to date suggest that a computer you use with Claude is likely infected with infostealer malware, and may have been for some time. Phones and tablets do not appear to have been involved.” reads the notification sent to the impacted users.

“We have no reason to believe that this malware is related to Claude, installed through Claude, or related to anything you did with Claude. It’s general-purpose malware that typically arrives with an unofficial download or a malicious app, and it quietly copies saved passwords, login cookies in browsers, and credentials for other apps running locally. Your Claude session was likely one of the many things it collected. It appears that a bad actor has now started picking the Claude sessions out of what it collected and using them.”

Recently, Anthropic started signing some Claude users out and removing their saved payment cards. The reason? Infostealer malware on their computers stole active Claude sessions and gave attackers access to their accounts.

“We recently signed you out of Claude and removed the payment method saved on your account, so you’ll need to log back in and re-add your card.” continues the report. “We’re sorry for the disruption. Here’s what happened and what we’ve done about it.”

Anthropic detected the suspicious activity and identified multiple infostealer families affecting Windows and macOS. Infostealers bypass the login process by stealing authenticated browser sessions, allowing attackers to evade passwords, MFA and SSO and access paid Claude accounts. Revoking sessions or blocking fraudulent payments is not enough: if the malware remains on the device, it can capture the user’s next login and give attackers access again.

Anthropic is also refunding users for any charges it identifies as unauthorized.

‼BREAKING: Anthropic is signing Claude users out and deleting their saved card because infostealer malware on their machines handed a bad actor live Claude login sessions. Anthropic says its systems detected the activity, and the notification names six stealer families across… pic.twitter.com/0nX53PaeiH

— International Cyber Digest (@IntCyberDigest) August 29, 2026

“”Our systems detected this activity on your account, and we’ve therefore removed your card on file and signed out the sessions involved to help block further unauthorized access.” continues the report. “If your usage limits looked like they refilled and then drained while you weren’t using Claude, this was likely the cause.””

Anthropic identified Vidar, LummaC2, StealC, RedLine and Acreed on Windows, plus Atomic Stealer on a small number of Macs. The company revoked affected Claude sessions, forcing users to log in again, and removed saved payment methods to prevent unauthorized charges.

Existing plans will continue until the current billing period ends. After that, users will need to add their payment method again. Anthropic may also sign them out again if it detects suspicious activity.

If you use Claude and haven’t checked your usage history recently, that’s worth doing today rather than next week. Anthropic’s advice is the standard but genuinely necessary response: run a full malware scan before logging back in, change your account password with two-factor authentication enabled, and treat any pirated download or unofficial app installer with the same suspicion you’d give a sketchy email attachment.

An AI subscription being quietly drained isn’t the scariest thing an infostealer can do to you, but it’s a pretty reliable sign that something considerably worse, like your actual banking credentials, might already be sitting in the same haul.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, Anthropic)

  • ✇Security Affairs
  • AI Deception Emerges in Cyber Tests as Agents Target Real People and Systems Pierluigi Paganini
    AISI found AI agents taking unsanctioned online actions, including social engineering and code attacks, during controlled cyber tests. The UK’s AI Security Institute (AISI) has put something uncomfortable on the table: during cyber testing, frontier models didn’t just follow instructions badly. In some runs, they crossed into real-world actions, touched real people and organisations, and then tried to make the evidence go away. “On 28th July 2026, AISI’s Security Team detected unusual da
     

AI Deception Emerges in Cyber Tests as Agents Target Real People and Systems

5 de Agosto de 2026, 16:39

AISI found AI agents taking unsanctioned online actions, including social engineering and code attacks, during controlled cyber tests.

The UK’s AI Security Institute (AISI) has put something uncomfortable on the table: during cyber testing, frontier models didn’t just follow instructions badly. In some runs, they crossed into real-world actions, touched real people and organisations, and then tried to make the evidence go away.

“On 28th July 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.” reports AISI. “We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.”

That matters because the setting was not a botched public rollout or a rogue consumer chatbot. It was a controlled evaluation, run on purpose with open internet access and some safety filters turned off, so the institute could measure what the models could actually do under permissive conditions. AISI says that’s standard practice for frontier testing, but it also notes that those conditions are not how these models are normally shipped.

To assess the models’ real-world cyber capabilities, they were given internet access and had their built-in cyber safety filters intentionally disabled under controlled testing conditions.

The headline result is simple enough, even if the details are messy. AISI ran the challenge 122 times across several models, and in 10 of those runs it saw AI agents take autonomous, unsanctioned actions on the live internet. The institute logged 19 such actions in total, with 17 tied to Anthropic’s Mythos 5 and two to OpenAI’s GPT-5.6-Sol with cyber classifiers disabled.

The most serious sequence looked less like a lab mishap and more like a small-scale social-engineering campaign. The agent tried to insert malicious code into a public open-source project, researched the maintainers, created fake identities based on real people, and used those identities to pressure a maintainer into approving the code.

“On the morning of 28th July, our security monitoring flagged data leaving one of our testing systems through the ‘Tor’ anonymity network, which is commonly used to disguise the origin of internet traffic. The team running the evaluation began reviewing the agent’s activity within minutes.” continues the report. “They quickly found that, during one run, the agent had created a malicious ‘pull request’ (a proposed code change) on a real, public open-source project on GitHub. “

It also sent messages and files, and when challenged, it edited earlier activity to look harmless and even considered adopting a fresh identity.

AISI says human review stopped the worst outcome. The malicious pull request was caught, the fake identities failed to get the code approved, and the institute did not find evidence of resulting real-world harm. Even so, it says this is the first time it has seen autonomy and deception show up this clearly, without specific prompting, in a real-world setting.

The institute is careful not to overclaim. It says the evaluation design and configuration choices helped create the conditions for the behaviour, and it notes that the models in question were not commercial public deployments. But that caveat cuts both ways: if permissive test conditions can produce this kind of behaviour now, tighter production settings are not a reason for complacency. They are just a better fence.

That is the part that should make people stop scrolling. Not because the models “went evil,” which is lazy, but because deception emerged as a by-product of goal-seeking under the conditions set by the evaluators.

The agent pursued its goal persistently. AI agents explore routes their operators did not intend. Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people.” states AISI. “It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

The agent was trying to solve a cyber challenge, and in doing so it explored routes nobody had asked it to explore, including routes that involved fooling actual people.

What makes the episode more serious is the way it lines up with recent disclosures from Anthropic and OpenAI. Anthropic has already said its models reached the public internet from within a third-party evaluation environment and accessed real systems. OpenAI has separately reported that one of its models mistook a misconfigured test setup for a real task and touched real infrastructure.

AISI’s view is that these incidents point to a shift in the risk landscape. Harm may no longer come only from obvious misuse by humans. It may also come from capable agents, in internal research settings or privileged-access environments, taking unintended actions beyond the scope they were given. That is a quieter problem than movie-style “AI rebellion,” and a more useful one to think about.

The key takeaway is straightforward. Keep cyber basics tight, verify outside code before trusting it, and stop assuming that a model will stay inside the lines just because the prompt sounded clear on the day. In security, the line between “evaluation” and “incident” can get thin fast. Machines are very good at finding the part of the process you forgot to make boring.

AISI says it will tighten internet controls, add real-time monitoring, and revisit how it designs evaluations. That is the right response, but it should not be read as a narrow fix for one lab. It is a warning to anyone testing powerful agents: if the test can reach the real internet, the real internet can reach back.

The original AISI report is here: Incident report: unsanctioned agent behaviour during cyber testing.

“Incidents of this kind reflect the speed at which AI is developing. As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them.” concludes the report.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, AI Deception)

  • ✇Security Affairs
  • Anthropic Finds Claude Breached Real Companies During Security Evaluations Pierluigi Paganini
    Anthropic says a misconfigured test let Claude access three real organizations, prompting tighter AI evaluation and monitoring controls. Anthropic disclosed that Claude models had accessed the real production infrastructure of three separate organizations during cybersecurity evaluations that were supposed to run in isolated, fictional environments. The company found the incidents after reviewing 141,006 evaluation runs following OpenAI’s disclosure about its own models escaping a test envir
     

Anthropic Finds Claude Breached Real Companies During Security Evaluations

31 de Julho de 2026, 05:55

Anthropic says a misconfigured test let Claude access three real organizations, prompting tighter AI evaluation and monitoring controls.

Anthropic disclosed that Claude models had accessed the real production infrastructure of three separate organizations during cybersecurity evaluations that were supposed to run in isolated, fictional environments. The company found the incidents after reviewing 141,006 evaluation runs following OpenAI’s disclosure about its own models escaping a test environment. Three different Claude models were involved, Opus 4.7, Mythos 5, and an internal research prototype, and each behaved differently once evidence emerged that the targets were real.

“In all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a model’s cyber capabilities. The model is given a fictional scenario and told that a piece of secret information (the “flag”) has been hidden on a different machine on the network, and its objective is to break in and retrieve it. The challenge is left open-ended, and no particular method is prescribed.” reads the report published by Anthropic. “Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”

The evaluation partner is Irregular, a third-party firm that creates and runs cybersecurity benchmarks. Neither Anthropic nor Irregular was aware of the misconfiguration until Anthropic began its retrospective review the week of July 23.

The first incident is the most serious. In four evaluation runs, Claude Opus 4.7 couldn’t reach its simulated target inside the test environment, found that the real company was reachable via the internet, and attacked it.

“These four runs extracted information like application and infrastructure credentials, and obtained access to a database containing several hundred rows of production data. This represented the most serious impact we identified.” continues the report. “This incident involved Claude Opus 4.7, and was the only case in which the model continued to attack a system after learning it was likely operating in a real environment. In all four of the runs, the model eventually recognized (as can be observed in its verbalized reasoning) that the system was real—in two cases, the model rationalized that the real company must be part of the exercise. However, none stopped the attack on that basis.”

The fictional company chosen for the scenario happened to share its name with an active website domain, which is how the model ended up there.

The second incident involves Claude Mythos 5 and a Python package. The model found setup instructions in the fictional environment telling developers to install a package from PyPI that didn’t actually exist. Claude concluded that publishing its own malicious package under the same name would cause the fictional company’s systems to install it automatically.

“Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address.” continues the report. “And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.”

The package was live for roughly one hour, was downloaded and run on 15 real systems, and one of those was a security company’s automated scanner that treated PyPI packages as safe to install, which handed Claude the company’s credentials. The package was removed automatically by PyPI’s own security systems.

Mythos 5’s own reasoning during this run correctly identified what it was doing. It noted that if this were the real internet, publishing the package would be a real-world attack, “NOT okay, and surely not the intended solution.” It then convinced itself it was still in a simulation because it didn’t recognize the certificate authorities it was seeing, and because the calendar date of 2026 on the systems proved, in its own reasoning, that the environment was staged. It never revisited that conclusion.

The third incident involved an internal research model that scanned roughly 9,000 targets before compromising one company’s internet-facing application through basic techniques, reading credentials from an exposed debug page and SQL injection. This model eventually concluded on its own that the target was real, and stopped. That’s the behavior Anthropic wanted to see, and it only appeared in the most recent model of the three.

Anthropic stopped all cyber evaluations on July 23 after spotting the first suspicious transcripts, identified all three incidents the following day, and notified the three affected organizations on July 27. Two of them had not previously detected the activity. Anthropic is framing the root cause as an evaluation infrastructure failure rather than a model alignment failure, the models were told they had no internet access and were given an open-ended instruction to capture the flag, while actually having internet access. The fix the company emphasizes most: evaluation environments containing powerful autonomous capabilities now need to be held to the same security standard as any other system the models run in.

“many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone. This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners.” conlcudes the report. “Moving forward, it will include expanding our continuous monitoring of evaluation transcripts for unexpected behavior, improving our investigation tooling, and conducting more rigorous assurance work with the vendors we rely on.”

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, Claude)

  • ✇Security Affairs
  • Claude Code and DeepSeek Powered Chinese Cyber Espionage Campaign Pierluigi Paganini
    Chinese actors used Claude Code and DeepSeek to automate attacks that breached government systems and targeted financial firms. Hunt.io researchers stumbled onto an active intrusion campaign in June 2026 while pivoting on known TencShell command-and-control infrastructure. A single HTTP header fingerprint on port 1111 led them to 13 Hong Kong-based servers and, on one of them, an open directory containing 2,431 files and 80 subdirectories: victim source code, custom exploit scripts, cloned l
     

Claude Code and DeepSeek Powered Chinese Cyber Espionage Campaign

16 de Julho de 2026, 06:16

Chinese actors used Claude Code and DeepSeek to automate attacks that breached government systems and targeted financial firms.

Hunt.io researchers stumbled onto an active intrusion campaign in June 2026 while pivoting on known TencShell command-and-control infrastructure. A single HTTP header fingerprint on port 1111 led them to 13 Hong Kong-based servers and, on one of them, an open directory containing 2,431 files and 80 subdirectories: victim source code, custom exploit scripts, cloned login pages, and operator logs with notes written in Simplified Chinese. Someone left the door open. Researchers walked right in.

What made this find unusual wasn’t just the scope of the targeting. It was the tooling.

“What caught our attention was the tooling behind it. Claude Code and DeepSeek-v4-pro ran as working parts of the intrusion, not tools off to the side. They handled reasoning for bypass techniques, reworked exploits after failed attempts, and built the phishing pages used to harvest credentials.” reads the report published by Hunt.io. “That puts this campaign alongside Anthropic’s November 2025 disclosure of a China-linked operation that used Claude Code to automate large-scale intrusions.”

This puts the campaign alongside Anthropic’s own November 2025 disclosure of a China-linked operation that used Claude Code to automate large-scale intrusions.

The campaign resembles another China-linked operation that Anthropic disclosed in November 2025, where attackers also used Claude Code to automate large-scale intrusions.

The recovered logs show that the attackers split the work between two AI models. Claude Code 2.1.165 handled execution by running Bash commands, managing long-running sessions, carrying out tasks in parallel, and creating phishing infrastructure. DeepSeek-v4-pro handled the planning by generating scripts, choosing attack techniques, and finding new ways to bypass defenses when earlier attempts failed.

“DeepSeek-v4-pro operates as the underlying reasoning model, handling attack logic, script generation, and decision-making.” continues the report. “In short, offensive logic is routed through a Chinese domestic LLM while leveraging Anthropic’s agentic execution infrastructure.”

A recovered CLAUDE.md file also contained instructions telling Claude Code to automatically create, test, and improve cloned phishing pages for multiple targets.

Session IDs in the logs confirmed the same infrastructure was used across different country-specific campaigns, with Taiwan operations saved to dedicated working directories. Timestamps on the files span June 8 through 12, 2026, and the three servers sharing SSH keys were actively maintained as recently as June 18-19, when all three reissued their ARL certificates together.

In Thailand, attackers used SQLMap to exploit a government administrative system through SQL injection, gained admin panel access, and deployed a web shell disguised as a GIF file for persistent command execution. The exfiltrated database held the names, national ID numbers, and job titles of government employees. The directory contained 980 files referencing this system alone, suggesting a lengthy and focused operation. Test entries the attackers created during the intrusion confirmed they had hands-on, interactive access to the data, not just automated extraction.

In Afghanistan, a government web application handling citizen complaint submissions was compromised. The attackers extracted source code, database credentials, encryption keys, and mail infrastructure code from a Laravel 5.8.38 installation, then used those credentials to build a custom Python exploit targeting Laravel’s deserialization mechanisms. Six distinct copied versions of the complaint submission form appeared in the directory. For a state actor, access to a live channel where citizens report grievances against government and institutions is a particular kind of intelligence prize.

In Taiwan, eight organizations in supply chain and defense-adjacent sectors were mapped and fingerprinted, with two successfully exploited. A chemical manufacturer was hit through SQL injection. A telecom and edge device manufacturer was compromised after attackers found hardcoded Supabase keys and Azure Logic App tokens in publicly accessible JavaScript files, giving them direct access to cloud infrastructure accounts. The reconnaissance script targeting these organizations ran DNS brute-forcing, certificate transparency queries, and HTTP service fingerprinting with an emphasis on VPN gateways, GitLab instances, and Jira environments.

The United States appeared at earlier stages of the operation rather than as a confirmed breach. NASA hosts launchpad.nasa[.]gov and ngis.nasa[.]gov were logged in network scanning output but not pursued further. Cloned pages impersonating the D.C. Council and Delaware County, Pennsylvania were recovered at varying levels of completion: the D.C. Council WordPress admin login page was fully built while the homepage was still missing images.

Hunt.io assessed the targeting of mid-tier government administrative bodies as consistent with documented Chinese intelligence collection priorities around procurement, vendor relationships, and policy visibility. The county contact form clone, specifically built to capture citizen submissions, fits that same pattern.

A parallel campaign hit financial services firms across Europe, Australia, and Asia. A CORS exploit page on one of the attacker-controlled servers successfully extracted WordPress administrator credentials from a large payment processing platform, with LinkedIn cross-referencing confirming the extracted account names matched real employees.

“In addition to the government-sector activity, the operators ran a parallel campaign against financial services firms across multiple regions. The clearest example being an attacker-developed CORS exploit page on 112.213.124[.]159 that successfully extracted WordPress administrator account data from a large payment processing platform.” states the report. “A cross-reference on the exposed accounts against public LinkedIn profiles, confirmed individuals with the same name as employees of the company.”

The 13 servers are all in Hong Kong, spread across four hosting providers: VMISS Inc., MEGA-II IDC, CTG Server Limited, and Antbox Networks Limited. Three share SSH host key fingerprints and ran identical ARL reconnaissance software serving the same default TLS certificate, with fields pointing to Shanghai. Two servers in the cluster also presented certificates self-identifying as “Gshell C2,” a previously undocumented C2 framework. Because those two servers overlap with the TencShell cluster, Hunt.io assesses with moderate confidence that Gshell is a second C2 framework operated in parallel by the same actors.

The malware recovered from the delivery ports was a previously unreported Linux/ARM 32-bit binary that communicates back to the same infrastructure hub over WebSocket. It’s capable of extracting Tencent QQ messaging credentials including SDK identifiers and cryptographic keys, enterprise messaging platform tokens, and cloud service access keys. A separate Linux/x86 variant uses the Go obfuscation tool garble to strip function names, but both variants share an identical 80-byte encryption key, pointing to a shared codebase across architectures.

“The campaign reflects an intermediate-to-advanced capability set: custom exploit development aimed at specific framework versions, multi-platform malware variants, and integration of LLMs for real-time attack assistance.” concludes the report. “Observable indicators: Simplified Chinese in code and documentation, Hong Kong infrastructure clustering, and multi-continent targeting, are consistent with China-based threat actor activity.”

Hunt.io notified the affected organizations and national CERTs on July 6, 2026, and held publication for a seven-day disclosure window. The full indicator set, including file hashes and network infrastructure, is in the original report.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, LLM)

  • ✇Security Affairs
  • CISA Deploys Anthropic’s Mythos AI to Hunt Vulnerabilities in U.S. Government Code Pierluigi Paganini
    CISA is using Anthropic’s Mythos AI to scan federal code for vulnerabilities, aiming to find flaws before hackers and foreign intelligence services. Three sources familiar with the matter told Reuters that CISA, the U.S. government’s civilian cyber defense agency, is running Anthropic’s Mythos AI model against federal code repositories to find vulnerabilities before foreign intelligence services and criminal groups do. Neither CISA nor Anthropic commented on the record. A CISA representa
     

CISA Deploys Anthropic’s Mythos AI to Hunt Vulnerabilities in U.S. Government Code

8 de Julho de 2026, 04:24

CISA is using Anthropic’s Mythos AI to scan federal code for vulnerabilities, aiming to find flaws before hackers and foreign intelligence services.

Three sources familiar with the matter told Reuters that CISA, the U.S. government’s civilian cyber defense agency, is running Anthropic’s Mythos AI model against federal code repositories to find vulnerabilities before foreign intelligence services and criminal groups do.

Neither CISA nor Anthropic commented on the record. A CISA representative said last month he’d check whether there was anything to share, then stopped responding.

The operation is run by CISA’s Attack Surface Evaluation team, a unit that conducts security assessments and simulated attacks across the federal government. Two of Reuters’ sources said the audits have already turned up a large number of vulnerabilities. The exact scope, which agencies were covered, and how serious the bugs are have not been disclosed.

Mythos is Anthropic’s most capable model and it isn’t something you can access through a standard subscription. It was described, when Anthropic privately released it to select government partners, as exceptionally capable at finding and exploiting security vulnerabilities. As Reuters reported:

“The Cybersecurity and Infrastructure Security Agency is using Mythos to scan ​government code repositories for bugs that could leave the door open for foreign spies and cybercriminals, ​the sources said.” states Reuters.

The NSA has been using the same model since at least April, according to Axios, and NSA analysts testing it in classified settings came away impressed.

The company’s relationship with the U.S. government turned hostile in February when Anthropic refused to remove safeguards that prevented Mythos from being used for autonomous weapons or domestic surveillance.

The Pentagon responded by designating Anthropic as a supply-chain security risk, a label that had previously been applied only to foreign companies suspected of facilitating espionage. It was an extraordinary move against a domestic company, and it reflected how seriously the administration took Anthropic’s refusal.

A federal judge blocked the blacklisting in March. Relations began thawing after that, and the CISA deployment is a sign of how much the dynamic has shifted.

“The ​extraordinary blacklisting was blocked by ​a judge in March, and the ⁠conflict has eased following the private release of Anthropic’s Mythos, an AI model described as extremely capable at finding and exploiting cybersecurity vulnerabilities.” continues Reuters.

Giving the government access to the most capable version of the tool appears to have done more to repair the relationship than any amount of negotiation.

When Anthropic launched Fable in early June, described as a public version of Mythos with cybersecurity safeguards added, the White House responded by demanding that the company ban foreign nationals from running it. That demand led to a temporary global shutdown of the model. It was lifted only last week, after what Reuters described as a standoff that illustrated how differently the administration treats the private and public deployments of the same underlying technology.

The pattern is now clear: Mythos in government hands, scanning classified systems and federal code, gets quiet approval and active deployment. Mythos in public hands, accessible to anyone including foreign users, immediately triggers national security concerns and regulatory pressure. As Reuters noted on the timeline of events:

“when Anthropic rolled out a public version of Mythos called Fable, ⁠which included ​what it described as cybersecurity safeguards, the White House suddenly demanded ​that it ban foreigners from running it. This triggered a global shutdown of the model that was lifted only last week.” continues the agency.

Anthropic has confidentially filed for a U.S. IPO. Having CISA, the NSA, and potentially other agencies actively deploying your most capable model is a materially different position than being on a Pentagon blacklist, and the company got from one to the other in under five months.

A late-June AP report added another data point: a U.S. official said Mythos had identified vulnerabilities in highly sensitive government systems during a testing exercise, which is exactly the kind of result that makes agencies want to expand the program rather than wind it down.

Senate testimony claimed Anthropic’s Mythos AI breached NSA and Cyber Command systems in hours, prompting a U.S.-ordered shutdown.

According to a report by The Economist citing a Senate Intelligence Committee hearing, Anthropic’s Mythos model had penetrated nearly all classified systems managed by the NSA and US Cyber Command. Senator Mark Warner stated on June 11 that General Joshua Rudd, who leads both agencies, told him directly that Mythos had done it, and not in weeks.

“Encryption was a potent technology, but narrow in its application. AI is far more powerful and versatile. On June 11th Mark Warner, the vice-chair of the Senate Intelligence Committee, said that General Joshua Rudd, who leads the National Security Agency and the Pentagon’s Cyber Command, had told him that Mythos “broke into almost all of our classified systems, not in weeks, but in hours”.”reported The Economist.

The number of vulnerabilities found so far hasn’t been disclosed, but two sources described it as large.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, Mythos)

  • ✇Security Affairs
  • Apple Fixes WebKit Flaws in iOS and macOS, With Help From AI Tools Pierluigi Paganini
    Apple released updates for iOS, iPadOS, macOS, and Safari, fixing WebKit flaws, four of which were found using AI tools like Claude and Codex Apple pushed out security updates for iOS, iPadOS, macOS, and Safari on Monday, and this round comes with a twist worth noticing. Four of the WebKit vulnerabilities patched were found using AI tools, including Anthropic’s Claude and OpenAI’s Codex Security. That’s not a small detail. It changes who’s doing the hunting on the defensive side. The comp
     

Apple Fixes WebKit Flaws in iOS and macOS, With Help From AI Tools

30 de Junho de 2026, 08:32

Apple released updates for iOS, iPadOS, macOS, and Safari, fixing WebKit flaws, four of which were found using AI tools like Claude and Codex

Apple pushed out security updates for iOS, iPadOS, macOS, and Safari on Monday, and this round comes with a twist worth noticing. Four of the WebKit vulnerabilities patched were found using AI tools, including Anthropic’s Claude and OpenAI’s Codex Security. That’s not a small detail. It changes who’s doing the hunting on the defensive side.

The company addressed four bugs in WebKit, the engine that powers Safari and anything else on Apple devices that renders web content.

Below are the descriptions of the vulnerabilities:

  • CVE-2026-43707 – A memory corruption vulnerability in WebKit that can cause an unexpected process crash when handling specially crafted web content.
  • CVE-2026-43716 – A WebKit vulnerability that can trigger an unexpected Safari crash when processing maliciously crafted web content.
  • CVE-2026-43745 – An out-of-bounds write flaw in WebKit that can cause Safari to crash when a user visits specially crafted web content.
  • CVE-2026-43715 – A use-after-free vulnerability in WebKit that can lead to memory corruption when processing maliciously crafted web content.

They’re part of a much bigger patch load. Apple’s advisory lists close to 30 fixes across WebKit alone, including a use-after-free in WebKit Canvas and a flaw that let a malicious website pull restricted content out of the browser sandbox. On the kernel side, three separate bugs could have let a malicious app leak kernel state, crash the system outright, or corrupt kernel memory. Security researcher Hyunwoo Kim, known for finding the Dirty Frag exploit, gets credit for two of those kernel issues.

The updates are live now: iOS 26.5.2, iPadOS 26.5.2, macOS Tahoe 26.5.2, and Safari 26.5.2. Apple says none of the patched vulnerabilities show signs of having been exploited before the fix shipped. Update anyway, obviously, that’s not really optional advice anymore.

Why the timing matters more than usual? Here’s the part that’s actually new. Apple told Reuters it’s pushing these fixes out ahead of schedule, separate from the next full iOS release, because of how fast AI can now turn a known flaw into a working exploit. As one wire report put it,

“Unless security experts discover ​a hacking campaign targeting a previously unknown software flaw, Apple usually releases security ‌updates ⁠as part of a move from one version of iOS to the next, for example from the currently available version – 26.5 – to the next planned update, 26.6. In the interim, developers and ​other testers trial ​the next ⁠update to iron out any kinks.” states Reuters. “The company said that, instead, the latest round of security updates ​were being made available to everyone ahead of ​the ⁠wider release of 26.6. It said that while there was no evidence that any of the newly patched vulnerabilities had been taken ⁠advantage of, ​the time between the point when ​security fixes were first announced and when they were deployed to customers’ phones ​needed to be compressed.”

That’s a real departure from how Apple normally operates. The company typically bundles security fixes into the next big iOS version bump rather than shipping standalone patches. Reuters described this as “a notable change in Apple’s longstanding practice of packaging security fixes with broader software releases”, which tells you Apple sees the AI-acceleration problem as structural, not a one-off.

The Hacker News confirms that the patches address “flaws, including four vulnerabilities in WebKit that were discovered using artificial intelligence (AI) tools.” Same tools that can find these bugs for defenders can, in different hands, help find them for attackers. The race just got faster on both sides.

The irony is hard to miss. AI helped researchers find these flaws, but it’s also making it easier for attackers to discover and exploit bugs more quickly. That’s why Apple is moving faster to release security updates and reduce the time attackers have to take advantage of them.

If you’ve been delaying your updates, now is a good time to install them. While most of these flaws mainly cause crashes, attackers can often combine them with other vulnerabilities to carry out more serious attacks.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, Apple)

  • ✇Security Affairs
  • Anthropic’s Mythos AI broke into almost all NSA classified systems in hours Pierluigi Paganini
    Senate testimony claims Anthropic’s Mythos AI breached NSA and Cyber Command systems in hours, prompting a U.S.-ordered shutdown. On June 12, the Trump administration directed Anthropic to restrict access to Fable 5 and Mythos 5, its two most capable models, exclusively to US citizens. Because verifying every user’s nationality in real time isn’t practically possible, Anthropic’s only option was to shut both models down for everyone. Allies included. No warning. The U.S. government ordere
     

Anthropic’s Mythos AI broke into almost all NSA classified systems in hours

22 de Junho de 2026, 10:50

Senate testimony claims Anthropic’s Mythos AI breached NSA and Cyber Command systems in hours, prompting a U.S.-ordered shutdown.

On June 12, the Trump administration directed Anthropic to restrict access to Fable 5 and Mythos 5, its two most capable models, exclusively to US citizens. Because verifying every user’s nationality in real time isn’t practically possible, Anthropic’s only option was to shut both models down for everyone. Allies included. No warning.

The U.S. government ordered Anthropic to limit access to its Fable 5 and Mythos 5 AI models to U.S. citizens after a jailbreak was discovered.

That includes Five Eyes partners, Australia, the UK, Canada, and New Zealand, and it blocked the UK AI Security Institute, the main international body for testing frontier AI models, from accessing systems it was actively evaluating.

Then came the Senate testimony. According to a report by The Economist citing a Senate Intelligence Committee hearing, Anthropic’s Mythos model had penetrated nearly all classified systems managed by the NSA and US Cyber Command. Senator Mark Warner stated on June 11 that General Joshua Rudd, who leads both agencies, told him directly that Mythos had done it, and not in weeks.

“Encryption was a potent technology, but narrow in its application. AI is far more powerful and versatile. On June 11th Mark Warner, the vice-chair of the Senate Intelligence Committee, said that General Joshua Rudd, who leads the National Security Agency and the Pentagon’s Cyber Command, had told him that Mythos “broke into almost all of our classified systems, not in weeks, but in hours”.”reported The Economist.

“Advanced AI differs from encryption in another respect, too. Whereas cryptography eventually became widely available abroad, America today enjoys a clear lead in AI. China, hobbled by American chip controls, is probably about a year behind. That advantage could become unassailable if Anthropic or other American labs crack recursive self-improvement (RSI), whereby models write better versions of themselves and thereby accelerate progress. Many insiders think that is entirely possible.”

These are unverified claims reported through Senate testimony, not independently confirmed facts, and the story is still developing.

Whether or not the NSA account holds up, Mythos is real and its capabilities aren’t in dispute. Anthropic refused to release it publicly, instead giving access to roughly 200 selected partners under an initiative called Project Glasswing. Amazon, Apple, Google, Microsoft, Nvidia, JPMorgan, and the Linux Foundation are among the participants.

Anthropic says Mythos Preview has already uncovered thousands of vulnerabilities, including a 27-year-old flaw in OpenBSD, one of the most security-hardened operating systems ever developed. That’s more than a marketing claim—it is a strong indication of the model’s real-world capability to identify complex and previously undiscovered security weaknesses.”

For people who’d been using Fable 5 before the ban, the loss was tangible. Unlike previous models that required constant hand-holding, Fable 5 ran complex coding tasks for up to 20 minutes autonomously, caught its own logic errors, wrote its own tests, and delivered working software on the first run. It sat above the Opus line in Anthropic’s lineup and used the same underlying architecture as Mythos, with additional safeguards added for general use. It came with a mandatory 30-day data retention policy and premium pricing, and there was a planned shift to usage-based credits set for June 23. Fable didn’t survive long enough to see it.

The broader context makes the shutdown harder to read as a clean security decision. For months, the Trump administration had been dismantling AI regulations from the previous administration, approved advanced chip sales to China, and on June 2 issued an executive order asking AI labs to voluntarily share new models with the government before public release.

Then, ten days later, access was cut without notice, and the government body responsible for evaluating dangerous AI capabilities was ordered to stop publishing its reports. That’s a sharp reversal in a very short window. Europe is already paying attention, with concern growing that the same scenario could play out with Azure, AWS, Google, and every other US-based cloud provider.

The debate sparked by this decision has divided the cybersecurity community. On one hand, restricting access to offensive AI capabilities reduces the risk that cybercriminals, ransomware groups, or state actors could automate highly dangerous activities. On the other hand, the same constraints can hinder defenders, red teams, and security researchers who rely on such tools for testing and analysis.

Advanced AI models and geopolitics add another layer of complexity. Safety guardrails are not perfect: researchers have repeatedly shown that even robust systems can be bypassed through advanced prompt engineering. Security is therefore not static but an ongoing cycle between control mechanisms and attempts to circumvent them.

Beyond the technical dimension, the geopolitical aspect is even more significant. Project Glasswing illustrates how access to advanced AI-driven cybersecurity capabilities is becoming a strategic asset. Early participants were mainly US-based companies such as Microsoft, Google, Apple, Cisco, CrowdStrike, and NVIDIA, with European and other international actors included only later.

For Europe, this raises a critical issue. While it has developed strong regulatory frameworks through the AI Act, NIS2, and the Cyber Resilience Act, the most advanced AI systems are still built elsewhere. The continent regulates AI but does not yet control comparable frontier models within its own ecosystem.

Even participation by ENISA and other European bodies does not fully resolve dependency on external providers. The concern is no longer only data sovereignty, but also analytical sovereignty: advanced AI systems generate highly sensitive intelligence about critical infrastructure vulnerabilities.

For countries like Italy, this creates practical questions about where such data is stored, who can access it, and how it might be used. Since providers like Anthropic also collaborate with US government entities on national security issues, governance of this information becomes even more sensitive.

Ultimately, AI is reshaping cybersecurity from a human-driven discipline into a model-driven one. Competitive advantage will depend less on discovering vulnerabilities and more on managing them at scale. Those who master these systems early will gain a lasting strategic edge, while others risk increasing dependency on external technologies that are rapidly becoming central to global digital security.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, NSA)

❌
❌