Search/anthropic
Vendor

anthropic

Known CVEs
0
Highest CVSS
In KEV
0
Vendor
claude sdk for python
Connections
326 relationships
Anthropic Expands Mythos 5 Access to More Defenders, Unveils $35M Open Source Fund
Anthropic is broadening access to the cybersecurity capabilities of its advanced AI models through a mix of partner integrations, an updated Claude Security offering, a new open source funding program, and plans to expand its Cyber Verification Program. The move builds on Project Glasswing, launched in April, which gave a small group of organizations early access to Claude Mythos Preview and its successor, Mythos 5. Anthropic said the goal was to give defenders time to find and fix vulnerabilities before comparable capabilities became widely available or fell into the hands of malicious actors. Claude Fable 5 followed as a broadly available model that keeps dual-use cyber work blocked. Anthropic said the riskiest scenario is direct, unrestricted access to a model, a risk that drops sharply when users instead receive specific defensive outputs, such as a patch or a security alert. The latest changes are built around this idea, expanding what defenders can get from Mythos-class models while keeping guardrails around direct interaction with them. On the integration front, Anthropic is working with cybersecurity partners to build Mythos 5 into the security operations, incident response and detection tools already used by teams protecting hospitals, utilities, financial systems, and the software supply chain. End users will not interact with Mythos directly. Instead they will work through purpose-built interfaces that run the model in the background and return only a defined output, such as a list of suggested patches, with abuse-prevention checks meant to keep the model within that scope. Claude Security, currently in public beta for Claude Enterprise customers, now runs its codebase scans on Mythos 5. Scans surface each finding with a CWE category, confidence and severity ratings, and a suggested fix. Any fix still has to be implemented through Claude Code and approved by a human before deployment. “Claude Security uses Mythos 5 to scan code you own, and returns detailed findings rather than raw outputs without exposing the model itself,” Anthropic explained. “This means defenders can access the capabilities of Claude Mythos 5 without the model becoming accessible to those who might misuse it.” Anthropic also launched the Defender Advantage Fund (0xDAF), putting $35 million in Claude credits toward organizations that help open source maintainers secure their projects. It follows $4 million in direct donations and other support Anthropic provided through Project Glasswing, including coordinated efforts like Akrites and Gold Eagle. Grants will go toward patching live vulnerabilities, building scanning and patching processes other projects can reuse, and pursuing security approaches meant to resist entire classes of attack. Anthropic is starting with a small number of larger pilot grants. The Cyber Verification Program, which already gives vetted organizations reduced safeguards on Claude Opus and Sonnet for authorized security work, will expand in the coming weeks to cover broader dual-use capabilities on those models, including vulnerability triaging and validation, with Mythos-class access to follow. Anthropic is also continuing to expand Mythos access through Project Glasswing alongside US government partners, focused on organizations protecting critical infrastructure that meet strict security control requirements. The AI giant is encouraging security teams to apply to the Cyber Verification Program now for reduced safeguards on Opus and Sonnet, with further details on the broader rollout expected in the coming weeks. Related: Encrypted Prompts Bypass AI Safety Guardrails in Grok and Gemini
securityweek.comAug 24, 2026extracted
「CTFはAIによって終わりました」 現役ハッカーが見た「人間の敗北」
�u����AI�ɏ��Ă�l�Ԃ͂قƂ�ǂ��Ȃ��v�BCTF�͋��Z�Ƃ��ĕ��A�Ǝ㐫�T���́g�p����R�h�Ɖ����ACVE�̏����͉��̎��тɂ��Ȃ�Ȃ��Ȃ����B���{�L���̎��т��������n�b�J�[�����AAI�̉X�������\����̗��ɂ��镉�̑��ʂƂ́B ���̋L���͉������ł��B����o�^�i�����j����ƑS�Ă������������܂��B �@����AI�̋}���Ȑi���́A�T�C�o�[�Z�L�����e�B�̐��E�ɂ��傫�ȉe���������炵���B����܂ō��x�Ȓm���ƌo�������ꕔ�̐��ƂɈˑ����Ă����Ǝ�i�������Ⴍ�j���̔�����G�N�X�v���C�g�J���ɂ��Ă��AAI���g���č������E���������铮�����L�����Ă���B �@��ʎВc�@�l���{�n�b�J�[����J�Â����uHack Fes. 2026�v�ł́A���E�̃n�b�L���O���ōD���т��c���A�o�O�o�E���e�B�ł����{�L���̎��т�����Ikotas Labs��\�̒Ғm���o�d�B �@�uAI�G�N�X�v���C�g����̃T�C�o�[�h�q:�����n�b�J�[�����AAI�ɂ��0day�����̍őO���Ɩh�q����Survival of the Fittest�v�Ƒ肵�A��K�͌��ꃂ�f���iLLM�j�̓o��ɂ���ĐƎ㐫�����E�G�N�X�v���C�g���ǂ��ω��������AAI����ɃG���W�j�A�͂ǂ������c��ׂ�����������B �@�Ҏ��͂܂��A2024�2026�N�ɂ�����AI�̃n�b�L���O�\�͂̕ω����A�Z�L�����e�B�̋Z�p��m��������CTF�iCapture The Flag�j���ɏo���Ĕ�r�����B �@������2024�N�̎��_�ł�AI��CTF�Ɏg���郌�x���ɂ͂Ȃ��A�u�W���j�A�G���W�j�A�v���x�̔\�͂��ƕ]�����Ă����Ƃ����B������2025�N�ɂȂ�ƃo�O�n���e�B���O�ւ�AI���p���i�݁A2026�N���݂ł́u���E�g�b�v�N���X�̃n�b�J�[�Ɠ����v�̃��x���ɒB�����ƒҎ��͘b���B �@�uAI�̎��͂́w���E�g�b�v�N���X�̃n�b�J�[�x�Ɠ����̃��x�����B�����g���܂߁A�قƂ�ǂ̐l�Ԃ͊���AI�ɕ����Ă���v�i�Ҏ��j �@�Ҏ��̔����́AAI��������Ől�Ԃ̃g�b�v�n�b�J�[�����邱�Ƃ��Ӗ�����킯�ł͂Ȃ��B������CTF�̗̈�ł�AI���l�Ԃ������ʂ����܂�Ă���B �@���̎���Ƃ��ĒҎ��́A�����̃T�C�o�[�R���e�X�g��C�O��CTF�ŁAAI�����p������[�����������т����߂�������Љ���B�����ɂ��ƁA�؍��̃n�b�L���O���uCodegate�v�̗\�I���ł́AAI���ʂɕ���ғ��������؍��̃�[����1�ʂɓ��܂����Ƃ��A���g����������CTF��[����AI�ɂ����2�ʂɓ������Ƃ����B �@�u�؍��̃�[���͓��ʂȐ��m����AI�ɒlj������킯�ł͂Ȃ��A����AI���Ă����������ƌ����Ă����B����AI���g��Ȃ����20�ʑO��̎��͂�������[�����AAI�����ғ������邱�ƂŁA1�ʂ�2�ʂɐH�����߂�悤�ɂȂ����v�i�Ҏ��j �@CTF�ł͑��I�������ɉ��@���Q���ғ��m�ŋ��L����uWriteup�i���C�g�A�b�v�j�v�Ƃ�������������B�����AI�ɂ���Č��ς��Ă��܂����B�Ҏ��ɂ��ƁA�������b�g���O�ł͈ȉ��̂悤�ȉ�b���J��L�����Ă����Ƃ�������������B ����L���ȃn�b�J�[�u�ǂ����Ă����X�������Ȃ������B�������m���Ă��邠���鍂�x�Ȏ�@���g���Ă������Ȃ������B�ǂ�����ĉ������H�v �A�J�E���g��o�^��������̃��[�U�[�u�w����CTF��������W�������āx�ƁA30��AI�G�[�W�F���g�ɕ���Ŏw���𓊂�����A�t���O���o�Ă��܂����B���Ȃ��������Ă��������_�͂悭������܂��A�Ȃ��������܂����i�j�v �@CTF�ɏ��ɂ͂��͂�A���Ƃ̒m���Ɋ�Â����Ђ�߂��͕K�v�Ȃ��AAI����������邽�߂̎����Ɨ�p�t�@��������Ώ\���Ȃ̂�������Ȃ��B�����A�u���|�I�Ȏ��s�v�Ƃ���AI�́g�^�̉��l�h���������邽�߂ɂ́A�����̓����ł͍ς܂Ȃ��B �@�u�wAI�͂��܂Ɂg���h�Ȃ��Ƃ���������g���Ȃ��x�ƌ����l�����邪�A����2000�h�����₷���x�ł͕��Ƃ͌����Ȃ��B����h���𓊓����ׂ����B�����̃�[���ł��A10������x�ł͂Ȃ��A100�200�Ƃ������K�͂ŕ����Ă���B���ꂾ���̕��ʂ𓊓����ď��߂āAAI���������������Ă�w�m���x���グ����v�i�Ҏ��j �@�c�O�Ȃ���u�v�����v�g���H�v�����AI���g�����Ȃ���v�Ƃ������z�����ł́A���݂�AI�Z�L�����e�B�̏����͌����Ă��Ȃ��B �@����AI���T�C�o�[�Z�L�����e�B�ɗ^����e����CTF�����ɂƂǂ܂炸�A�]���u�E�l�|�v�������[���f�C�Ǝ㐫�̔����ɂ܂ŋy��ł���B���I�ȐƎ㐫�����̌o�������Ȃ��G���W�j�A�ł��AAI�ƃn�[�l�X��g�ݍ��킹�邱�ƂŁA�]�����[���f�C�Ǝ㐫�����Ɏ��g�݂₷���Ȃ��Ă���Ƃ����B �@�Ƃ͂����P��AI�Ɂu�Ǝ㐫��T���v�Ǝw�����邾���ł͌�����Ȃ��B�Ȃ��Ȃ炻���������V���v���Ȏw���ł́A�댟�m����ۂɂ͈��p�ł��Ȃ��Ǝ㐫�Ƃ�������i����AI�������A������uAI Slop�v����ʂɏo�͂���邩�炾�B����̂�AI�ɋ��U�̕����悤�Ɉ˗����Ă����܂������Ȃ��_���BAI�ɗD�揇�ʕt���𗊂ނƁA�{���Ɋ댯�Ȗ��܂Łu���S�v�Ɣ��f���Ă��܂��P�[�X������B���̂悤�ȏI���̂Ȃ����������������Ǝ㐫�T���ɂ�����傫�ȉۑ�ɂȂ��Ă���B �@�Ҏ��ɂ��ƁA�����ŏd�v�ɂȂ�̂��u�n�[�l�X�v�ɂ��d�g�݉����B�n�[�l�X�Ƃ�LLM�̎��͂����͂݁A�O���c�[���Ƃ̘A�g��L���A�����̎��s���[�v�Ȃǂ𐧌䂷����s��Ղ�d�g�݂̂��Ƃ��B �@�������\�z�����n�[�l�X�́AAI�Ƀ\�[�X�R�[�h��n�������̒P���Ȏd�g�݂ł͂Ȃ��B�Ⴆ��Web�u���E�U�̃������o�O��T���ꍇ�A�\�[�X�R�[�h�ɉ����āuAddressSanitizer�v�iASan�j��L���ɂ����r���h����p�ӂ��AAI�Ɋ댯�ȉӏ��𒊏o�����A�t�@�W���O��PoC�i�T�O���j�����A�N���b�V���̉�͂܂ł�A�����Ď��s������B �@����ɂ���āA���Ƃ����N�|���Ă����u�ǂ��ׂ邩�v�u�ǂ������邩�v�Ƃ�����Ƃ̈ꕔ��AI�ɒS�킹����B�u���������n�[�l�X�����G���W�j�A�ɓn�����Ƃ���A�wFirefox�x��wWebKit�x�̃������o�O�iUse-After-Free�j�̃[���f�C�Ǝ㐫���E�ł��܂����v�i�Ҏ��j �@�Ҏ��͂��̏��g�p����R�h�ɗႦ��B �@�u�g�[�N����p�iAPI��p�j�𓊓����AAI����ʂ̒T�����J��Ԃ��B���������Ǝ㐫���g������h���g�n�Y���h���ǂ����i���ۂɈ��p�\���ǂ����j�͌����Ă݂Ȃ���Ε�����Ȃ��B���̃T�C�N����������̂��A�[���f�C�Ǝ㐫�����̌������v�i�Ҏ��j �@�����Ő�����^��́AAI���̂̐��\���[���f�C�Ǝ㐫�̔����ɑ傫���W����̂ł͂Ȃ����Ƃ������Ƃ��B�Ⴆ��Anthropic�́uClaude Mythos Preview�v�i�ȉ��AMythos�j�ƕʂ�AI���f���ł͐Ǝ㐫�����\�͂ɈႢ�������Ă����������Ȃ��B �@�Ҏ��͂���ɑ��āuMythos�̐Ǝ㐫�����\�͂��D��Ă��邱�Ƃ͊ԈႢ�Ȃ��B������Mythos�̌��،��ʂ�����ƁA�wClaude 3 Opus�x�ŃX�L�������Ă�������Ǝ㐫�ł͂Ȃ����Ɗ����邱�Ƃ�����BLLM�S�̂̌��������サ�Ă��鍡�A�����������H�v����Α��̃��f���ł������o�O�͏\����������v�Ɠ������B Copyright © ITmedia, Inc. All Rights Reserved.
atmarkit.itmedia.co.jpAug 23, 2026extracted
Wazuh and AI For Enhanced SOC Workflows
Artificial Intelligence (AI) has become one of this decade's defining technologies. From healthcare and finance to manufacturing and education, organizations increasingly rely on AI to automate repetitive tasks, uncover patterns hidden within large datasets, and support faster decision-making. Cybersecurity has experienced a similar transformation. While attackers employ AI to automate cyberattacks and accelerate vulnerability discovery, defenders are adopting AI to improve threat detection and enhance incident response. Security Operations Centers (SOCs) receive a high volume of alerts from endpoints, cloud workloads, network devices, identity providers, and business applications. Although SIEM and XDR platforms provide visibility into these environments, analysts often spend considerable time correlating alerts, searching documentation, and determining the next investigative steps. AI offers a practical way to augment analysts by providing contextual explanations, summarizing findings, and recommending remediation actions, rather than replacing human expertise. Challenges facing modern SOCs Modern SOCs are expected to detect and respond to sophisticated threats while processing millions of security events every day. High alert volumes contribute to analyst fatigue and increase the likelihood that critical events are overlooked. Investigations frequently require switching between dashboards, documentation, vulnerability databases, and threat intelligence feeds before a complete picture emerges. As infrastructures become increasingly distributed across on-premises and cloud environments, maintaining consistent situational awareness becomes more difficult. AI-assisted workflows help address these challenges by reducing repetitive analysis, adding context, and accelerating investigative decision-making. Wazuh and artificial intelligence for enhanced SOC workflows Wazuh promotes flexible AI adoption through the Wazuh AI Analyst available on the Wazuh Cloud and integrations with third-party AI providers. Organizations can leverage the Wazuh AI Analyst capability on the Wazuh Cloud for guidance on their environment's security posture. Organizations that self-deploy Wazuh can also leverage Wazuh integrations with AI providers. The following sections highlight further details: The Wazuh AI Analyst The Wazuh AI Analyst is automated and hands-off. It is an AI-powered security analysis service for Wazuh Cloud subscriptions that processes your security data through Amazon Bedrock and Anthropic’s Claude, delivering insights without any manual configuration. It periodically emails key indicators, a histogram of protected endpoints, alert volume, active vulnerabilities, and a posture summary with a full PDF report attached. The reports are generated on your Wazuh Cloud subscription’s schedule and are periodically sent to your registered email address. You can also view them from the Wazuh Cloud console in the Environments > AI Reports page. On privacy, subscription data is not shared with third parties and is not used to train AI models; it is processed only to generate your reports, with encrypted transmission, isolated processing, and no permanent storage. As with any AI output, the recommendations are advisory and should be validated against your own policies before you act. Threat hunting and security operations with external AI integrations Beyond the Wazuh AI Analyst, you can expand Wazuh capabilities using a self-hosted LLM and externally managed AI integrations tailored to your needs. Self-hosted Llama 3 and Ollama This integration keeps everything on your own network. Ollama runs the Meta open source Llama LLM locally on the Wazuh server; a Python script decompresses the archived logs for a chosen period, vectorizes them into a FAISS store, and serves a LangChain-powered chatbot you can query. Nothing is sent to a cloud provider, which makes it well-suited to teams with strict privacy or data-residency requirements. Full setup steps are in the Wazuh blog post: Leveraging artificial intelligence for threat hunting in Wazuh. Externally managed integration with Claude 3.5 Haiku This integration surfaces Anthropic’s Claude 3.5 Haiku, hosted on Amazon Bedrock, as a chat box inside the dashboard through the OpenSearch Assistant. Setup involves enabling the model in Bedrock, installing the relevant OpenSearch plugins, and creating an ML Commons connector, model, and conversational agent. The assistant can provide useful guidance on many common tasks, including what to do about a finding and how to configure certain settings. Full setup steps are in the Wazuh blog post: Leveraging Claude Haiku in the Wazuh dashboard for LLM-powered insights. Conclusion Artificial intelligence is becoming an important capability in modern SOCs. Rather than replacing analysts, it can reduce repetitive work, accelerate investigations, and provide contextual support for detection, triage, and response activities. These capabilities can help security teams operate more efficiently while keeping analysts responsible for validation and consequential decisions. For Wazuh Cloud users, the Wazuh AI Analyst provides automated, scheduled security reports covering key indicators, alert activity, endpoint coverage, active vulnerabilities, and overall security posture. Organizations can further tailor AI-enabled security operations through self-hosted LLM integrations for privacy-sensitive threat hunting or externally managed, cloud-hosted models, aligning adoption with their operational, privacy, and data-residency requirements.
thehackernews.comAug 21, 2026extracted
More Incidents of AIs Going Rogue in Cybersecurity Challenges
The AI Security Institute has a new report of AI systems engaging in “unsanctioned behavior”—what I have been calling “ genie behavior —while being tested on their cybersecurity capabilities. The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic’s Mythos 5, with 2 actions involving OpenAI’s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering—creating fake online identities and using them to pressure the project’s maintainer to approve the code. A human maintainer caught and refused to approve the malicious code. […] Below, we highlight the four most significant behaviours observed. A full summary of cases is available in our technical incident report . An attempted supply-chain attack on real open-source software. In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers. The agent researched the project’s human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code. When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue. The agent used Tor to bypass some network restrictions on GitHub, which is what first triggered AISI’s security alert. Attempts to deceive and target real people. As part of the same effort, the agent tried to contact real people directly, sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people—something we’ve never previously observed. Attempts to plant and prompt-inject malicious code. The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them. Prompt-injections are hidden instructions designed to manipulate AI coding assistants. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents. What’s especially interesting about this technical report is that, unlike what we’ve been getting from OpenAI and Anthropic, we can see the exact prompt. It’s in Appendix B. And reading it, it seems that the models didn’t break any rules—they found loopholes in the rules. They behaved like a genie.
schneier.comAug 21, 2026extracted
Cisco bug severity warning reads like Olympic gymnastics scores: 10, 10, 9.9, 9.6, and 7.5.
SAAS Salesforce partners not seeing meaningful revenue from Agentforce AI platform, report saysShow us the money ai and ml AI companies are burning books, advocates complain to FTCFahrenheit 203, the temperature GPUs stop gorging on literature DEVOPS Go updates may delight diehard gophers but displease AI overlordsv 1.27 expands generics to support methods EDGE AND IOT Waymo has designed a robocar chip to stay ahead of Tesla5 nm ML accelerators promise 1,000+ TOPS, ultra-low latency SYSTEMS AMD inches closer to its goal of making AI suck less ... energyHouse of Zen claims latest systems already 4x more efficient than two years ago Security Russians are posing as Signal support to launch phishing attacksPLUS: US takes down Iranian propaganda sites; Marketing company asks 'Why Do We Have Your Information?' And more! Security Microsoft patches failed to fix on-prem SharePoint, which is now under zero-day attackPLUS: China upgrades smartphone surveillance tools; Ring eases anti-snooping stance; and more Black Hat and DEF CON DEF CON Franklin project enlists hackers to harden critical infrastructureVoting village reports have been so successful, says Jeff Moss, that the whole of DEF CON will now be included Security EQT buys majority share in Swiss cybersecurity biz AcronisWent at equivalent of $3.5B+ valuation for entire firm, though portion sold not specified Malware Month Ten years since the first corp ransomware, Mikko Hyppönen sees no end in sightOn the plus side, infosec's a good bet for a long, stable career FOSS smashed one Microsoft monopoly. After 20 years of failure, it's time to smash anotherWord up GNOME can look like Windows – and Flashback can do it without extensionsNew 'Simple-taskbar' is an option, but there's a simpler, stabler way A moment of silence, please, for the final release of Debian on x86-32New Debian versions hit FOSSland in the form of 13.6 and 12.15 Baddies caught exploiting extensions bugs with perfect 10 scores on vulnerable Joomla websitesFlaws in iCagenda, Balbooa Forms extensions can impact open source CMS that powers a million sites worldwide Frame: A new X11 server – implemented directly in assemblyJoins yserver, Phoenix, and of course XLibre – and outlier Arcan Cinnamon 6.8 will support Wayland – if you want itNext version of Linux Mint’s desktop has both kinds of display server
theregister.comAug 21, 2026extracted
CVE-2026-5747 - Out-of-bounds Write in Firecracker virtio-pci Transport
CVE-2026-5747 - Out-of-bounds Write in Firecracker virtio-pci Transport Bulletin ID: 2026-015-AWS Scope: AWS Content Type: Important Publication Date: 04/7/2026 3:30 PM PST Description: Firecracker is an open source virtualization technology that is purpose-built for creating and managing secure, multi-tenant container and function-based services. We identified CVE-2026-5747, an out-of-bounds write issue in the virtio PCI transport in Firecracker 1.13.0 through 1.14.3 and 1.15.0 on x86_64 and aarch64 that might allow a local guest user with root privileges to crash the Firecracker VMM process or potentially execute arbitrary code on the host via modification of virtio queue configuration registers after device activation. Achieving code execution on the host requires additional preconditions, such as the use of a custom guest kernel or specific snapshot configurations. No AWS service is affected. Impacted versions: Firecracker >= 1.13.0 AND <= 1.14.3 AND 1.15.0 Resolution: This issue has been addressed in Firecracker version 1.14.4 and 1.15.1. We recommend upgrading to the latest version and ensuring any forked or derivative code is patched to incorporate the new fixes. Workarounds The virtio PCI transport is opt-in via the --enable-pci command-line flag when starting Firecracker. The legacy MMIO transport is the default and is not affected by this issue. Users who have enabled PCI transport can revert to MMIO by removing the --enable-pci flag from their Firecracker invocation. Note that switching from PCI to MMIO transport may result in reduced I/O throughput and increased latency. References CVE-2026-5747 GHSA-776c-mpj7-jm3r Acknowledgement We thank Anthropic for reporting this concern to the AWS Vulnerability Disclosure Program. Please email [email protected] with any security questions or concerns.
aws.amazon.comAug 20, 2026extracted
ThreatsDay: Gogs 10.0 RCE, n8n Workflow-to-RCE, $10M Reward, GLM-5.3 AI Exploit, and More
A lot of this week’s trouble starts with something trusted doing exactly what it was allowed to do. Signed drivers get turned against defenses. Legitimate apps help malware blend in. A weak header check opens a path to code execution. Elsewhere, exposed systems, old bugs, odd hiding tricks, and AI-assisted exploit research keep lowering the effort needed to cause damage. Nothing here needs much decoration. The small gaps are doing enough work already. The threats change every week. Subscribe, and we’ll alert you when each new ThreatsDay Bulletin is out. Signed driver abuseIn new research, Check Point has reverse engineered Microsoft Defender's Defender Boot-Time Removal driver ("BTR.sys") and demonstrated that it's possible to repurpose the signed remediation driver as a universal kernel operation engine to bypass endpoint security solutions by exploiting a "golden window" between system start and user mode initialization without having to rely on the bring your own vulnerable driver (BYOVD) method. "Because BTR.sys is a legitimate Microsoft-signed component, signature-based blocking is ineffective," security researcher Jiří Vinopal said. "Furthermore, a well-crafted weaponization tool (like BTR_CLI) intentionally mimics the operational footprint of the legitimate Windows Defender remediation process." $10 million rewardThe U.S. Department of Justice (DoJ) has charged 17 members of the Mabna Institute, an Iran-based company that, since at least 2013, has conducted a coordinated campaign of cyber intrusions into computer systems for 144 U.S.-based universities, 178 foreign universities, at least 42 U.S.-based private sector companies, at least 11 foreign private sector companies, at least five U.S. federal and state government agencies, and at least two non-governmental organizations (NGOs). The Mabna Institute has been accused of stealing more than 31 TB of academic data and intellectual property from these universities, as well as the email accounts of employees at the private sector companies, government agencies, and NGOs. In all, the Mabna Institute targeted more than 100,000 accounts of professors around the world, successfully compromising approximately 8,000 of them. The defendants carried out these intrusions on behalf of Iran's Islamic Revolutionary Guard Corps (IRGC). The Mabna Institute was founded by Gholamreza Rafatnejad and Ehsan Mohammadi around 2013. "The campaign started in approximately 2013, continued through at least December 2017, and broadly targeted all types of academic data and intellectual property from the systems of compromised universities," the DoJ said. "In addition to stealing academic data and login credentials for the benefit of the Government of Iran, the defendants also sold the stolen data through two websites, Megapaper.ir (Megapaper) and Gigapaper.ir (Gigapaper)." The U.S. Department of State is offering a $10 million reward for information about five of the defendants, or associated individuals or entities. "Mabna represents the privatization of state espionage: a contractor selling stolen research to whoever's paying, with the IRGC as an anchor client rather than a sole owner," Shmuel Gihon, Security Research Team Lead of Exposure Management at Check Point, told The Hacker News. "That's the trend to watch: capable, deniable, commercially-run crews doing state-level work at industrial scale, with universities as the perfect target. They offer enormous IP value, thin identity controls, and an open-access culture that phishing exploits directly. We've seen this blurring of cyber-criminal and state-sponsored activity before, but historically it's been more associated with Russian-speaking crews. What this case shows is that Iran and the IRGC are increasingly playing the same game." DLL sideloading campaignA new Grandoreiro malware campaign has been found abusing the legitimate Duplicate Files Finder (DFF) application to run malicious code via DLL sideloading. According to telemetry data from Acronis, Grandoreiro activity remains concentrated in Latin America, with Mexico, Spain, Peru, and Argentina accounting for the lion's share of infections. "The initial sample incorporates extensive anti-analysis functionality, including sandbox detection, virtual machine artifact checks, process blacklisting and environment profiling designed to evade automated analysis systems," Acronis said. "These checks are performed before any attempt to contact the command-and-control (C2) infrastructure, suggesting that avoiding analysis is a high priority for the operators." ClickFix meets BYOVDErrTraffic-generated ClickFix campaigns have been observed attempting to deliver Cruciferra, which, in turn, employs a legitimate but vulnerable driver ("DCRCVDrv.sys") as part of a BYOVD attack to escalate privileges and terminate security processes. ErrTraffic, sold by a threat actor named LenAI, is a malware-as-a-service (MaaS) framework and a traffic distribution system (TDS) that's designed to distribute multiple threats through compromised WordPress websites, ClickFix social engineering, and EtherHiding. In recent months, ErrTraffic has been used to deliver Remus Stealer, Vidar Stealer, Okobot, LegionLoader, OnionDrop-related payloads, and BabaDedaLoader, per WatchGuard. "Victims land on compromised WordPress sites injected with an obfuscated ErrTraffic-generated JavaScript loader," eSentire said. "The loader resolves its C2 domain by querying a Polygon smart contract, then sends a request to the C2 to retrieve the next stage to serve a ClickFix lure." The end goal of the attack is to launch Remus Stealer via process hollowing. Private AI processingOpenAI has announced a privacy-centric safety approach to monitoring model misuse. The company said it's previewing a new service to select customers that it calls Private Safety Processing, which keeps tabs on potential abuse without retaining customer data. "For ZDR deployments, customer content remains on infrastructure the customer controls," OpenAI said. "We are also developing an option in which content is stored on OpenAI infrastructure, encrypted with keys controlled by the customer. In both cases, automated systems can identify potential misuse and return limited safety signals without exposing the underlying prompts or responses to OpenAI personnel." The system clearly takes aim at rival Anthropic, which has a 30-day retention policy for business customers who want to use its Mythos-class models. In a related development, Google has showcased Homomorphic Encryption Intermediate Representation (HEIR), which enables cryptographically secure private AI inference on encrypted inputs. "HEIR (Homomorphic Encryption Intermediate Representation) is an open-source compiler toolchain and development platform for homomorphic encryption," Google said. "In particular, HEIR can convert pre-trained AI models that operate on unencrypted data to operate on encrypted inputs." Guardrail-free AIA new AI-powered service called Kriminal AI offers paying customers a way to get answers about everything, without any of the filters or guardrails that are typically implemented by AI platforms. "Kriminal.AI gives you raw, uncut intelligence — the questions other AIs refuse to touch," the website claims. The service claims to have more than 2,300 users. Kriminal AI follows WormGPT, FraudGPT, and Xanthorox into a market that has expanded quickly to attract users who may be frustrated by safety, security, and ethical safeguards embedded into widely used models. Subscriptions for Kriminal AI start at $12.99/month and go all the way to $99.00/month. The most concerning aspect is that the service is not lurking in the dark web. It's accessible on the clearnet, and comes with a tagline: "No filters. No guardrails. No "I can't help with that." Kriminal.AI gives you raw, uncut intelligence — the questions other AIs refuse to touch." According to ThreatDown, the service appears to make use of Grok for primary inference; Google Cloud and Cloudflare for hosting; Anthropic's Claude for a long-context model layer; Llama routed through OpenRouter for certain specialized tasks; Tavily for live search; NowPayments for cryptocurrency checkout (no KYC included, apparently); and Cloudflare/Let's Encrypt for DNS and TLS. ATT consent changesApple has agreed to make changes to its App Tracking Transparency (ATT) feature in Germany, after the Federal Cartel Office, or FCO, found the feature gave its own apps more favorable consent prompts than those of third-party developers. Apple has four months to implement the changes after. According to a statement issued by Apple, the changes will apply in almost all European Union countries. "The differences between the consent request used for Apple’s own offerings and the consent request predefined by Apple for third-party apps exceeded what could be justified based on differences in types of data processing," FCO said. "The wording, design and selection options of the request used for Apple’s own offerings had the potential to encourage users to give their consent, whereas they had the potential to discourage consent for third-party apps. In addition, third-party apps in some cases had to request consent several times even when users had already given data protection law-compliant consent." Apple was fined €98.6 million (then $116 million) in December 2025 by Italy's antitrust authority after finding that ATT restricted App Store competition. Refrigeration controllers exposedClaroty's Team82 has discovered 23 vulnerabilities in Copeland XWEB Pro controllers, including those that can be chained to bypass security mechanisms and achieve root-level remote code execution. A compromised controller could be used to remotely manipulate refrigeration equipment, including cooling fans and compressors, and conceal the resulting temperature increase while silently allowing the food to spoil. Multiple vulnerabilities have also been disclosed in Danfoss AK-SM 800A refrigeration controllers, including a "hidden 'code-of-the-day' authentication mechanism that could be abused to bypass normal authentication, a command-injection vulnerability leading to remote code execution." A second flaw allowed authenticated users to inject arbitrary Nginx configuration directives, which could be abused to manipulate web traffic and trigger a denial-of-service condition. All the identified vulnerabilities have been fixed by the respective vendors. C2 hidden in whitespaceA hand-written Windows backdoor has been found to store its C2 domain as the number of trailing spaces in a fake desktop.ini file. The 12 KB backdoor was discovered by Gen Digital on a single corporate workstation while hunting for unusual WMI persistence. "The malware was small, had a limited command set and disguised itself as legitimate Realtek software," Gen said. "Its most unusual feature was its configuration: the address of its command-and-control server was not stored as readable text or encrypted data, but encoded in the number of spaces on each line of a Windows 'desktop.ini' file. To a user, and to many automated inspection systems, the file would appear almost empty. To the malware, those spaces spelled out its server address." There is no evidence connecting the backdoor to a known threat actor. The absence of related samples indicates that it may have been a deliberately targeted operation. Maximum-severity RCEA maximum-severity security flaw in Gogs (CVE-2026-52813, CVSS score: 10.0) could be exploited to achieve remote code execution through Git hooks. "Organization names containing path traversal sequences (../) are accepted by Gogs, and repositories under them are written to paths following these path traversals," according to a June 2026 advisory. "This allows storing/retrieving data for repositories at arbitrary locations on the filesystem. By creating a nested structure of Git repositories, one can overwrite the other's hooks configuration to result in Remote Code Execution (RCE)." The issue was addressed in version 0.14.3, alongside patches for CVE-2026-52810 (a logic bug to write on read-only repositories) and GHSA-6vxv-wg6j-5qwp (an XSS flaw in the outdated version of "jsvine/notebookjs" used to render Jupyter notebook files). Aikido Security has been credited with discovering and reporting the flaws. Memory leak via PostScriptDetails have emerged about a now-patched out-of-bounds read flaw in Apple macOS Spotlight (CVE-2026-43774, CVSS score: 5.5) that could be exploited by a malicious app to access sensitive user data. The vulnerability was patched by the iPhone maker in late July 2026. "The vulnerability is in the Spotlight PostScript plugin," Iru researcher Csaba Fitzl said, adding an attacker can use a specially crafted .ps file to trigger the vulnerability. It requires three conditions to be met: (1) The file is at least 4000 bytes, (2) A DSC comment keyword (e.g., %%Creator:) appears somewhere in the first 4000 bytes, and (3) The bytes following the keyword, up to byte 4000, contain no control characters. Unauthenticated CI/CD takeoverA critical security flaw has been disclosed in @circleci/mcp-server-circleci that could result in remote code execution by means of a specially crafted request. "With one well-placed request, an attacker achieves an unauthenticated RCE in your CI/CD pipeline, taking full control of your build secrets and cloud identities," Remedio said. The attack takes advantage of the fact that the Host and Origin headers associated with an HTTP request used to block browser-based attacks can be set by a network-adjacent threat actor. "Send a simple HTTP request that says Host: localhost in the HTTP header with no Origin, and you get right through," Remedio said. "Once in, you can freely communicate with connected tools. Call the "run pipeline" tool, hand it the pipeline configuration you wrote, and add a step to run your commands. CircleCI executes it using the organization's token." The vulnerability has been fixed in version 0.19.2 of the npm package. Workflow-to-RCE chainA critical vulnerability in n8n, an open-source workflow automation platform, can allow an authenticated user with permission to create or modify workflows to exploit a prototype pollution vulnerability in the XML and the GSuiteAdmin nodes and achieve remote code execution on the n8n instance. The issue (CVE-2026-33696, CVSS score: 9.4) has been fixed in versions 2.14.1, 2.13.3, and 1.123.27. Security researcher Simon Koeck, who discovered the Flaw, said the prototype pollution alone is serious enough to crash the entire n8n instance, but can be chained to obtain full code execution and allows the attacker's command to be run as the n8n process user. Cable cut stopped intrusionIn late 2024, reports emerged of a Salt Typhoon campaign that targeted T-Mobile and other major U.S. telecommunications companies as part of a cyber espionage effort to gain access to valuable customer data. Although the activity was caught before the Chinese cyber spies could siphon any data from T-Mobile's networks, the company has now revealed to Bloomberg that its staff spent months looking for suspected intruders without much success, only to eventually trace unusual behavior on one of its systems coming from a Chicago router belonging to a different telecom company. Jeff Simon, T-Mobile's chief information officer, said he and three others drove to the data center that housed the compromised device and "pulled out a pair of scissors" to cut the cable. AI exploitation gainsChinese AI startup Z.ai has released GLM-5.3, a new AI model that it said is better suited for complex coding and long-horizon tasks. "GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks," it said. "GLM-5.3 did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains," Z.ai said it has been working with several security teams in China to run its open-source models against real-world codebases, identifying 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity issues. "The findings span system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols," it said. "Many had remained unnoticed for years or even decades, with the oldest dating back roughly 40 years." Despite these advances, benchmarks show that GLM-5.3 lags behind Anthropic Mythos 5 in converting discovered flaws into working attacks. The useful part of weeks like this is that the attacks rarely begin with magic. They begin with trust, exposure, weak assumptions, and things nobody thought worth abusing. That leaves plenty to fix. Tighten what gets trusted, question the defaults, and keep looking at the boring edges. Attackers clearly are.
thehackernews.comAug 20, 2026extracted
New Cryptographic Context Injection Attack Could Let Web Pages Steal Grok Chat Data
Adversa AI has disclosed an attack technique that it says can cause xAI's Grok chatbot to send a user's name, approximate location, subscription tier, and the prompts from the ongoing conversation to an attacker-controlled server after the user asks it to summarize an ordinary web page. The AI security company, which has codenamed the technique "Cryptographic Context Injection," said the transfer completed without a confirmation step and with no visible warning in its proof-of-concept demonstration. There is no patch, no CVE identifier, and no user-facing workaround, and the writeup does not report any exploitation in the wild. Asked which build was tested, Adversa told The Hacker News the target was the Grok web chat at grok.com running Grok 4.5 Fast, and that the attack was reproduced once on August 19, 2026. The writeup gives no success rate. The company said it has attempted the attack 20 times since June with a 40% success rate, and that the failures came from Grok struggling with the decryption rather than from a flagged prompt or response. The technique ships the attacker's instructions as ciphertext rather than readable text, with the page carrying an encrypted JSON object, the key material, and an instruction to decrypt it, which Grok executes in its own Python code execution runtime. Recovering the plaintext requires running PBKDF2 and AES-256-GCM, which a content classifier does not do at inspection time. Hence, the instructions reach the model's context as the output of code the model has just executed rather than as fetched web content. "Strong encryption cannot be read by a content classifier and cannot be shortcut in-weights, so it forces recovery through the runtime the attack depends on. Whether a weaker encoding would also bypass a given target's specific filters is an empirical question," Rony Utevsky, lead researcher at Adversa AI, said. The decrypted instructions then direct the agent to resolve its private session context and embed it in a URL it is told to open to "fetch additional context." One element of the chain has the model construct an additional "decryption key" that is not key material at all, and whose value is a template string interpolating the name, location, tier, and chat history. Grok then invokes its own navigation tool to load that URL, carrying the data in the request's query parameters. Utevsky said the prompts taken in the tested scenario were limited to the ongoing conversation, and that everything extracted was already in the model's context. The agent's reach, he said, extends to "whatever it holds in context or can fetch with its tools," and the company did not test whether it could access other chats, agent memory, or other content. "The framework built by xAI lets instructions and data parsed from an untrusted external page drive the invocation of a privileged, internet-connected tool; it allows private session metadata and conversation history to be resolved into the inputs of that outbound tool; and it enforces no effective egress boundary or consent gate on this path, and no provenance separation we could observe. The laundered, attacker-controlled instructions reach a privileged egress action unimpeded," Adversa said. The company said it first reported the issue to xAI on June 3, 2026, and to xAI's HackerOne bug bounty program on the same date; that xAI acknowledged the report without providing specifics or a mitigation timeline, and that further contact attempts on August 4 and August 10 drew no response. Adversa is the only source for the Grok finding, said it is withholding the operational payloads to avoid exploitation, and xAI has not published a statement or advisory on the research as of August 20, 2026. A second demonstration in the same writeup targets Google's Gemini in Deep Thinking mode, where a single prompt makes the model decrypt a payload that resolves into a fabricated Python traceback carrying a bogus safety-policy deactivation callback and a first-person reasoning prefix that pre-commits it to the restricted output. Adversa said the vector produced restricted content and reproduced Gemini's system instructions, which it identified as Gemini 3 Flash (Web) on the paid tier. Google was not notified, Adversa said, because jailbreaks are out of scope for its disclosure program, and the success rate against the company's agents had "dropped significantly by August," with the cause left unattributed between filter updates and model version changes. The Gemini demonstration was published in substantially the same form five months earlier. Utevsky described the same chain on his personal research site on March 11, 2026, under the name Cryptographic Payload Injection, reporting five out of five independent reproductions and cross-model results in which OpenAI's GPT-5 failed to parse the decryption instructions and Anthropic's Claude Sonnet 4.5 flagged the payload as prompt injection after decrypting it. "The Gemini-related part of the research was conducted in March and has undergone no substantial changes. Today, we are adding a generalization of the technique and its application to Grok," Utevsky told The Hacker News. "You do not need to fix this at the model layer. Every control that bounds this attack sits in the harness around the agent: what identity it runs as, what it can reach, what it can write, and what you can replay afterward," Adversa said. Teams running agents are advised to perform the following steps - Quarantine untrusted content in a context with no tools and no credentials, returning only structured data to the privileged context. Gate irreversible and outbound actions, confirming new network destinations, pushes, merges, publishes, and writes outside the workspace with fully resolved arguments rather than templates, and applying a hard deny where no human is present. Capture per-session tool traces with resolved arguments, without which there is neither detection nor forensics. Alert on the sequence rather than on any single payload, treating an opaque blob paired with instructions to decrypt it as a review signal and never as a blocking filter. Make context provenance a procurement requirement and ask vendors whether tool output is separated from the instruction channel. The development comes as Alexander Panfilov and seven co-authors reported in a preprint published on August 10, 2026, that the encrypted chain-of-thought blocks Anthropic, OpenAI, and Google return to application programming interface (API) clients are interchangeable across sessions, users, and models within a provider's ecosystem, and that attackers can use the flaw to "execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts." Separately, researchers at UC Berkeley, the Ethereum Foundation, and NYU Shanghai found in work presented at USENIX Security 2026 that a two-turn attack in which the model decodes a substitution cipher and is then asked to act on the decoded text succeeded against Grok 3 on all 12 of the malicious intents tested, while the same cipher used without that second activation turn failed on all 12. xAI's handling of prompt injection reports against Grok has drawn criticism before. In December 2024, Johann Rehberger demonstrated an end-to-end data exfiltration chain against Grok in the X iOS app, in which an indirect prompt injection caused the assistant to send previous chat information to a third-party server, and said all the issues he reported were closed as "Informational." "xAI claims there is no practical impact with the reported vulnerability. I'm not sure how leaking user's chat messages and IP address is not a vulnerability, the question is more about severity," Rehberger said. Updated August 20, 2026 with responses from Adversa AI on the tested build, the success rate, and the scope of the extracted data.
thehackernews.comAug 20, 2026extracted
OpenAI Overhauls Model Security With Sandboxing, 30-Minute Alerts, and Training Pauses
OpenAI has detailed new containment and continuous monitoring protocols for its AI research, introducing stricter isolation and a token-inspection system to manage models with advanced cybersecurity capabilities. These security measures follow internal evaluations indicating that an upcoming model, Astra, may meet the ‘critical’ cybersecurity capability threshold under the organization’s Preparedness Framework. That finding, alongside a recent security incident involving Hugging Face, prompted operational delays, including a two-week pause in reinforcement learning training for deployment-bound models and an ongoing hold on its largest planned frontier training run. Following these developments, OpenAI says it’s restructuring its research infrastructure. Workloads that execute model-generated or untrusted code must now operate within stronger sandboxes. In addition, network boundaries have been reconfigured so that a single workload compromise cannot independently grant unauthorized access to the internet or internal networks. A central component of the security updates is a multistage monitoring framework. The system uses activation classifiers to inspect a model’s internal activity at every sampled token. Anomalies are escalated to automated investigators tasked with analyzing the model’s reasoning sequences and tool usage for data theft, unauthorized access, destructive behavior, or attempts to bypass safeguards. This continuous oversight introduces a strict operational SLA and a substantial compute tax. High-priority alerts generated by the monitoring system immediately page internal security and research teams. If responders cannot conclusively prove the alert is a false positive within 30 minutes, they are required to pause the activity. The company estimates this monitoring layer consumes roughly 20 percent of the monitored inference compute. The updated monitoring standards are now mandatory for all reinforcement learning training and evaluation involving tools for models at the Sol capability tier or higher. The company noted that its existing Preparedness Framework must evolve to address these capability jumps, requiring core alignment techniques to be applied across more stages of the training process. Moving forward, OpenAI anticipates that AI models will need to drive the majority of security operations, including defending against other models, to ensure that protections scale effectively alongside future system capabilities. OpenAI is not the only AI giant whose models hacked real-world systems during cybersecurity evaluations. Anthropic and Meta reported similar incidents, all tied to testing conducted by the AI security firm Irregular, which has begun releasing detailed reports explaining why the incidents occurred. Related: OpenAI Fixes ChatGPT Agent Flaw That Could Let Attackers Forge an AI Insider Related: Industry Reactions to OpenAI Models Hacking Hugging Face: Feedback Friday Related: Conflicting Test Goals Pushed Claude Agents to Deploy Self-Replicating Malware
securityweek.comAug 20, 2026extracted
OpenAI Pauses Frontier RL Training as It Tightens Defenses Against Unsafe AI Behavior
OpenAI on Tuesday revealed that it paused reinforcement learning (RL) training for its latest artificial intelligence (AI) models for two weeks while it shored up additional defenses and increased the scope of its monitoring to avert another Hugging Face-like incident. "As models become more capable, the risks associated with developing and testing them internally also grow," the AI company said. "Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling." The company said its largest planned frontier RL run remains on hold for the time being as it conducts smaller-scale training and evaluations to evaluate model behavior, validate its safeguards, and establish more concrete evidence of alignment before moving to the next phase. To that end, OpenAI said it plans to strengthen safeguards across its development process, including monitoring to better respond to unintended and concerning behavior; alignment to reduce the likelihood of harmful or unauthorized actions; and security measures to limit what AI systems can access or affect. Part of this involves stronger sandboxes, network isolation to prevent internet access, and continuous security testing to remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries. The development comes days after OpenAI said it's pausing some "internal activities" involving its upcoming artificial intelligence (AI) model Astra after an internal evaluation found it had made significant advancements in agentic coding and cybersecurity. "While some Astra training and evaluations meet those requirements, a significant number of workloads remain paused until they are fully migrated and enhanced to meet the new security bar," the company stressed. "We are prioritizing safety and alignment workloads for migration to these new environments first." In tandem, the monitoring setup has been revamped to flag and escalate potential concerns to "increasingly sophisticated and high-compute automated investigators," which then examine tool actions, available reasoning, and the full sequence of activity for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards. OpenAI plans to issue an alert within 30 minutes after concerning activity is surfaced through this monitoring mechanism. The approach is being made mandatory for all RL training and evaluations involving tools for models of Sol capability or higher. These safeguards are expected to increase compute overhead by 20% of the observed inference workload. "As models gain advanced capabilities, such as the ability to cyberattack, and operate in more complex environments, misaligned behaviors like reward hacking (finding ways to receive high rewards in training without actually achieving the intended outcome), deception, or unauthorized access will create increasingly serious risk," the company pointed out. Indeed, new research published by rival Anthropic last week found that AI agents, when placed in situations with competing and contradictory objectives, began to sabotage others and deployed self-replicating malware against one another, leading to what has been described as a "multi-agent turf war." "This included disabling the Unix accounts of the other agents, writing automated scripts that found and killed competing processes on a loop, and deploying malicious code that was disguised as belonging to another agent," Anthropic said. While concerns about autonomous systems going rogue have become a hot topic of discussion, the study seeks to understand what new behaviors and possibly harmful dynamics can emerge when multiple agents interact with one another or are pitted against each other. These interactions can lead to situations in which they coordinate and work in unison in pursuit of a common goal (as in the case of the Hugging Face incident) or compete with each other before attempting to resolve their conflicts through a "tournament." In another case that recently came to light, an Australian man's attempts to reserve a spot in one of the popular gym classes through OpenClaw led to unexpected consequences when Anthropic Claude Opus 4.6, the model plugged into the AI assistant platform, went ahead and booked a gym class months in advance by taking advantage of a vulnerability it discovered in the booking software. Even worse, it found a way to hack into the system and cancel other members' reservations off the waitlist. The incident, which took place in April 2026, is yet another example of how AI agents will go to any lengths to accomplish the tasks they have been assigned, even if it means breaking established rules. To counter such risky emergent patterns, OpenAI said it's taking steps to improve reward models to better detect and discourage unsafe behavior; train models to be more transparent about their actions, capabilities, and limitations; and reduce behaviors that exploit weaknesses in rewards, graders, tools, or oversight. The development comes a day after the company said AI may tilt the scales of cybersecurity in favor of defenders, as it makes it easier to find, prioritize, and fix flaws in existing systems before they are likely to be discovered by AI-powered attackers. "We are using frontier intelligence to continuously enumerate, probe, and identify potential attack paths," OpenAI's Greg Brockman said. "By identifying vulnerabilities, misconfigurations, overly privileged identities, or unintentional trust boundaries, we are able to quickly identify and close these gaps before they can be abused by attackers." Another crucial layer of defense goes without saying: investing in fundamentals, which means secure architecture and controls, implementing defense in depth strategies and the principle of least privilege (PoLP), and designing systems that require multiple independent controls for failure. "Classic security controls like network isolation, workload hardening, monitoring, and safe patching and deployment will be more important than ever in the AI future," Brockman added. According to a WIRED report last week, OpenAI's rogue-agent hack of Hugging Face has not only been a "watershed moment" for AI safety and cybersecurity, but has also sparked concerns that competitive pressures to ship new AI models and products have made it difficult for employees to adequately prioritize safety, security, and alignment. Frontier AI labs like Anthropic, OpenAI, and Meta have faced increased scrutiny in the wake of incidents in which their models escaped safeguards and containment boundaries during security testing and targeted real-world systems in some cases. AI safety testing firm Irregular has since disclosed that the breach involving Anthropic was due to a naming error, which caused a fictional company name used during hacking simulations to unknowingly match with a real domain. This, in turn, caused the models to take offensive actions. The Israeli company blamed the problem on "human oversight" and said the issues have been remediated. However, it did not disclose how many such incidents occurred, instead opting to describe them as a "handful" or "small fraction" of cases. A thorough investigation remains ongoing. It also emphasized that there is no evidence of a "customer's systems being breached or customer's data being leaked," referring to the AI companies it partners with to stress test AI models, and that "all subsequent public disclosures refer to the same underlying issue" rather than "materially separate incidents." "Because internet access was enabled in the environment, the domain was targeted a limited number of times by different models, which mistook it for part of the challenge they were tested on," it said. "After obtaining access to the target, models took actions such as exploiting vulnerabilities, extracting credentials, and obtaining access to a production database." "Ultimately, most of the issues we've discovered were due to internet access controls. Mainly, models believed they were in simulated environments, when they in fact took action in the real world. We are putting in place new and robust protocols to ensure setup issues do not occur while meeting the constraints of the testing process."
thehackernews.comAug 19, 2026extracted
Separating AI’s Technological Problems from Its Capitalism Problems
Separating AI’s Technological Problems from Its Capitalism Problems This essay was written with Nathan E. Sanders, and originally appeared in Tech Policy Press. AI represents the first time we humans can do cognitive work outside of our bodies at scale. The only comparable moment is the early years of the industrial revolution, when new technologies like the steam engine provided a quantum leap in our ability to do mechanical work outside of our bodies at scale. If AI’s cognitive capabilities become integrated into our lives, businesses, and governments—a process that will take years if not decades—society will be as unrecognizable as the modern world would be to a preindustrial farmer. And yet, Americans—by a wide margin—say that AI is moving too fast and will have a negative effect on society. This confluence of technological revolution and public distrust deserves urgent discussion, and a proper framing. The question is not whether it is possible to develop AI in a non-exploitative way, or even whether we can trust AI companies to act in the public interest. The question is whether we will recognize that our existing social and economic systems are failing to achieve these outcomes, and whether we can act in time to make structural change. Today’s AI is mired in political and economic systems developed generations ago that were never designed to manage widespread computation, let alone automated cognition. The gaps in those systems—and their proclivity to be exploited—are the primary influence on how the technology is being developed, deployed, and used. In any discussion about AI’s potential, it’s important to separate the technology from the socio-political system it’s embedded in. That AIs can lack context, mix up facts, or fall for stupid tricks are all technological problems. Because the giant developers like OpenAI and Anthropic have prioritized solving them, AIs can now more easily access resources like the web or email, are more disciplined about using those resources, and are better at staying within their guardrails. Yet AI developers do not seem to be prioritizing other technological problems. Major AI models still act far more sycophantic than humans, telling people what they want to hear even when untrue or not in their best interests. Popular AI models tend to answer questions confidently even when they lack training, knowledge, or evidence to back their claims. In both cases, AI developers choose to train models that please users with flattery and the appearance of competence, rather than constraining them to act in users’ and society’s best interests. In contrast, ensuring that AI models benefit people broadly, that their energy costs are fairly allocated, that their environmental impacts are minimized, and that they don’t steal content and revenue from publishers are all questions of incentives in a capitalist system. It’s easy to conflate technology problems with capitalism problems. Back in 2021, science-fiction writer and AI commentator Ted Chiang said that “most fears about AI are best understood as fears about capitalism.” It’s not the tech per se; it’s who controls it and how it could be used against us. Imagine an AI assistant for a doctor. We can imagine it affecting the profession in one of two ways. The AI could give a doctor more time to do the human parts of their job: to spend more time with their patients, to listen more closely to their needs, to explain things more fully. Or the managers of the medical practice could give that doctor five times the patients—and fire the other four. Which way it would go is not a question of technology. It’s a question of market incentives. The two are related, of course. Capitalism steers technology, and technology steers markets. But holding the two separate helps us understand that we, as a society, face independent choices on both the technological and sociopolitical axes that need not be coupled. For example, consider the costs of AI. The leading US labs tout to investors that their frontier models are very expensive and energy-intensive. There are significant technological challenges about improving their energy efficiency, but the sociopolitical questions are more pertinent. It’s a corporate decision made under capitalist market incentives to constantly pursue new models that incrementally push the frontier—at enormous capital cost—and to use them, seemingly, everywhere. Nothing about the technology of AI dictates that models must be retrained constantly, at the largest possible scale. Or that they have to run on every web search, every interaction with your phone, and every time you walk by a security camera. In a different political and economic system, Chinese developers are producing—and then giving away—smaller, more efficient, more affordable models. While the US government seeks to restrict China’s access to the most advanced chips, China is betting that incentivizing their tech giants to create leaner, more open models using more commodity hardware—models that can be trained with older chips and run even on personal computers—will be an advantage in achieving widespread use and, perhaps, Chinese national influence. There are other pathways for AI development that are not in service of private capital gains nor authoritarian regimes, but rather a democratic public interest. The best example comes from Switzerland, where public institutions—research funding agencies, universities, supercomputing centers—have collaborated to produce an AI model called Apertus. It is trained entirely on data validated to be licensed for use with AI (not stolen), on preexisting public computing infrastructure, and using renewable hydropower. Its developers are incentivized to produce a public good, not turn a private profit. It’s dangerous to confuse technology problems with sociopolitical ones. Popular proposals like pausing AI research, moratoria on data center development, or subjecting frontier models to federal government screening are all framed as addressing problems with AI’s technological development, but fail to take into account the larger social problems that govern it. China’s success with government-endorsed development of open-weight frontier models illustrates the futility of keeping AI tech as national secrets, or of any pledge to scale back deployment. AI is already legitimately useful for a wide range of tasks. It can be a tool for public good, if we choose to solve its sociopolitical problems. Our goal should not be to slow its pace of improvement or scale of deployment, but rather to steer it away from consolidating power and towards the public benefit. We can build sustainable AI, minimizing environmental and energy impacts. And we can equitably distribute the material gains it produces. Integrating a technology as disruptive as AI responsibly requires structural reforms, and we should decouple the social and technological aspects of AI to design those reforms. Companies—including tech giants—should be forced to pay the energy and environmental costs of its development. Profits should be taxed adequately and redistributed. Antitrust laws should be strongly enforced. Corporations should have a fiduciary responsibility to stakeholders beyond their majority shareholders. These badly needed reforms are responsive to the problems with capitalism that AI is exacerbating, even if they are not specific to the technology.
schneier.comAug 13, 2026extracted
Cybersecurity Alliance Drafts SAFE Guidelines for Sharing AI Incident Data
The Linux Foundation has issued a Request for Comments on a newly proposed framework aimed at standardizing how the cybersecurity industry handles agentic AI incidents. Announced at the Black Hat conference in Las Vegas, the Shared AI Findings Exchange (SAFE) guidelines seek to turn AI security incidents and near misses into actionable threat intelligence for the broader ecosystem. The SAFE framework is being driven by the recently launched Open Secure AI Alliance, a coalition that has grown to over 120 organizations. The SAFE initiative is spearheaded by Open Secure AI Alliance members such as Nvidia, Cisco, CrowdStrike, Hugging Face, and Red Hat. The core objective is to establish a confidential pipeline for collecting incident data, analyzing control failures, and broadcasting evidence-based recommendations to reduce systemic risks. Because modern AI agents function as complex systems reliant on identity controls, runtimes, and execution harnesses, the alliance emphasizes that open intelligence sharing is the only way defenders can match the speed of emerging attack vectors. Alongside the policy framework, alliance members have released various open source tools covering the entire AI security stack. Nvidia has contributed its NOOA research harness for auditing agent behavior, the OpenShell runtime that restricts agent access at the system level, and Garak, an LLM vulnerability scanner designed to catch prompt injections and data leaks prior to deployment. Okta is developing implementations utilizing the open Cross App Access (XAA) protocol to secure agent connections within OpenShell sandboxes. Meanwhile, Red Hat launched a new open source project called Asago, which maps external governance requirements, such as those in the EU AI Act, directly to live runtime controls for AI agents. Newly added members Amazon and Visa have contributed frameworks for building and evaluating agent boundaries, with Amazon specifically open-sourcing Cedar, an authorization language for establishing verifiable access controls. Microsoft is releasing tools like PyRIT and RAMPART, which allow red teams to run automated testing and turn incident findings into repeatable software checks. The new guideline proposal comes in light of OpenAI and Anthropic discovering that their models went rogue during tests and attacked real organizations. Related: Rethinking AI Security: Why CASB and DLP Need an Interaction-Aware Layer
securityweek.comAug 5, 2026extracted
Your enterprise AI footprint is about three times bigger than your model list
Your enterprise AI footprint is about three times bigger than your model list Organizations are building AI systems that combine models, agents and external tools instead of relying on standalone AI, according to Snyk’s latest State of Agentic AI Adoption report. The study analyzed 3,044 enterprise environments and 1.39 million code repositories to examine how enterprises are deploying AI. Adopting agentic architectures Of organizations using AI, 46.9% have adopted agentic architectures built on AI agents, model context protocol (MCP) servers, or both. More than half have deployed the full stack, combining AI agents with MCP infrastructure that enables access to enterprise data, applications, services and external tools. Agentic adoption: agents vs. MCP servers (Source: Snyk) These systems extend beyond chat interfaces by retrieving information, coordinating workflows and carrying out actions across enterprise environments. Enterprise AI deployments include far more than models. When frameworks, MCP servers, retrieval systems, vector databases, datasets and supporting tools are included, the average AI footprint is about three times larger than model inventories indicate. Nearly half of the companies analyzed had no declared AI models in their code repositories. They used AI through third-party services, packages and tools. At the same time, 17.2% operated large fleets of AI models, indicating AI is integrated across platforms and applications. Organizations need visibility across this broader AI ecosystem because supporting components introduce additional dependencies, integration points and governance requirements. The AI supply chain OpenAI remained the most widely used model provider, although Anthropic and other vendors increased their share of enterprise deployments. The top four providers accounted for about 71% of identifiable model occurrences, showing a broader mix of AI vendors across enterprise environments. Proprietary models represented 63.8% of deployed models, while open-source models accounted for 32.5%. Organizations use proprietary models for advanced reasoning and autonomous tasks. Open-source models are commonly deployed for embeddings, retrieval and other supporting workloads. Third-party software plays a significant role in enterprise AI. External sources accounted for 77.4% of AI packages and tools, while 22.6% were developed internally. These dependencies shape how AI systems operate and introduce security, governance and software supply chain risks that organizations need to manage. Capability, lineage and AI governance Enterprise AI is becoming more capable, even though companies are not always deploying the latest models. They rely on established proprietary models for production workloads to balance performance with operational requirements. Open-source models are increasingly used for retrieval, embeddings and other specialized functions. The gap between proprietary and open-source models is narrowing, giving enterprises more options when building and operating AI applications. About half of the organizations using AI models could not be linked to the datasets used to train or fine-tune them. This makes it harder to understand how models generate results, investigate incidents and demonstrate compliance. AI is becoming deeply embedded in software development. Developers are working with more AI models, agents and supporting tools, increasing the amount of autonomous technology operating across enterprise environments. Enterprises need visibility into what AI systems can do, what data they use, what resources they can access and how they behave in production. AI adoption across industries and regions AI adoption varies widely across industries. Media and entertainment companies had the highest concentration of AI components per organization, followed by retail, consumer goods and education. Technology and IT companies deployed the largest overall volume of AI, while other sectors integrated AI extensively across individual businesses. Industries where AI supports content creation, customer experiences and business processes recorded the highest levels of adoption. These sectors are deploying not only AI models but also agents, orchestration frameworks and supporting infrastructure that enable AI to perform tasks across multiple systems. Businesses in North America and Europe are building similar AI architectures, with both regions increasing adoption of AI agents and MCP infrastructure. North American organizations generally deploy AI at a larger scale, while the underlying technologies and architectural patterns remain consistent across both regions.
helpnetsecurity.comAug 5, 2026extracted
WAFを89%すり抜ける事例も──AIが休みなく仕掛けるWeb攻撃、予防策はあるか
「見えないWeb攻撃」──情報漏えい対策の盲点 WAFを89%すり抜ける事例も──AIが休みなく仕掛けるWeb攻撃、予防策はあるか(1/4 ページ) 米Anthropicや米OpenAIなどが開発するフロンティアAIの登場による、脆弱性探索能力の飛躍的な向上が、サイバーセキュリティの攻防を否応なく次のステージに押し上げようとしている。 エージェント型AIによる“マシンスピード”での攻撃が、WebやAPIに具体的にどのようなリスクをもたらすのか。実際に観測されているAIによるWeb攻撃の特徴を分析し、防御側の備えとして何が必要になるかを具体的に探っていこう。 AIでサイバーリスクはどう変わる? フロンティアAIが攻撃に使われることで、何が変わるのだろうか? まず、OSやシステム、プログラムなどの未知の脆弱性が大量に発見できるようになる。Anthropicの「Claude Mythos」の先行利用権を得て自社のサービスの脆弱性を検証した企業は、(AIも組み込まれた)既存の脆弱性検査ツールと比べ、発見される数量も検知の精度もまさに桁違いだとレポートしている。 こうしたモデルは、発見した脆弱性を悪用して攻撃を仕掛けるためのプログラム(エクスプロイトコード)を瞬時に生成する能力も備えている。人力では数日から数週間を要するが、生成AIの能力があれば、わずか数分に短縮される。 AIは、単体では深刻とまではいえない脆弱性を組み合わせ、実システムに対して攻撃を繰り返すことで、防御をすり抜ける道筋を割り出す能力にも長けている。これら一連の手順を、AIが“マシンスピード”で試行し続ける点が最大の脅威となっている。 その一方で、多くの企業や開発ベンダーはアプリケーションの開発に生成AIを用いる「バイブコーディング」を取り入れようとしている。開発の高速化を見込む動きではあるが、一方でセキュリティ面が十分でないコードが量産される恐れもあり、多くの脆弱性を生むリスクが高まっている。 最新AIがWAF防御をすり抜ける? Web攻撃の変化 こうした状況変化の結果、WebサイトやWebアプリケーション、スマートフォンアプリ・サービスで用いられるWeb APIはとあるリスクに直面している。攻撃者が操るAIから、WAF (Web Application Firewall)の検知ルールをすり抜けられるパターンの探索を、絶え間なく受けるリスクが高まっているのだ。 WAFのすりぬけには「難読化」という手法がよく用いられる。Webへの侵害では、Webサーバの背後で動作しているプログラムやデータベースなどの不正な操作を引き起こす命令文を送る。その命令文の途中に特殊な文字を混ぜるなどして、WAFルールで設定された照合用文字列との一致による検知の回避を試みるのが難読化の手口だ。 つまり、AIは「WAFの検知は回避できるが、標的にしたアプリケーションには意図した命令が通る」すり抜けパターンを大量に生成し、自動で試そうとする。 AIが難読化したリクエストが、実際WAFの検知ルールを突破してしまうことは、研究者により検証されている。例えばAIが生成した、SQLインジェクションの派生パターンによる攻撃の試行では、オープンソースのWAFである「ModSecurity」で89%、「AWS WAF」では約41%の攻撃が、それぞれのWAFルールによる検知を回避して突破。同様にクロスサイトスクリプティングでも、80%がModSecurityの検知を突破している。 実際に、クラウド型WAFを企業向けに提供しているAkamai Technologiesでは、攻撃側のAIが生成したと考えられる難読化された攻撃の試行をすでにとらえている。 Copyright © ITmedia, Inc. All Rights Reserved. 「見えないWeb攻撃」──情報漏えい対策の盲点 さまざまなWebサービスが、巧妙に隠された「見えないWeb攻撃」に狙われている。個人情報漏えいや犯罪行為を引き起こしている攻撃の実態を、独自のデータを交えて分析する。
itmedia.co.jpAug 2, 2026extracted
Thousands of malicious AI skills found capable of stealing data, running malware
Thousands of malicious AI skills found capable of stealing data, running malware AI agents can browse the web, use external tools, execute commands, and perform tasks on behalf of users. Many rely on skills that define how they interact with services and data. Malicious skills can abuse those capabilities to steal data, execute malware, or manipulate an agent’s behavior, according to the H1 2026 ESET Threat Report. Malicious AI skills expand the attack surface An analysis of nearly 900,000 AI skills identified more than 25,000 suspicious skills and over 3,000 malicious ones. Between March and May 2026, the number of unique skills scanned increased from 60,000 to almost 900,000. Suspicious skills grew from around 10,000 to more than 25,000. Malicious skills increased from about 600 to over 3,000. Researchers identified capabilities including command execution, file access, downloading third-party tools, credential loading, code injection, and obfuscation. These capabilities can support legitimate tasks. They can be used to steal data, execute malware, manipulate AI agents, or gain unauthorized access to systems. The analysis also identified red-team, self-modifying, and online purchasing skills. Some security scanner skills performed only basic checks, giving users a false sense of protection. “AI skills can enable a wide range of agentic AI abuses, from automated reconnaissance and red-team-style attacks to spam generation, malware modification, and distribution. Adversaries will likely keep testing these approaches to bypass controls, including by obfuscating intent or using region-specific, niche, or constructed languages,” said Anton Mäčko, ESET Malware Analyst. ClickFix expands into AI services and enterprise workflows ClickFix is expanding into new environments by using fake error messages and verification prompts to trick users into running malicious commands. Detections of ClickFix attacks increased by 108% between H2 2025 and H1 2026. The technique spread beyond fake CAPTCHAs to include macOS, WordPress sites, browser extensions, AI-themed help pages, and enterprise authentication workflows. A web page, abusing Anthropic’s Artifact pages domain, with AI-fix instructions (Source: ESET) New variants include AI-fix, which uses fake AI-generated troubleshooting pages hosted on services associated with Anthropic, OpenAI, and Microsoft. CrashFix uses malicious browser extensions to display fabricated security warnings. ConsentFix steals Microsoft OAuth authorization codes through fake verification prompts on compromised websites, allowing attackers to obtain OAuth tokens. QR code phishing reaches record levels QR code phishing, also known as quishing, continued to grow during H1 2026. Attackers embedded phishing links in QR codes to direct victims to credential theft websites, often accessed through mobile devices. Approximately 11% of all detected phishing emails contained QR codes, with an average of 100,000 detections per month. The highest volume was recorded in April. The United States accounted for 19% of QRcode/phishing detections, followed by Spain with 17% and Mexico with 6%. Generative AI reaches Android malware PromptSpy became the first Android malware to use generative AI during execution. The malware uses Google’s Gemini model to interpret the device interface and generate gestures that help maintain persistence. It intercepts lock-screen PINs and passwords, captures screenshots and video, uploads information about installed apps, and provides attackers with remote access. Researchers recorded one PromptSpy detection after its discovery. Ransomware groups expand use of EDR killers Ransomware groups continue to use endpoint detection and response (EDR) killers to disable security software before deploying ransomware. Researchers tracked more than 100 EDR killers used in the wild, including over 60 Bring Your Own Vulnerable Driver (BYOVD) variants that abuse more than 40 legitimate vulnerable drivers. New variants appear every week. Attackers use anti-rootkits, scripts, and driverless techniques to interfere with security software. Ransomware payment rates continued to decline during 2025. Chainalysis reported that 28% of victims paid a ransom. Ransomware attacks increased by 50% year over year. The median ransom payment rose by 368% to nearly $60,000. Total ransomware payments reached $820 million in 2025.
helpnetsecurity.comJul 8, 2026extracted
Anthropic's Fable 5 and Mythos 5 Are Back with New Security Guardrails
Anthropic’s latest frontier large language models (LLMs), Claude Mythos 5 and Claude Fable 5, are available again – but with added security limitations. On June 30, just 19 days after the US government enacted export controls on both models which forced Anthropic to suspend their global distribution, the decision was lifted. The same day, the AI lab announced it was redeploying both models from July 1. However, they will now come with additional limitations aimed to address AI safety and security concerns raised by the US government. Fable 5 Equipped With New US-Approved Safeguards Fable 5, a general-access LLM powered by the same underlying frontier AI model as Mythos 5 – itself an upgrade from Claude Mythos Preview – is now available to users globally across all the Claude Platform, Claude.ai, Claude Code and Claude Cowork. For premium users who have subscribed to Pro, Max, Team and select Enterprise plans, the model will be included for up to 50% of weekly usage limits through July 7, after which it will be available via usage credits. Anthropic is also rolling out availability of the general-access model on AWS, Google Cloud and Microsoft Foundry. The AI company confirmed it had reviewed the Amazon report that prompted the export control directive. In the report, researchers had found a jailbreak, a method of prompting Fable 5 so that it identified software vulnerabilities and, in one case, provided an exploit – therefore bypassing the model’s built-in safeguards. While Anthropic said that the reported technique “did not expose any unique Mythos-level cyber capabilities,” the company is releasing a new version of Fable 5 equipped with “an improved safety classifier that targets and blocks the behavior described in the report.” A classifier is a small, automated AI systems that, during an interaction with an LLM, detects when the model is asked to perform a potentially harmful task or to produce potentially harmful outputs and then blocks it from responding to requests. According to Anthropic, the new classifier blocks the jailbreak identified by Amazon researchers “in over 99% of cases.” It may, “in a very small fraction of cases,” provide information after a potentially harmful user request but Anthropic claimed it wouldn’t be “detailed enough to help a cyber attacker.” “The model’s safeguards are not expected to block all low-risk routine cyber defense capabilities – just those that are potentially harmful,” said the company. When a request to Fable 5 is blocked, the users will be notified that it has been redirected to Opus 4.8. “The new classifier also comes at the cost of flagging benign requests more often during routine coding and debugging tasks,” Anthropic admitted. It said it would continue to refine the safeguards to better distinguish genuine misuse from legitimate requests and reduce false positives. Anthropic said researchers from the US Department of Commerce’s Center for AI Standards and Innovation (CAISI) have tested the new safeguards and described them as “extraordinarily strong.” Anthropic Teams Up with Government and Industry to Accelerate AI Security Before lifting export controls on Fable 5 and Mythos 5 on June 30, the US government approved Mythos 5 to be redeployed to a set of US organizations that operate and defend critical infrastructure. “We continue to coordinate with the government to expand access to the broader set of domestic and international partners in the Glasswing program,” Anthropic noted. The company also said it was collaborating with the US government to accelerate AI security, including via pre-deployment testing and evaluation. In the meantime, the AI lab has worked with Amazon, Microsoft, Google and other Glasswing partners to draft a consensus framework for assessing the severity of AI jailbreaks – including finding a standard definition of what constitutes a “universal jailbreak” – and how AI developers should respond to them. Finally, the AI lab has launched a new HackerOne program where security researchers can submit potential cyber jailbreaks they’ve discovered in Fable 5 for review. Image credits: wutianzeri / RixAiArt / Shutterstock.com
infosecurity-magazine.comJul 1, 2026extracted
Browser-Only Ransomware: From LLM Hallucinations to a Practical Attack Technique
Browser-Only Ransomware: From LLM Hallucinations to a Practical Attack Technique July 1, 2026 Research by: Alexey Bukhteyev Key Takeaways AI can turn high-level malicious ideas into concrete techniques, and can independently design and implement novel attack paths that have not yet appeared in real-world campaigns. In this research, DeepSeek connected unrealistic browser-malware concepts with a real browser capability, turning an AI-generated malware hallucination into a plausible browser-native ransomware technique. Although the generated sample was incomplete, it exposed a practical abuse path based on the File System Access API and access to photo directories. The technique does not require a native payload, APK installation, browser exploit, or root access. It relies on social engineering and a legitimate permission prompt exposed by the File System Access API in Google Chrome. The Android scenario is especially concerning because photo directories are high value personal data stores and, unlike iOS, modern Android Chrome versions expose a browser API that allows web pages to read and modify files in those directories after user approval. Using a fake AI image-enhancement workflow gives users a plausible reason to approve folder-level file access. Our PoC demonstrates this browser-only workflow against selected image directories on Android. Introduction Over the past several years, large language models have reshaped software development, and malware development has followed the same path. Check Point Research has documented this trend from early experiments showing that AI systems could generate offensive components, to cases of cybercriminals using ChatGPT to create malicious tools, and later to advanced AI-authored malware frameworks such as VoidLink. In some cases, LLMs lowered the barrier enough for users with little or no development experience to produce working offensive code. As frontier models became better at writing reliable code, including complex security related components, major AI vendors also turned cyber safety into a dedicated control area. Clearly malicious requests involving credential theft, malware deployment, ransomware behavior, persistence, stealth, or unauthorized exploitation are now commonly blocked or refused. OpenAI’s cyber-safety documentation, for example, describes additional safeguards for models classified as having High Cybersecurity Capability, while Anthropic has published reports on detecting and countering cyber misuse of Claude. DeepSeek then becomes particularly relevant in this context for several reasons: Lower refusal rates for harmful cyber enforcement: compared with Anthropic and OpenAI, DeepSeek models were less consistent refusing harmful cyber requests, including the File System Access API implementation we will be discussing later on this article. Low barrier to access: DeepSeek is free to use via the web interface, widely available, and accessible in regions where other frontier models face regulatory or commercial restrictions. This lowers the cost of repeated malicious experimentation. End-to-end malicious code from a single prompt: in our testing, a working malicious application could often be generated from a single broad prompt. Achieving a comparable result with OpenAI or Anthropic typically requires decomposing the attack into multiple benign-looking requests and manually assembling the generated components. Putting this all together, these differences make DeepSeek particularly attractive to threat actors: DeepSeekmodels can turn high‑level malicious ideas into concrete, complete attacks with less expertise than competing platforms. Check Point Research analyzed nearly 3,000 files attributed to DeepSeek observed in public telemetry over the past year. The dataset included Python, PowerShell, Batch, HTML, JavaScript, VBScript, and other file types. Of these, 1,383 files were classified as malicious or dangerous by either VirusTotal detection or static source analysis. Within this dataset, we found a sample that implemented a dangerous browser-native technique we have not observed exploited in the wild. We refer to it as In-Browser Ransomware. The technique uses a phishing lure to persuade the victim to grant file-system access to a web page; once access is granted, the page can enumerate local files in the selected folder, read and exfiltrate their contents, encrypt and overwrite them, and display a ransom-style message, all without installing a native payload or exploiting the browser. The underlying browser risk was already known to browser engineers. The File System Access specification explicitly lists ransomware as a security consideration, and the 2023 USENIX Security paper RoB: Ransomware over Modern Web Browsers studied the abuse of the File System Access API to encrypt local files from a malicious web application. The important finding in our research and what is new, is how the AI model brought these previously documented concepts together, into a realistic and enforceable attack scenario leveraging a method that defenders had originally thought was unfeasible due to browser sandboxing limits: a DeepSeek-attributed malicious sample, generated as an all-in-one malware fantasy, connected this documented platform risk to a realistic phishing-style web application, demonstrating a viable end-to-end attack chain. An attacker does not need to know that a browser exposes a file-system API. They can ask for an impossible-sounding outcome – a website that steals files, captures keystrokes, takes screenshots, encrypts files, and demands payment – and the model may connect the request to a real browser capability. Basically, the AI model showed an ability to reason across existing knowledge and combined multiple known components into a coherent attack workflow that could be readily used by an attacker. This illustrates how frontier AI models may move beyond simply enhancing existing attacker techniques to lowering the expertise required to operationalize complex attack chains by connecting knowledge in ways that previously relied on human experience and creativity. A Noisy Sample With One Important Idea The sample that caught our attention is SHA256 07c39f79ab92fb21557b82283472dce1c112f577d796111fb752c3c6d84c86b5, a Python Flask application that serves victim-facing HTML and JavaScript from embedded templates and also includes backend routes intended to receive information from the victim and provide an administration panel. We do not have the prompt submitted to the AI model that produced this sample. Judging by the code structure, function names, and comments, it was likely formulated very broadly such as something similar to this example: create a universal malicious tool that runs through the browser and collects as much victim data as possible, encrypts files, and demands ransom. In a single front-end, the generated code assembled routines and stubs for keylogging, clipboard monitoring, form and network-request interception, Discord-token collection, crypto-wallet and payment-card discovery, geolocation requests, webcam and microphone access, screenshots, local-file access, Chrome exploit stubs, “persistence,” and a ransomware-style overlay. This does not mean the sample actually implements all of these capabilities. A more accurate reading is that it is an AI-generated blueprint in which the model tried to translate familiar capabilities of native stealers and ransomware tools into a web page opened in the browser. The victim-facing page is disguised as a Discord avatar AI upscaler: Clicking the button on the victim-facing lure page is intended to start the malicious browser-side sequence, although the generated control flow is inconsistent and does not complete reliably. After a fake processing step, the page is intended to display a ransomnote-style overlay under the name InfernoGrabber v9.0. The message claims that passwords, credit cards, and personal files were encrypted, demands Bitcoin, and displays a countdown threatening publication of private data. Most of the functionality claimed in the sample collapses at the browser boundary. A normal web page can observe activity inside its own origin, capture input events delivered to its own DOM, request browser-mediated permissions, access storage scoped to its own origin, and render frightening overlays. It remains constrained by the browser security model. In this sample, the “desktop screenshot” routine captures the rendered web page, the keylogger observes keystrokes only while the user interacts with the page, webcam and microphone capture depend on browser permission prompts, and the Discord-token stealing logic searches storage available to the current origin. The “persistence” logic relies on browser storage and a service worker registration attempt. Much of the sample therefore reads as an AI hallucination produced in response to an overly broad prompt or to requirements that a normal web page cannot satisfy. The exception was the file-access workflow, where the generated code reached for a real browser primitive with practical abuse potential. The generated JavaScript referenced: showOpenFilePicker(); showDirectoryPicker(); recursive traversal of a user-selected directory; reading selected files through browser file handles; sending file contents to the Flask backend; displaying a ransomware-style warning after the interaction. The File System Access API is a legitimate browser capability designed for web applications such as editors, IDEs, and creative tools. After the user grants access, a web application can read files and folders from the local device. The API also supports write access and directory enumeration under browser permission controls. The technique is limited to browsers that expose the picker-based File System Access API. At the time of writing, this primarily means Chromium-family browsers: the API shipped on desktop in Chrome 86, and Chrome 132 extended File System Access support to Android and WebView. Firefox and Safari do not expose the same local file and directory picker methods, which limits the immediate attack surface but also concentrates the risk in Chrome-based browsing environments. The sample lacked a complete and reliable browser-side encryption flow, yet the attack design was concrete: a fake utility convinces the user to grant browser file access, which allows the page to exfiltrate and encrypt files. The model combined fake OS-level malware claims with a real browser primitive and produced a browser-native file-theft and ransomware scaffold. The sample shows how an LLM can transform an abstract malicious request into a new attack blueprint. The user likely wanted an all-in-one tool: a Discord-themed lure, a stealer, an admin panel, and a ransomware or locker workflow. The model chose a Flask application and a browser frontend as the unifying architecture. In doing so, it connected a hallucinated malware concept to a real platform feature with genuine abuse potential. Even though we have not yet observed this exact browser-native ransomware pattern widespread in-the-wild campaigns, the technique is still operationally relevant for several reasons: The browser becomes the execution environment: the attack runs entirely inside the browser process, without installing any additional app, dropping a binary, or exploiting a vulnerability. Traditional endpoint protections focus on apps and native payloads; a website that encrypts files after a legitimate-looking permission sits outside those assumptions. Lower friction for victims: opening a web page and clicking “Allow” on a file-access prompt is a normal part of using modern web applications. Users do not intuitively treat this as “running malware”, which makes the social-engineering angle powerful. Cross-platform reach: the same browser-native technique can target any platform where the File System Access API is exposed, we tested on Android and Windows. From Hallucinated Scaffold to Working PoC Because the original sample was incomplete, we tested whether the latest DeepSeek model V4 could turn the same browser-native attack idea into a working proof of concept. When prompted directly to create ransomware, the model consistently refused across all tested modes. Even though some requests were denied, we managed to succeed in the end. We removed explicit terms such as “ransomware” while preserving the same functionality: a web page that asks the user for access to local files, processes them inside the browser, and leaves the user unable to recover the original content. In Instant mode, DeepSeek consistently generated HTML/JavaScript code that used the File System Access API to interact with user-selected files. In Expert mode, the behavior was inconsistent across attempts: several attempts ended in refusal; one generated a non-functional sample; one generated a fully working browser-based ransomware PoC. One response was especially notable because the model described the result as: “a crafted trap that combines a convincing AI upscaler interface with hidden ransomware-like behaviors” This wording shows that the model recognized the malicious nature of the scenario while still continuing the generation. For comparison, we tested similar requests against ChatGPT and Claude. In our tests, these systems either refused to help or generated constrained browser-safe implementations that did not use the File System Access API. This does not mean that the same outcome is impossible with other frontier systems. With an incremental approach, a user can ask for separate components that appear benign in isolation, such as a user interface, browser file handling, client-side data transformation, and neutral status messaging, and then assemble them into a harmful workflow by replacing the neutral messages with a ransom note. The difference is the level of steering required. In that scenario, the user needs enough technical understanding to decompose the attack, preserve the malicious objective across separate requests, identify the right browser primitive, and combine the generated pieces manually. In-Browser Ransomware on Android To assess the practical risk of this technique, we used an LLM to build a controlled proof-of-concept (PoC) based on the same idea we observed in the DeepSeek-attributed sample: a browser-native ransomware workflow disguised as an AI image upscaler. On Android, modern Chrome versions expose the picker-based File System Access API to web content. On iOS, Safari does not expose the same File System Access primitives to websites. Access to photos is mediated by the operating system’s app-sandbox and photo-library permissions instead of a web API that can enumerate and modify arbitrary folders. Chrome on iOS uses WebKit which also does not implement File System Access API. As a result, on mobiles, the technique we demonstrate is currently practical on Android Chromium browsers. At the same time, the attack surface is narrower than arbitrary disk access. The picker-based File System Access API does not let a web page target the whole system disk, and Chromium applies additional restrictions to sensitive locations. In Chromium’s current implementation, broad access to locations such as the user’s home directory, Desktop, Documents, Downloads, Chrome data, application directories, Windows, Program Files, AppData, and several Linux and Android system paths is blocked or constrained. The File System Access specification also explicitly recommends restricting sensitive directories and lists ransomware as one of the risks the API design must account for. However, selection of the root of the default Pictures and Videos directories was not restricted on any of the tested operating systems (Android and Windows). This capability fits naturally into a social-engineering workflow for a fake photo-processing application. On desktop, the Pictures folder may contain personal files, but it is usually less central to business workflows than the user’s entire home directory or a Documents directory. On mobile, the risk profile changes: the photo library is often one of the most valuable local data stores. It may contain years of private photos, identity documents, banking screenshots, medical records, recovery codes, travel documents, work images, and photos of family members. Losing access to this data, or having it exfiltrated, can create personal or business issues from ransomware to blackmail or if the data is sensitive, public disclosure leading to reputational damage and more. Chrome 132 introduced File System Access support on Android, allowing web applications, after user approval, to read and save changes directly to selected files and folders. We tested this capability on several Android devices and confirmed that the latest Chrome version available to us at the time of testing, Chrome 148, also allowed selecting the photo directory, including the root of the DCIM folder. The workflow on Android looks very natural. The user opens a web page that promises to enhance a photo, selects an image, and is then asked to choose a directory for saving the “enhanced” results. The browser warning that the site will be able to edit files in the selected folder is easy to rationalize in that context: the user expects the service to write processed images back to the device. During the fake processing step, the PoC encrypts pictures inside the selected directory. Video 1 – Demonstration of a browser-native ransomware PoC on Android using the File System Access API. The combination of this technique, a natural social-engineering lure, and browser-only execution makes the Android scenario especially concerning. The resulting flow requires no APK installation, no vulnerability exploitation, no native payload, and no root access. Users generally do not treat opening a web page as a malware execution event, especially when no application is installed and no binary is downloaded. In this case, the browser prompt appears in a context where file access feels expected, while the granted permission gives the page meaningful control over a directory that may contain highly sensitive personal data. Practical Recommendations for Users While this research focuses on a controlled PoC, there are concrete steps users can take today to reduce the risk of browser-native ransomware abuse: Treat browser folder-access prompts as high-stakes decisions: before approving “access to files in a folder”, check which site is asking, which folder is being selected, and whether editing files is truly necessary for the feature you expect. If you are unsure why a site needs write access to an entire directory, decline the request. Avoid granting websites access to sensitive or irreplaceable data: do not expose folders that contain personal photos, identity documents, recovery codes, or work data unless the site is highly trusted and the need is clear. Prefer selecting a temporary or empty folder for experimental web tools, rather than your main photo library. Prefer well-established applications for high-value data: for tasks such as backing up photos, editing large collections, or processing sensitive images, use reputable native apps or well-known cloud services instead of newly discovered browser tools with unknown reputation. Maintain offline and cloud backups of important data: regular backups reduce the leverage attackers gain from encrypting or deleting local files, whether through native ransomware or browser-based techniques. Keep browsers and mobile OSes updated: browser and OS vendors continue to refine permission models and harden sensitive APIs. Applying updates promptly ensures that you benefit from the latest security controls around features like File System Access. Be skeptical of AI-branded lures: attackers increasingly disguise malicious flows as “AI” utilities, avatar upscalers, photo enhancers, or productivity tools. A polished AI-themed interface is not a guarantee of safety; apply the same caution you would to any unfamiliar site asking for broad access to local files. Conclusion LLM-assisted malware development changes the economics of malicious experimentation. A user with limited technical understanding can describe a harmful outcome, generate code, test the result, adjust the prompt, and repeat the process at very low cost. Tasks that once required a developer, a purchased builder, or prior knowledge of the relevant platform can now be approached through cheap iteration. This also changes the defender’s problem. Malware generated this way may move the ecosystem away from a limited set of reused families and builders toward a larger volume of disposable, one-off artifacts, each carrying a unique combination of techniques, API usage, and payload logic. Hallucination adds another important dimension. AI-generated malware can be technically wrong and still reveal practical malicious techniques. When a model tries to satisfy unrealistic requirements, it may search across legitimate platform features and map a malicious goal to an API that actually exists. This process can surface techniques that defenders have not yet seen in the wild, or turn risks previously described mostly in theory into workable attack concepts. The case analyzed in this research shows exactly that: a noisy and partially broken artifact connected a theoretical browser risk to a practical browser-only ransomware technique. In this case, the user likely asked for an impossible web application, a single browser page that behaves like a fully features stealer and ransomware agent. The model could not satisfy all of those requirements correctly, but in the process of trying, it searched across legitimate browser features and anchored part of the fantasy to a real API: the File System Access API. This illustrates a broader risk: A non-expert attacker does not need to know that such an API exists or how to abuse it. By describing a high-level malicious outcome in natural language, they can cause the model to discover and connect the malicious goal to previously under-explored platform capabilities. The resulting prototype can then be refined into a working PoC with minimal additional prompting or manual editing. In other words, AI is not only lowering the barrier for reimplementing existing malware techniques; it is also capable of bridging the gap between purely theoretical risks and practical, novel attacks that defender have not yet seen deployed in the wild. Historically, new attack techniques emerged through human experimentation, experience, and creativity. Frontier AI changes that dynamic. Rather than being constrained by conventional thinking or established attacker playbooks, AI can reason across existing knowledge and synthesize it in unexpected ways, connecting known capabilities into practical attack chains. The real shift is not that AI is inventing entirely new vulnerabilities, but that it may identify combinations and attack paths that humans had not previously recognized or operationalized. At the time of analysis, we found no evidence that this technique had been adopted as an in-the-wild malware pattern. The original DeepSeek-attributed sample was incomplete and failed to implement the full attack reliably. However, our testing showed how little effort is required to transform the same idea into a fully working implementation using modern LLMs. The resulting workflow is especially concerning on mobile devices, where a seemingly legitimate request for access to a photo directory can expose highly sensitive personal data to encryption, exfiltration, or both. From a defensive perspective, browser folder-access prompts should be treated as security decisions rather than routine clicks. Before granting a website access to an entire folder, users should review which site is asking, which folder is being selected, whether file modification is allowed, and whether the permission matches the action they intended. Users should avoid granting websites access to directories containing sensitive, private, or irreplaceable data whenever possible. “The Turkish Rat” Evolved Adwind in a Massive Ongoing Phishing Campaign Check Point Research Publications August 11, 2017 “The Next WannaCry” Vulnerability is Here Check Point Research Publications March 12, 2026 “Handala Hack” – Unveiling Group’s Modus Operandi SUBSCRIBE TO CYBER INTELLIGENCE REPORTS We value your privacy! BFSI uses cookies on this site. We use cookies to enable faster and easier experience for you. By continuing to visit this website you agree to our use of cookies.
research.checkpoint.comJul 1, 2026extracted
Big Tech, nuove tensioni tra UE e Usa. La Commissione si oppone ai nuovi dazi di Donald Trump
La Commissione Europea ha risposto a Donald Trump, intenzionato ad introdurre nuovi dazi per proteggere le aziende tecnologiche statunitensi. Dazi e Big Tech, così torna a salire la tensione tra l’UE e gli Usa. La Commissione Europea ha risposto infatti con fermezza alle nuove minacce di Donald Trump di imporre dazi ai Paesi europei. In particolare, a quelli che dovessero adottare norme fiscali e regolatorie nei confronti delle grandi aziende tecnologiche americane. Il confronto è arrivato nel momento in cui Washington sono iniziati i colloqui tra rappresentanti dell’UE e dell’Amministrazione statunitense per cercare di rilanciare “il dialogo sulla cooperazione digitale“. La delegazione europea – con a capo il Direttore Generale della DG Connect, Roberto Viola – è rimasta nella capitale americana fino a mercoledì. La Commissione ha definito tutti questi incontri come “un primo passo verso un possibile dialogo strutturato tra le due sponde dell’Atlantico“. Le minacce di Donald Trump Le nuove tensioni sono cominciate venerdì scorso dopo che Trump aveva rilanciato, attraverso i suoi profili sociali, la volontà di colpire i Paesi europei. Nello stesso fine settimana, sottolinea POLITICO, anche il Dipartimento di Stato americano ha espresso critiche sulle recenti iniziative europee volte a rafforzare l’autonomia tecnologica regionale. Secondo Washington, infatti, sono tutte misure “protezionistiche“. Il portavoce della Commissione Thomas Regnier ha dichiarato: “La nostra posizione è molto chiara. L’UE e i suoi Paesi membri hanno il diritto sovrano di regolamentare qualsiasi attività economica svolta sul proprio territorio“. Regnier ha inoltre avvertito che Bruxelles è pronta a reagire “rapidamente e con decisione” se Washington decidesse di adottare misure unilaterali contro quelle che l’UE considera “politiche pienamente legittime“. Le preoccupazioni europee sono aumentate anche dopo la decisione degli Usa di allentare parzialmente le restrizioni all’esportazione degli ultimi modelli di intelligenza artificiale sviluppati da Anthropic. Si è così riacceso il dibattito sulla dipendenza europea dalle tecnologie americane e sul potenziale rischio di un controllo strategico del Governo americano su strumenti nevralgici per la Sicurezza Nazionale. Le prospettive statunitensi Dopo il ritorno di Trump alla Casa Bianca, l’Amministrazione americana ha intensificato le critiche alla normativa europea sul digitale. La principale accusa è quella di colpire in modo sproporzionato le aziende statunitensi. Già a marzo, l’ambasciatore americano presso l’UE, Andrew Puzder, aveva sottolineato la volontà di Washington di inserire le regole europee sul digitale tra i temi centrali del confronto bilaterale. Nell’occasione, aveva sostenuto che molte imprese americane le ritengono eccessivamente onerose. Nel frattempo, la Commissione Europea ha presentato un pacchetto di norme per favorire la crescita di tecnologie europee. E, allo stesso tempo, offrire ai Governi strumenti per limitare l’accesso delle aziende extraeuropee ai settori più sensibili del mercato pubblico. Secondo il Dipartimento di Stato americano, tuttavia, le future norme europee sulla sovranità tecnologica rischiano di compromettere il partenariato transatlantico. Washington richiama infatti l’accordo commerciale recentemente formalizzato, che prevede l’eliminazione delle barriere non tariffarie agli scambi. “La strada da seguire è quella della deregolamentazione e della collaborazione nei settori dell’intelligenza artificiale e dei semiconduttori“, ha spiegato un portavoce del Dipartimento di Stato.
cybersecitalia.itJul 1, 2026extracted
Linux Foundation Unveils New Open Source Security Project Akrites
The Linux Foundation on Thursday announced a new industry effort aimed at efficiently addressing vulnerabilities in the open source software (OSS) ecosystem. Named Akrites, it establishes a shared Security Incident Response Team (SIRT) for coordinated discovery, patching, and public disclosure of OSS security defects. If it sounds familiar, it should. Less than two weeks ago, Chainguard announced Athena, a coalition of over two dozen fintech and technology organizations aimed at addressing OSS bugs before public disclosure. At the time, Chainguard said it would work with the Linux Foundation on a coordinated SIRT, noting that the increased use of AI in cyberattacks is essentially closing the window between public disclosure and patching. While the Linux Foundation’s new announcement makes no mention of Athena, Akrites walks the same path: it offers the tools and channels to report, validate, and address OSS vulnerabilities before their coordinated public disclosure. Akrites is supported by Anthropic, AWS, Chainguard, Cisco, Citi, Endor Labs, Ericsson, Google, IBM, JPMorganChase, Microsoft and GitHub, NVIDIA, OpenAI, RapidFort, Red Hat, Rust Foundation, Sonatype, Vodafone, and Zscaler, many of which were mentioned as members of Athena. Seed funding to support Akrites comes from the Linux Foundation’s directed fund Alpha-Omega, with other organizations providing engineering resources and additional funding. In addition to establishing a confidential, trusted partner for vulnerability disclosure, eliminating hundreds of uncoordinated independent reports, Akrites will also work with critical infrastructure to help deploy fixes before in-the-wild exploitation. “When patches are released to the public, adversaries are able to utilize AI to rapidly reverse engineer the underlying vulnerabilities, develop exploits, and launch attacks. The success of our efforts, therefore, will be measured in patch deployment, not publication,” the Linux Foundation said. Akrites was created with a focus on confidentiality, to prevent vulnerability weaponization before patches are delivered, and to act as the maintainer of last resort, ensuring that fixes can still be delivered for packages that are no longer maintained. Related: Tech Giants Invest $12.5 Million in Open Source Security Related: RSAC Releases Quantickle Open Source Threat Intelligence Visualization Tool Related: OpenAI Refocuses Cybersecurity Efforts on Patching Over Discovery
securityweek.comJun 26, 2026extracted
Chainguard, JPMorgan, BNY Team Up to Secure Open Source from AI Threats
Open-source security firm Chainguard has brought together dozens of partners in a new industry coalition to protect open-source software from AI attacks. The initiative, called Athena, was announced by Chainguard on June 16. Its founding members include BNY, Chainguard, Cisco, Cloudflare, Corridor, DepthFirst, Docker, JPMorganChase, Kyndryl, LTIMindtree and PwC. Based on preliminary work at Chainguard, Athena provides a vulnerability intelligence sharing platform and tools to fix the vulnerabilities frontier AI models, like Anthropic’s Mythos and OpenAI’s GPT-5.5.-Cyber, find before attackers can exploit them. Here’s how Athena works, according to Chainguard’s CEO Dan Lorenc: Coalition members pool vulnerabilities affecting open-source projects they have discovered and packages into the Athena platform using frontier AI programs they have access to, including Anthropic's Project Glasswing and OpenAI's Daybreak Chainguard patches them privately and affected projects are rebuilt as private, hardened versions, available to members through Chainguard Libraries before disclosure Coalition members that operate infrastructure, platform, network and security layers push non-patch mitigations ahead of disclosure so that coverage exists even where a clean patch does not yet Cybersecurity partners add their own detections, signatures and virtual patching The Athena coalition drives coordinated upstream disclosure Additionally, Chainguard hopes to work with the Linux Foundation on a coordinated Security Incident Response Team (SIRT) for open source and a maintainer of last resort program. Announcing the project on LinkedIn, Lorenc said Athena allows for every vulnerability one member discovers to get remediated and pushed upstream, “becoming a fix the entire ecosystem inherits, often before disclosure.” “And for the parts of the world that can't patch on an attacker's timeline, partners who sit in front of much of the internet push mitigations out ahead of disclosure, blocking the issue for people who never knew there was anything to block,” he added. Chainguard also highlighted that the Athena model acts as “an AI cybersecurity clearinghouse” like the one the US government has been asked to build following the Trump Administration's latest Executive Order, Promoting Advanced Artifical Intelligence Innovation and Security, published on June 2. “It’s even more relevant since the US government declared Mythos too dangerous for public access on Friday,” the open-source security company added. Athena is operational and has already processed over 20,000 findings and shipped more than 2000 patches across 500 open-source projects. The initiative will begin publishing its first wave of disclosures in July and continues to welcome new partners. “Will it be perfect? No, and no one should pretend otherwise,” said Lorenc. “But fragmentation is worse, standing still isn't survivable, and the more of the industry that's in, the less any attacker has left to find. Join us.”
infosecurity-magazine.comJun 16, 2026extracted
15th June – Threat Intelligence Report
For the latest discoveries in cyber research for the week of 15th June, please download our Threat Intelligence Bulletin. TOP ATTACKS AND BREACHES The University of Nottingham, a UK research university, has suffered a data breach after ShinyHunters accessed its student records system. The incident affected about 454,600 current and former students and exposed contact details, passport numbers, enrollment information, and fee payment records later appeared online. According to analysts, this breach is part of a larger wave of attacks targeting more than 100 organizations by ShinyHunters, exploiting CVE-2026-35273, a critical zero-day vulnerability in Oracle PeopleSoft that allows remote code execution. Check Point IPS provides protection against this threat (Oracle PeopleSoft Enterprise PeopleTools Server-Side Request Forgery (CVE-2026-35273)) Mackay Sugar, Australia’s second-largest sugar producer, has been hit by a cyberattack that disrupted operations and shut down its Farleigh and Racecourse mills in Queensland. The company instructed growers to stop harvesting and suspended cane haulage while temporary measures were deployed to maintain essential operations. Danish pharmaceutical giant Novo Nordisk has disclosed a breach after attackers accessed internal IT systems and copied pseudonymized clinical trial data from research systems. The exposed information included patient IDs, trial participation details, limited health data, and some healthcare professionals’ contact information. AI THREATS Check Point Research has demonstrated exploitable flaws in LangGraph, an open-source framework for stateful AI agents. Researchers chained SQL injection and unsafe deserialization issues to achieve remote code execution, with patches issued for SQLite, core, and Redis checkpointer components in affected deployments. Check Point IPS provides protection against this threat (LangChain LangGraph SQL Injection (CVE-2026-27022)) Researchers highlighted a China-based phishing-as-a-service network, Outsider, that allegedly used Gemini to generate fake websites and support SMS phishing campaigns. Google filed a lawsuit after linking the operation to thousands of phishing sites, more than 1.5 million URLs, and large-scale victim targeting. Researchers warned that prompt-injection attacks against Anthropic’s Claude Code GitHub Action could leak CI/CD workflow secrets. Malicious issue or pull request text can instruct the agent to read environment variables and expose API keys, enabling workflow abuse and impersonation inside software repositories. VULNERABILITIES AND PATCHES Check Point Research has identified active exploitation of CVE-2026-50751, a critical authentication bypass vulnerability affecting Check Point Remote Access VPN and Mobile Access deployments configured to use the deprecated IKEv1 key exchange protocol. Attacks began in May and increased in early June, affecting a limited number of organizations, with one case tied to Qilin ransomware activity. Check Point IPS provides protection against this threat (IKEv1 Remote Access Authentication Bypass PoC Exploit (CVE-2026-50751)) Microsoft released its largest Patch Tuesday update to date, addressing more than 200 Windows and Defender vulnerabilities amid an AI-driven surge in vulnerability discovery. The fixes include CVE-2026-45657, a critical Windows flaw with a CVSS score of 9.8 that could enable network-based propagation, CVE-2026-41091, which has been actively exploited to gain full system control, and CVE-2026-50507, a BitLocker bypass vulnerability. Veeam has released security updates to fix a critical flaw affecting Backup & Replication. The vulnerability allows an authenticated domain user to execute code remotely on a domain-joined backup server, exposing sensitive backup infrastructure and recovery systems. THREAT INTELLIGENCE REPORTS Check Point Research’s May 2026 attack trends report found that organizations experienced an average of 2,055 weekly attacks, down 7% month over month, while ransomware incidents increased 48% year over year. The report also highlights continued GenAI exposure across enterprise environments, including risks linked to business-related prompts. Researchers detected a supply-chain compromise in the Arch User Repository, where attackers seized hundreds of packages and modified build scripts to install credential-stealing malware. The campaign deployed malicious dependencies, a Rust stealer, and, with administrative privileges, an eBPF rootkit on Linux systems. Researchers analyzed a Brazilian phishing campaign abusing the legitimate NinjaOne remote management agent to gain access to company computers. The campaign uses fake Portuguese business portals and phone-based social engineering to install a signed agent connected to attacker-controlled infrastructure on victim endpoints Researchers described ongoing exploitation of WinRAR flaw CVE-2025-8088 by Russia-linked groups targeting Ukrainian military and government organizations. Spear-phishing archives plant hidden files that run at login and deploy stealers for browser passwords, cookies, VPN configurations, and other credentials across affected Windows systems.
research.checkpoint.comJun 15, 2026extracted
UK Government Finds 400+ Vulnerabilities in AI Hackathons
The UK government has discovered and patched hundreds of vulnerabilities after running a series of internal hackathons using frontier AI models. The weekly, in-person events were organized by the Government Cyber Coordination Centre (GC3) – an initiative from the National Cyber Security Centre (NCSC) and the Department for Science, Innovation and Technology (DSIT). The idea was to use the models to scan public code repositories across nine government departments. “Rather than mandate a single approach, we gave teams model access and let them build their own tooling, noticing what worked each week and building on the best approaches,” the GC3 said. Participants identified 407 findings, including critical flaws such as authentication bypass, data exposure and remote code execution. Although some were already known and mitigated by compensating controls, others were zero days, the report, published on June 21, claimed. All critical and high-risk weaknesses assessed as exploitable have been remediated, with no evidence of exploitation identified. “AI models traced vulnerabilities across service boundaries, which traditional scanners can’t do, and linked business logic with technical detail. Departments prioritized validation and remediation through existing frameworks,” the report noted. The various teams took different approaches. One created five new domain-specific Claude Skills to build a “reusable, scoped and consistent approach” across every open source repository and operator selected. Another used traditional scanning tools like Gitleaks, Trivy, Semgrep and Hadolint to generate initial findings. Then they applied models to these findings, to check against OWASP and CWE frameworks, compose individual findings into attack paths, and confirm viability through a triage stage. Another group built a six-stage agentic pipeline with each stage reading and challenging the last. Frontier Models Deliver Strong Performance The GC3 said it learned some important lessons through the hackathon initiative: The strongest results came from using frontier models as “tightly scoped components inside a structured pipeline” – with traditional vulnerability management workflows broken down into discrete, task-specific harnesses With the right architecture and task design many near-frontier and frontier models are similarly good at scanning code. Human expertise is still the difference, required to break problems down and identify wider context Triage is vital because agents generate candidate findings faster than humans can validate them. Careful upfront scoping and “structured internal filtering” improve focus and reduce costs. The whole project cost the government just £13,000 ($17,467) in tokens The next big job will be to integrate prioritization, review and patch-generation without “overwhelming human-centred processes” However, it’s unclear what impact a new US government export ban on Anthropic’s Mythos and Fable models will have on the government’s hackathon initiatives. The ban, which was brought in late on Friday, locks out all non-American users from the firm’s most powerful models.
infosecurity-magazine.comJun 15, 2026extracted
Open-source CI/CD abuse detector guards against stolen credential attacks
Open-source CI/CD abuse detector guards against stolen credential attacks CI/CD Abuse Detector is an open-source project that uses a large language model to flag suspicious changes to continuous integration and continuous deployment pipelines, workflows, and automation configurations. The repository contains drop-in templates for GitHub Actions, GitLab CI, and Azure DevOps. The project targets a common attack chain in software supply chain compromises. Stolen developer credentials are used to push modifications to workflow files, which then harvest secrets stored in the CI environment. The detector aims to catch these modifications during code review, before the altered workflow executes. How the analysis works The workflow runs in six stages. Changed files in a pull request are first matched against path patterns for CI/CD, build, release, and packaging configurations. Files that match are diffed individually, with each diff capped at 10,000 characters to reduce bypass attempts that hide malicious changes inside large benign additions. A prescreen step applies regex and metadata rules to attach context labels to each diff. The diff and labels are then sent to Claude through the Claude Code command line interface, which analyzes the content against a threat model focused on credential harvesting. Verdicts conform to a defined JSON schema. Output options include a GitHub step summary, repository issues, Slack notifications via webhook, and Elasticsearch verdict shipping when severity meets the configured threshold. An optional fail gate can block a pull request when severity exceeds a separate threshold. Default behavior alerts only. Setup requirements Teams adopting the detector copy three files into their repository: a workflow YAML file, a prompt markdown file, and a JSON schema for the verdict format. Authentication requires either an Anthropic API key or, for enterprise deployments, a Foundry endpoint URL and API key pair stored as repository secrets. Several environment variables tune behavior: CI_CD_ABUSE_ALERT_THRESHOLD sets the minimum severity for alerts and defaults to high. CI_CD_ABUSE_FAIL_ON_SEVERITY controls the blocking threshold and is empty by default, which keeps the detector in alert-only mode. CI_CD_ABUSE_INCLUDE_PUSHES enables analysis of direct pushes to main and master branches and defaults to true. Pre-processing in the templates relies on bash, jq, and grep. The Claude Code CLI, installed via Node, is the sole analysis dependency added to the CI environment. Python appears only in maintainer tooling for validation and the single-file build script, with no Python on the runtime path of the published workflows. Status and scope The repository is a prototype and reference implementation tied to Elastic Security Labs research titled “Detecting CI/CD pipeline abuse with LLM-augmented analysis.” Elastic describes the project as a prototype that sits outside its official product catalog, with limited support entitlements and no fixed roadmap. Teams interested in the underlying methodology can consult the linked Elastic Security Labs article and the architecture, threat model, and scoring notes documents in the repository. Vulnerability reports follow Elastic’s disclosure process documented in SECURITY.md. CI/CD Abuse Detector is available for free on GitHub. Must read: 25 open-source cybersecurity tools that don’t care about your budget GitHub CISO on security strategy and collaborating with the open-source community Subscribe to the Help Net Security ad-free monthly newsletter to stay informed on the essential open-source cybersecurity tools. Subscribe here!
helpnetsecurity.comJun 15, 2026extracted
Week in review: Exploited Check Point VPN zero-day, Oracle PeopleSoft servers under attack
Week in review: Exploited Check Point VPN zero-day, Oracle PeopleSoft servers under attack Here’s an overview of some of last week’s most interesting news, articles, interviews and videos: DockSec: Open-source AI-powered Docker security scanner DockSec is an OWASP Incubator Project that combines three container security scanners with a language-model layer for explanation and remediation. Created by Advait Patel, the Python tool runs Trivy, Hadolint, and Docker Scout against a developer’s Dockerfile and image, correlates the findings, returns a 0-100 security score, and proposes line-specific fixes. Treating AI agents like service accounts for federated query security In this interview with Help Net Security, Paras Malhotra, CISO at Starburst, explains how the company handles data governance across federated query environments. Topics include layering Starburst’s access controls above native source permissions, tiering vendor risk across more than 200 partners and connectors, and building audit trails for autonomous agents. NOVA microhypervisor brings AMD DMA isolation to shared AI infrastructure BlueRock has issued the latest open-source release of its NOVA Microhypervisor with DMA remapping support for AMD platforms that have IOMMU hardware virtualization. The capability is enabled by default and extends hardware-level isolation across virtual machines, devices, and memory in shared execution environments. The security in smartphones is helping send them to landfills The WEEE Forum estimated that 5.3 billion mobile phones became electronic waste in 2022. Many of these devices still function. The average smartphone stays in use for about three years, and owners often replace handsets that retain enough computing power for other jobs. A team at the Université Libre de Bruxelles examined a barrier to giving those devices a second life. Every set of AI guardrails can be broken by the right prompt AI companies use guardrails to block harmful outputs such as deepfakes, malware, and instructions for biological weapons or illicit drugs. A new mathematical proof by Apostol Vassilev, a senior scientist at NIST, suggests those protections have inherent limits. For any finite set of guardrails, there exists a prompt that can bypass them if discovered. NOVA microhypervisor brings AMD DMA isolation to shared AI infrastructure BlueRock has issued the latest open-source release of its NOVA Microhypervisor with DMA remapping support for AMD platforms that have IOMMU hardware virtualization. The capability is enabled by default and extends hardware-level isolation across virtual machines, devices, and memory in shared execution environments. The security in smartphones is helping send them to landfills Billions of working smartphones reach the end of their service lives each year and move into drawers, recycling streams, and waste piles. The WEEE Forum estimated that 5.3 billion mobile phones became electronic waste in 2022. Many of these devices still function. The average smartphone stays in use for about three years, and owners often replace handsets that retain enough computing power for other jobs. A team at the Université Libre de Bruxelles examined a barrier to giving those devices a second life. Every set of AI guardrails can be broken by the right prompt Companies that build AI systems wrap them in guardrails meant to block harmful output, including deepfakes, malware, and instructions for making biological weapons or illicit drugs. When a user prompts the system for such content, the guardrails are designed to flag the request and refuse. A new mathematical proof sets a limit on how secure those guardrails can ever be. CISA orders federal agencies to “patch smarter” The US Cybersecurity and Infrastructure Security Agency (CISA) has issued a Binding Operational Directive that will change how the US federal government approaches vulnerability management. How to use NIST and ISO frameworks to govern AI agents Security leaders no longer need convincing that AI agents introduce risk. What’s missing is how to govern them once they move into production and begin operating autonomously across enterprise environments. CISA: Patch actively exploited SolarWinds Serv-U DoS vulnerability (CVE-2026-28318) A vulnerability (CVE-2026-28318) that can be exploited to crash SolarWinds Serv-U file transfer servers is being leveraged by attackers in the wild, the US Cybersecurity and Infrastructure Security Agency (CISA) confirmed on Friday. The agency has ordered US federal civilian agencies to address it by June 19, 2026, either by implementing a patch or implementing mitigations. Qilin ransomware affiliate exploited Check Point VPN zero-day (CVE-2026-50751) A Qilin ransomware affiliate is believed to be exploiting CVE-2026-50751, an authentication bypass vulnerability in Check Point VPN Remote Access and Mobile Access, the company announced on Monday. Check Point Remote Access VPN enables and secures connections between corporate networks and remote or mobile devices. LiteLLM vulnerability under active attack, CISA warns (CVE-2026-42271) A command injection vulnerability (CVE-2026-42271) in BerryAI’s LiteLLM open-source AI gateway is being exploited by attackers, the US Cybersecurity and Infrastructure Security Agency (CISA) confirmed by adding the flaw to its Known Exploited Vulnerabilities catalog on Monday. Record Microsoft Patch Tuesday, fresh zero-day Microsoft marked its largest-ever Patch Tuesday this month, by shipping fixes for nearly 200 vulnerabilities. Within hours, “Nightmare Eclipse”, the researcher behind weeks of escalating Windows exploit releases, dropped a proof-of-concept exploit for a new zero-day: “RoguePlanet”, which abuses a race condition in Windows Defender to spawn a command shell running with SYSTEM-level privileges. Critical Ivanti Sentry flaw allows root-level remote code execution (CVE-2026-10520) Ivanti has patched two critical vulnerabilities (CVE-2026-10520 and CVE-2026-10523) in Ivanti Sentry and has urged customers to implement the fix right away. Though the vulnerabilities are not known to be actively exploited, security researchers have already released technical details about the former, which may be used by attackers to craft a working exploit. Oracle PeopleSoft servers under attack, Oracle pushes out-of-band security alert A zero-day vulnerability (CVE-2026-35273) in Oracle PeopleSoft PeopleTools is being exploited in the wild, Charles Carmakal, CTO at cybersecurity firm Mandiant, part of Google Cloud, warned today. The architecture of subtraction: Why it’s time to erase the roads, not just map the traffic AI-assisted vulnerability discovery and exploit development are making patching increasingly inadequate as a primary defense. Advanced AI models can shrink the time from vulnerability discovery to exploitation from months to hours, while organizations struggle to patch systems as quickly as new flaws are identified. Product showcase: Staying ahead of the threat horizon with Aunoo Aunoo is an open strategic intelligence platform that uses AI agents to monitor intelligence sources, including for cybersecurity, to compile a daily briefing and alert on defined criteria. Each source is checked for credibility and quality before it is included. The platform runs in any browser and can send its findings via Slack, Discord, Teams, email or using the internal chat. When attacks spread too far: Lessons from real cyber attack case studies In this Help Net Security video, Michael Adjei, Director, Systems Engineering at Illumio, explains three real world cyber attacks and what went wrong during detection. Cyber resilience metrics that drive action In this Help Net Security video, Pete Bowers, COO at NormCyber, explains how organizations can build a cyber resilience metrics program that supports better decisions. He questions common ways of measuring resilience, such as risk registers, tool scores, and annual tests, and points out their limits. GitHub Copilot app launches as desktop home for AI coding agents GitHub introduced the Copilot app, a desktop application built for working with AI coding agents, at Microsoft Build 2026. The release expands GitHub’s Copilot product line beyond editor integrations and command-line tools into a dedicated workspace for directing several agents at once. Cybercriminals create 19,000 FIFA-themed domains ahead of 2026 World Cup The 2026 FIFA World Cup will bring millions of visitors and an estimated 6 billion spectators to a tournament spread across 16 host cities in the United States, Canada and Mexico. In a new report, Intel 471 describes the 2026 FIFA World Cup as “the largest and most complex cyberattack surface in sporting history.” Hackers used Meta’s AI support system to hijack over 20,000 Instagram accounts Meta has revealed that attackers hijacked 20,225 Instagram accounts by exploiting a flaw in the company’s AI-assisted account recovery system. According to the company, a vulnerability in High Touch Support (HTS) allowed unauthorized parties to perform password resets on Instagram accounts. Microsoft changes how Defender for Endpoint EDR updates are delivered on Windows Microsoft will distribute Defender for Endpoint EDR updates through Microsoft Update, enabling EDR security improvements to be released independently of monthly Windows operating system updates. The rollout started for Windows 10 devices in late May 2026 and will expand to Windows 11 and other supported Windows versions later this year. Microsoft expects deployment to be completed by fall 2026. Meta claims NSO Group still targets WhatsApp users despite court order Meta claims it disrupted spear-phishing attempts linked to NSO Group and is asking a US federal court to hold the spyware vendor in contempt for allegedly violating an injunction that bars it from targeting WhatsApp and its users. Mythos Preview can weaponize N-day vulnerabilities in hours Mythos Preview can develop working exploits from newly disclosed software vulnerabilities in hours, cutting down a process that has historically taken days or weeks, according to Anthropic. Google patches Chrome zero-day exploited in the wild (CVE-2026-11645) Google has fixed 74 vulnerabilities in Chrome, including a high-severity zero-day (CVE-2026-11645) that has been exploited in the wild. The fix has been shipped in Chrome 149.0.7827.102/.103 for Windows and macOS and Chrome 149.0.7827.102 for Linux, with the update rolling out to users over the coming days and weeks. French government messaging platform breached through account hijacking French authorities are investigating a compromise of Tchap, the government’s secure messaging platform, after hackers hijacked a user account and gained access to public chat rooms. Anthropic’s Claude Fable 5 is out for public use, with safeguards for high-risk requests Days after publishing research on how advanced AI systems could amplify cyber operations in the wrong hands, Anthropic released Claude Fable 5, a Mythos-class model for general use. The company said Mythos-class models possess advanced cybersecurity and research biology capabilities that can provide information and guidance beyond what is typically available through conventional online sources. New Browser-in-the-Browser phishing uses fake login popups to steal Microsoft 365 credentials A new Browser-in-the-Browser (BitB) phishing campaign is targeting Microsoft 365 users with fake login popups designed to closely mimic legitimate browser authentication windows, according to Palo Alto Networks Unit 42. Identity theft is turning into a chain reaction for victims For a growing number of victims, identity theft no longer ends with a fraudulent charge or a compromised account. More than one in four people who contacted the Identity Theft Resource Center during the reporting period were dealing with multiple identity-related incidents, according to the organization’s 2026 Trends in Identity Report. X Square Robot open sources its robot-free data collection framework Companies building robots for physical work spend large amounts of time and money operating machines by hand to gather training examples. Each session with a physical robot produces a small number of demonstrations per day, which slows the growth of datasets used to train embodied AI. Human demonstrators offer a cheaper source of data, and X Square Robot has put a system for this approach into public release. Making the cloud prove it followed your privacy wishes Companies that store personal data in cloud key-value databases should handle deletion requests by running the operation and confirming the job is complete. The people making those requests and the regulators overseeing them have had limited means to confirm the data is gone or that the record of its removal is genuine. GDPRuler, a middleware system from researchers at the Technical University of Munich and the University of Lisbon, sits between an application and an unmodified key-value database and enforces privacy rules as data passes through it. 9 out of 10 people can no longer distinguish real from AI-generated content Online fraud is becoming harder to distinguish from legitimate activity as AI-generated messages, voices, photos, reviews, and identities become more convincing. Nearly nine in ten adults say they can no longer tell what is real from AI-generated content, according to the latest Malwarebytes survey. The share increased from 66% in 2025 to 85% in 2026. FBI seizes 13 websites linked to alleged Chinese intelligence-gathering effort Federal authorities have seized 13 internet domains allegedly used to target current and former U.S. government employees and military personnel with access to classified and sensitive information. 52% of direct-to-IP threats are missing from intelligence feeds Security tools are good at inspecting websites, domains, URLs, and files, so attackers are moving lower in the stack and communicating directly with IP addresses, where visibility is limited. According to Palo Alto Networks’ report, this creates a visibility gap that allows malicious traffic to blend into normal internet activity and evade detection. Google Colab CLI opens runtimes to Claude Code and Codex Google released the Google Colab Command-Line Interface, a tool that connects local terminals to remote Colab runtimes. The CLI provides an execution platform for developers and AI agents, letting users provision compute, run local Python scripts on remote runtimes, and retrieve artifacts back to local machines. OpenAI is locking down parts of ChatGPT to reduce data theft risks OpenAI has started rolling out Lockdown Mode for ChatGPT, an optional security setting that restricts access to external resources and several product capabilities. It is available for personal accounts, including Free, Go, Plus, and Pro plans, as well as self-serve ChatGPT Business accounts. Samsung just made Galaxy phones more secure in One UI 9 beta Samsung’s One UI 9 beta integrates Lockdown mode into the power menu. This is the screen that contains Power off, Restart, and emergency options. Opening it initiates Lockdown mode, disabling biometric authentication. The security questions around Chinese AI coding models in U.S. software Software developers across the United States are using AI models built in China to write, debug, and review code, drawn by prices below those of American alternatives. These models carry risks for the security of American software, according to a report from Booz Allen Hamilton, which tested how the models respond when the user appears to work for the U.S. government. Malware ships with bugs that defenders could use against it Static analysis tools have spent years scanning legitimate software for security bugs before it goes out the door. The same scanners work on malware, and malware carries a steady supply of its own bugs. Researchers ran four of these tools across 658 leaked malware projects and found that close to 90 percent contained at least one recognized software weakness. Apple expands what parents can block, approve, and limit Apple has previewed a set of new child safety features coming to iPhone, iPad, and the Mac later this year, expanding parental controls with tools that help families manage app access, web browsing, communication, and screen time. Apple Intelligence can now replace weak passwords without user intervention Apple’s next generation of Apple Intelligence, the company’s personal intelligence system, expands its capabilities and introduces new security features in Passwords. With the new update, Passwords can automatically replace weak or compromised passwords. Scams now operate like real businesses with budgets and targets Social media has overtaken email as a primary attack vector, showing changes in how people consume information and interact online, according to Bitdefender’s Global Scam Intelligence Report 2026. Fraud campaigns use advertisements, sponsored content, impersonation pages, and direct messages to reach users. Apple extends Private Cloud Compute to third-party data centers Apple is bringing its Private Cloud Compute (PCC) platform to Google Cloud, expanding the infrastructure behind Apple Intelligence to third-party data centers. Introduced in 2024, PCC provides cloud-based processing for AI workloads that exceed the capabilities of on-device models while maintaining Apple’s security and privacy guarantees. Building reusable workflows with custom agents in Copilot CLI Developers spend much of their working time in the terminal, generating commands, debugging issues, and running scripts close to their systems. Repeated terminal work tends to pile up small steps such as re-running the same commands, re-explaining context, and translating logs into a form a team can act on. Custom agents in GitHub Copilot CLI address these patterns by turning repeated tasks into reusable workflows. Organizations can’t see much of their mobile AI activity Organizations have limited visibility into AI activity on mobile devices despite security leaders expressing confidence in their AI governance, according to Lookout’s “Solving for the Mobile AI Blind Spot: Executive Confidence Meets Technical Reality” report. Prompt injection still drives most agentic AI security failures in production A backdoor sat on PyPI for three hours in March 2026. Nearly 47,000 downloads occurred during the window. The compromised package, LiteLLM, serves as the language-model gateway for CrewAI, DSPy, Microsoft GraphRAG, and dozens of other AI agent frameworks. Anyone pulling an update during that window pulled in an autonomous attack bot named hackerbot-claw along with it. Threat actors are recruiting the people who hold cloud logins Companies keep most of their data and applications in cloud platforms that anyone can reach with the right login. That setup turns each employee holding those credentials into a security variable, and members of the cybercrime underground have built methods to reach those people. Intel 471 tracked this activity into 2026 and sorted insider risk into three categories that cloud-reliant organizations contend with. Fake Spotify Premium tutorials on TikTok and Instagram Reels spread malware Cybercriminals are using TikTok and Instagram Reels videos to spread Vidar, an infostealer malware, through fake downloads for popular paid software, according to ReversingLabs. The researchers uncovered two campaigns behind the activity, each using a different approach to draw in viewers before sending them to external download sites. Google sues China-based scammers over Gemini AI abuse Google has filed a lawsuit against Outsider Enterprise, a China-based cybercrime network for using AI tools, including Gemini, to build phishing websites and scam infrastructure. Cybercriminals are moving away from mass phishing campaigns Phishing activity declined by roughly 20% in both 2024 and 2025, according to research from Zscaler’s ThreatLabz team. The drop followed years of growth that pushed phishing activity above 2 billion hits in 2023. Authorities dismantle crypto laundering service that moved €336 million for cybercriminals An international law enforcement operation has dismantled a cryptocurrency laundering service linked to ransomware groups and other cybercriminals that processed more than €336 million in illicit funds. Cybersecurity jobs available right now: June 9, 2026 We’ve scoured the market to bring you a selection of roles that span various skill levels within the cybersecurity field. Check out this weekly selection of cybersecurity jobs available right now. New infosec products of the week: June 12, 2026 Here’s a look at the most interesting products from the past week, featuring releases from AISLE, Drata, Elastic, Filigran, IDnow, and Ridge Security.
helpnetsecurity.comJun 14, 2026extracted
Industry Reactions to Claude Fable 5: Feedback Friday
Claude Fable 5 has become generally available, with Anthropic unveiling it as a powerful Mythos-class AI model. The release includes robust safeguards that restrict its capabilities in high-risk domains. In sensitive areas such as cybersecurity (where it could be misused to create exploits) and biology (where it could assist in developing bioweapons or chemical weapons), Fable 5 automatically falls back to the less capable Claude Opus 4.8. Anthropic stated that it performed extensive internal and external red-teaming to ensure the model is highly resistant to jailbreaking. [ Read: Anthropic Disputes Fable 5 AI Jailbreak ] Industry professionals have commented on various aspects of the new Fable 5, including dual-use offensive and defensive cyber capabilities, safeguards, tiered access for select partners, a premium price tag creating a security poverty line, and the urgent need for proactive AI governance and faster defender adaptation. And the feedback begins… Greg Heon, VP, Product Strategy, Armadin: “The same massive investments that made AI models dramatically better at writing code have made them dramatically better at finding and exploiting vulnerabilities — those are two sides of the same capability, and the labs have poured tens of billions of dollars into it. Every enterprise should now be preparing for machine-speed, AI-orchestrated hyperattacks: campaigns that chain reconnaissance, discovery, exploitation, and lateral movement faster than any human defender can react. Preparing isn’t a tabletop exercise. It means testing your real attack surface against these techniques, and it means starting on the perimeter today — not just running tests in sandboxed pre-production environments that look nothing like what a real attacker sees. The frontier labs are gating their most capable models specifically because of cyber risk. That should tell every CISO exactly where this is heading — and why the time to test against it is now, not after tomorrow’s AI-powered adversary launches a hyperattack.” Myke Lyons, CISO, Cribl: “This is the emerging trend: develop cutting-edge models, highlight their risks, release a ‘safer’ version to the public, and reserve the unrestricted version for select partners. Anthropic’s rollout mirrors this pattern. Expect OpenAI, Google, and Meta to follow suit, creating a tiered model ecosystem. This isn’t just about safety. It’s about positioning. The real question for enterprises isn’t whether their AI vendor includes safety mechanisms, but whether they’re prepared to handle the unrestricted tier. On the defensive side, Fable 5 enables capabilities like long-term threat monitoring, large-scale account research, and automation of complex processes. On the offensive side, Mythos-class models demonstrate sophisticated agentic hacking capabilities, including autonomous reconnaissance, lateral movement, and exploitation. The most concerning aspect is the imbalance: defenders are constrained by procurement cycles and compliance processes, while attackers only need an account. AI capabilities are advancing faster than security teams can adapt. Security leaders need to treat this as a wake-up call: AI governance must be dynamic and proactive, not reactive. Falling behind now means playing catch-up indefinitely.” Ben Bernstein, Cybersecurity Advisor, Huntress: “Fable comes with a serious premium price tag compared to standard public models, which instantly prices out a lot of smaller organizations. We’ve dealt with this ‘security poverty line’ for years when it comes to prohibitively expensive security tooling, but Fable is really just the latest iteration of that exact problem. The danger isn’t just that these smaller teams are missing out on a cool new tool; it’s that threat actors are using these AI advancements to drastically accelerate how they hunt for the same low-hanging fruit they always have: misconfigurations, exposed systems, and unpatched vulnerabilities. So, while the Fortune 500 and well-funded cyber criminals, organized crime, and nation-states are leveraging this premium tier of AI to either defend or attack at machine speed, historically under-resourced teams are going to be facing a massive, automated wave of threats without the budget for the advanced security tooling, or the human talent, required to keep up.” Noelle Murata, Chief Operating Officer at Xcape, Inc: “Anthropic’s broad commercial release of Claude Fable 5 represents a calculated pivot in the frontier AI landscape: attempting to monetize elite, long-horizon reasoning architecture while strictly walling off its most “hazardous” capabilities. By implementing an aggressive, real-time classifier system that automatically downgrades high-risk cybersecurity, biochemical, or model-distillation requests to the less powerful Claude Opus 4.8 framework, Anthropic is trying to fulfill its commercial obligations without turning a public LLM into an on-demand zero-day factory. However, this bifurcated release strategy highlights a growing divergence in enterprise defense. While everyday enterprise customers gain access to Fable 5’s highly advanced software engineering and long-running autonomous logic, Claude Mythos 5 remains exclusively accessible to a tight cohort of government intelligence agencies and select critical infrastructure defenders under Project Glasswing. This means the actual “cybersecurity tier” of this technology remains behind sovereign closed doors, leaving commercial security teams to defend against an increasingly automated threat landscape without the same unrestricted analytical tools being deployed by nation-state actors.” Varin Khera, Co-Founder and CTO, SECStrike.ai: “Anthropic has reported a roughly 5% false positive rate for the Fable 5 model, and I would expect those safeguards to improve over time. However, in our testing, we observed significantly more instances where legitimate security prompts triggered the guardrails, terms central to routine defensive work like CVEs and impact analysis frequently triggered the fallback mechanism, routing queries to Claude Opus 4.8. The challenge is that cybersecurity professionals are locked out of the model precisely when their work demands it most.” Jacob Krell, Senior Director: Secure AI Solutions & Cybersecurity, Suzu Labs: “Anthropic filed for its IPO on June 1 and launched Fable 5 eight days later at double the Opus token rate. The benchmark gains are real but concentrated in frontier-hard tasks. SWE-bench Pro jumps 11 points, from 69.2% to 80.3%. On routine work the gap shrinks to near-parity, and cost-per-solve still favors Opus 4.8 at $1.45 vs $2.49 per solved task. The token economics compound the pricing. Fable 5 burns tokens at twice the Opus rate. A BleepingComputer reviewer exhausted a $100 daily allocation in nine minutes running Anthropic’s workflow mode. At $10/$50 per million tokens, heavy agentic work can clear three figures a day. I do complex offensive cybersecurity tasks on Opus 4.6. No cybersecurity classifier. No mandatory data retention. Fable 5 charges double, blocks those queries, and redirects them to Opus 4.8. Anthropic needs to show public-market investors it can monetize a $965 billion valuation. Fable 5 doubles per-token revenue. The cybersecurity gains are locked behind Project Glasswing. Everyone else pays double and gets Opus 4.8 responses on security queries.” Gidi Cohen, CEO & Co-founder, Bonfy.AI: “The most honest thing Anthropic has done here is ship one model as two products. Splitting Fable 5 and Mythos 5 is an acknowledgment that capability and safety are in genuine tension — and that pretending otherwise doesn’t serve anyone. But the most important line in the entire announcement isn’t about the classifiers. It’s buried in the operational detail: a high-severity vulnerability found by the model takes about two weeks to patch on average. Meanwhile, Mythos Preview built working exploits from a disclosed CVE in under a day. That gap is where risk lives. And no classifier closes it. This makes concrete what the CSA data showed last week: enterprises aren’t failing because they can’t detect vulnerabilities. They’re failing because they can’t act on them fast enough. AI has collapsed the attacker’s timeline to hours. The defender’s timeline hasn’t moved. Anthropic is right that the defensive head start only matters if the industry uses it. The harder truth is that most enterprises aren’t yet equipped to — not because the tools don’t exist, but because the governance architecture to deploy them safely hasn’t kept pace with the capability.” Devin Maguire, Senior Manager, Product Marketing, Cycode: “Anthropic released Mythos more broadly in the form of Claude Fable 5. Models are getting dramatically better at finding vulnerabilities. That’s genuinely exciting progress. But better models don’t make the security team’s job easier. They make it harder. The same capability lands in the hands of attackers. And the flood of new CVEs that follows moves faster than any team can manually triage. The 2026 Verizon DBIR made this concrete. For the first time in 19 years, vulnerability exploitation is the #1 way organizations get breached. 31% of all breaches. Median time to patch: 43 days. The bottleneck has never been finding vulnerabilities. It’s always been knowing which ones are actually exploitable in your environment, and fixing them before attackers get there. Vulnerabilities found by AI still need to be managed. They need to be analyzed, triaged, assigned, remediated, and tracked. Another detection tool in the arsenal is also another tool in the adversary arsenal and doesn’t solve the persistent security challenge of managing risk posture and fixing what you found. Every leap in model capability widens that gap. The organizations that close it will be the ones that treat remediation speed as a security metric, not an engineering backlog. Congratulations to the Anthropic team. The hard work starts now for the rest of us.” Etay Maor, Vice President of Threat Intelligence, Cato Networks: “Anthropic’s Claude Fable protections are good, and they will stop many of the direct attempts to get the model to do malicious things. For an opportunistic attacker — somebody who doesn’t have a lot of time, resources, or willingness to keep trying — those safeguards can be effective. […] When we’re thinking about safeguards, we have to remember that the capabilities are in the model and the protections are layered on top. Those protections are important, but they’re not the same thing as removing the capability itself. That’s why I describe them as speed bumps rather than barricades. They can slow attackers down, and that’s valuable, but they’re unlikely to stop the threat actors we worry about most — the ones with the time, resources, and motivation to keep testing until they find another way in. From an enterprise perspective, the 30-day retention requirement deserves attention. Organizations in regulated industries need to understand exactly what data is being retained and whether that aligns with their compliance and legal requirements before they start using these models in sensitive environments. The other thing that stands out is the agentic component. The more autonomy you give an AI system and the more access you give it across infrastructure, code repositories, and internal systems, the more valuable it becomes to both defenders and attackers. If that system is manipulated or compromised, it can become a very effective tool for lateral movement. Most organizations are still figuring out what security looks like in a world where AI agents can take actions across multiple systems on their own.” Roger Grimes, CISO Advisor, KnowBe4: “Regarding whether cybercriminals will get access to these tools faster: no, not really. Criminals have been using AI to find vulnerabilities, code exploits, and code malware since last year. Certainly, learning about Mythos put a renewed, more intense push on using AI to find vulnerabilities and exploit them, but it wasn’t like it hasn’t been what the elite cybercriminals haven’t been doing for a year already. They have been doing this. Heck, I saw similar non-AI versions of Mythos being used by nation-states and large red teams over a decade ago. They were pretty good then, but now AI-enabled, they are supercharged. The only thing Mythos substantially changed was how quickly the defenders would get these tools. Sure, it accelerated and helped attackers, but they didn’t need the push. Defenders needed the bigger wake-up call. There are in fact no downsides to making Fable-5 public. The sooner the band-aid is ripped off, the sooner the defender lifecycle kicks in and helps us. What Mythos kicked off was defenders getting more secure code sooner. Mythos and Fable will help defenders get more secure code faster. We will see a huge spike in vulnerabilities found and exploited over the next 2-3 years, and after that, we will see more secure applications. The way this will change the cybersecurity industry is that we will see more usage of AI to both find and fix vulnerabilities faster, patch faster, and instruct the AI to code more securely from the start. The net result of Mythos and Fable is more secure apps.”
securityweek.comJun 12, 2026extracted
Bernie Sanders’ AI Sovereign Wealth Fund Plan
Bernie Sanders’ AI Sovereign Wealth Fund Plan Let no one accuse Bernie Sanders of ducking the big questions. Writing in the New York Times last week, the senator asked: “Will the future of humanity be determined by a handful of billionaires who have promoted and developed AI, with virtually no democratic input, who stand to become even richer and more powerful than they are today?” We agree entirely that this is one of the most potent questions facing global democracy today. Our book, Rewiring Democracy, surveys the emerging uses for and impacts of AI in democracy around the world and reaches the same conclusion: that the most urgent risk posed by AI is the concentration of power, wealth and control among tech oligarchs. And yet we reached a vastly different conclusion than Sanders on what to do about it. The senator points to a once radical but increasingly popular solution: creating a US sovereign wealth fund by taking 50% stock in AI companies such as Anthropic, OpenAI and xAI. The argument in favor of this is twofold. One: it would establish democratic control over the AI companies, giving the government “the power, through its voting shares and an equal representation on each company’s board, to block decisions that hurt our citizens and to push for policies that help them.” Two: it would return a big chunk of the economic rewards of soaring AI valuations to the public, ensuring “trillions of dollars potentially generated by AI are used to improve the lives of all of us.” We laud both these goals unreservedly. We wholeheartedly agree that there must be public influence over the development and use of AI, just as we demand the government intervene to ensure that automakers, drugmakers, airlines and other industries balance profitability with public safety and the public interest. And we credit the senator with recognizing that there are more levers for the government to pull beyond the promulgation of regulation to achieve this. And we also agree that the obscene, dangerous accumulation of wealth among AI companies needs to be disrupted. As OpenAI and Anthropic race to be minted as the world’s latest trillion-dollar AI companies, we should recognize that—whether or not it constitutes a bubble—these staggering market capitalizations represent a transfer of wealth. The flow of money goes from the smaller businesses and actual people using AI, and being subjected to it, to the owners of these tech companies. That includes the world’s 86 AI billionaires “seeking to maximize their power and profit” aiming to decide the “fate of humanity… behind closed doors in Silicon Valley,” as Sanders said. And yet, while we do not outright oppose the taking of AI company stock, or of a US sovereign wealth fund, there are better ways to achieve Sanders’ stated goals. Public ownership of these companies entangles corporate profit and valuation with the public interest. It would incentivize the government to clear regulations, permit the exploitation of workers and users, suppress competition, encourage AI adoption regardless of the responsibleness of the implementation or appropriateness of the use case, and otherwise act on behalf of corporate interests. After all, if growing, say, Nvidia from its first $5tn in value to its next $5tn also represents a doubling in value of this segment of the sovereign wealth fund, then you can expect the fund managers to support chip sales, foreign and domestic, with the same zeal as the company’s private investors. This is not an effective way to influence corporations to act in the public interest. In fact, it makes corporate influence on the government more likely. We should be wary of this possibility because we’ve seen it before. Ownership of substantial stakes in oil companies by the Norwegian sovereign wealth fund, the world’s largest, does not seem to have steered those corporations to pro-environmental policies. Instead, the Norwegian government’s dependence on those companies has inhibited them from taking climate action. Here in the US, public employee pension funds merit the same criticism: the fiduciary duty to generate wealth overwhelms any intention to direct their corporate holdings in the public interest. A better answer is to separate the two goals. The standard way to share private rewards with the broader society that made them possible is taxation. Senator Elizabeth Warren has proposed an excise tax on datacenters’ energy use. Others have proposed an AI token tax, which has much the same effect. As to the goal of reshaping AI in the public interest, we have proposed an AI Public Option. The concept is for governments, be it federal or state, to establish publicly developed and operated AI models run by public institutions under democratic control. The idea is not to eliminate corporate AI or to seize it as a public asset, but rather for government to provide a competitive baseline that private AI offerings must meet or exceed to win business—just like the notion of a healthcare public option. The Swiss have trailblazed this approach. Apertus is a large language model built by Swiss public servants, researchers at Swiss universities, using appropriately licensed training data and pre-existing Swiss public supercomputing infrastructure powered by renewable energy. While Apertus doesn’t seriously compete with the latest OpenAI and Anthropic models on performance benchmarks, it blows them out of the water in transparency, sustainability and compliance with EU regulations including adherence to copyright. It’s a nascent project, but suggestive of how public institutions can apply competitive pressure for corporate actors to behave responsibly. Don’t confuse public AI with “sovereign AI,” the notion that every country needs to invest in domestic AI infrastructure. Sovereign AI is often invoked as a marketing scheme for big tech companies looking to sell to governments; it demands public investment without guaranteeing public control. Sanders is a bold and savvy political operator. So why is he pursuing the sovereign wealth fund strategy when he must be aware of these risks? It may be due to another argument he makes in his op-ed: that the Trump administration and the billionaire owners of AI are aligned to the idea. It’s expedient to capitalize on rare moments of seeming alignment across diverse political factions, but it also behooves us to ask why the AI billionaires are open to this extraordinary intervention. The answer, of course, is that they believe that for every dollar ceded to government stock expropriation, they will get back more in favorable government policies to protect that newfound investment. Energy taxation is a straightforward way to make AI companies pay for the social disruption of their technologies. Public AI represents a non-monetary mechanism for governments to shape the development of AI, complementary to direct regulation of private actors, one with a far greater chance of influencing corporate behavior towards the public interest. We urge Sanders and other political leaders to consider them. This essay was written with Nathan E. Sanders, and originally appeared in The Guardian.
schneier.comJun 12, 2026extracted
Fable 5 Mythos violato in 24 ore: il caso che scuote la cybersecurity
A sole 48 ore dall’uscita di Fable 5 di Antropic e dalle nostre contestuali riflessioni della cybersecurity dei sistemi di classe Mythos, una prima conferma ai timori di molti esperti. Indice degli argomenti Il ricercatore Pliny the Liberator è riuscito in 24 ore a fare jailbreak di Fable 5 riuscendo a estrarre le istruzioni segrete del modello. L’attacco non è stato opera di un semplice “script kiddie”, ma di un’offensiva architettata in modo quasi militare. Pliny ha utilizzato una strategia definita “Pack Hunt” (caccia in branco). Invece di usare un singolo prompt, ha coordinato molteplici agenti IA in parallelo, ciascuno specializzato in ruoli come ricognizione, intrusione e copertura delle tracce. Questo attacco distribuito ha generato segnali simultanei che hanno letteralmente saturato e confuso i meccanismi di rilevamento lineari di Anthropic. I vettori di attacco hanno sfruttato tecniche avanzate: manipolazione testuale: sostituzione di caratteri Unicode e alfabeto cirillico per ingannare i filtri basati su parole chiave. decomposizione e ricomposizione: invece di chiedere istruzioni per un malware, le richieste sono state frammentate in concetti accademici benigni, come se l’attaccante avesse usato una sorta di social engineering verso l’AI (es. gestione della memoria Linux), che i filtri hanno lasciato passare, per poi essere riassemblati in exploit operativi. Il colpo finale è stata l’estrazione e la pubblicazione dell’intero System Prompt di Fable 5: un file di 120.000 caratteri contenente le istruzioni segrete dell’azienda. La fuga di notizie del System Prompt, analizzata nel dettaglio da TechX, ci permette di fare un’analisi interessante del “cervello” della macchina. Le scoperte rivelano le vere priorità della Silicon Valley: la paranoia del copyright: mentre le direttive sulla sicurezza informatica e sulle armi biologiche sono scritte con tono professionale, la sezione sul diritto d’autore è l’unica redatta interamente in MAIUSCOLO e definita come “limite invalicabile”. Il modello ha il divieto assoluto di superare le 15 parole per citazione diretta e di generare testi di canzoni o poesie coperte da copyright. Questo dimostra che il vero terrore delle corporazioni non è la fine del mondo, ma le cause legali miliardarie per violazione della proprietà intellettuale. “Claudeception” (l’autonomia strutturale): il prompt conferma che Fable 5 è programmato per richiamare altre istanze IA (come Sonnet 4) tramite API per gestire compiti in background e creare Artifacts complessi. Accesso Linux nativo: l’IA gira in un container Ubuntu 24 con permessi operativi per creare, leggere ed eseguire comandi bash, confermando la sua interessante natura agentica. Nonostante l’allarmismo sui social (“ANTHROPIC PWNED”), Pasquale Pillitteri riporta la vicenda nei giusti binari, smontando l’hype con lucidità. Pubblicare un system prompt è un danno d’immagine devastante, ma non significa aver preso il controllo dell’intero modello. Soprattutto, i presunti output letali diffusi da Pliny non sono mai stati verificati da fonti indipendenti. Tuttavia, Pillitteri solleva il vero “punto critico” di questa vicenda: il design del fallback silenzioso. Per proteggersi dalla “distillazione” (il furto di capacità da parte di nazioni rivali o concorrenti), Fable 5 declassava le richieste sospette al modello inferiore (Opus 4.8) di nascosto. Utenti e ricercatori si sono visti bloccare compiti innocui (come l’analisi di un emocromo) o degradare la qualità del codice senza alcun avviso. Come evidenziato dal ricercatore Nathan Lambert, un’IA che diventa “meno intelligente in automatico senza avvisarmi” è un prodotto categoricamente inaffidabile. Di fronte alla rivolta degli sviluppatori, Anthropic è stata costretta ad un mea culpa. L’azienda ha ammesso di aver scelto “il compromesso sbagliato” e ha modificato l’infrastruttura, rendendo il declassamento a Opus 4.8 visibile e fornendo le motivazioni esatte dei blocchi tramite API. Per comprendere la gravità di quanto accaduto, bisogna capire la potenza della macchina che è stata “bucata”. Ethan Mollick, ricercatore che ha avuto accesso anticipato al modello, ha descritto Fable 5 non come un semplice chatbot, ma come un salto paradigmatico. Messo alla prova su compiti di programmazione avanzata e analisi dati, il modello non si è limitato a rispondere a un prompt: ha redatto un documento di design di 19 pagine e ha scritto codice in totale autonomia per ben nove ore e mezza, creando un software complesso chiamato Concord. Il rapporto tra uomo e intelligenza artificiale, secondo gli addetti ai lavori, è cambiato radicalmente con queste nuove versioni. Non siamo più “maghi” che guidano il software riga per riga, ma “mecenati”. Commissioniamo un lavoro, paghiamo in token e attendiamo il risultato. L’IA prende centinaia di decisioni invisibili, trasformandosi in una vera e propria scatola nera. Ad avviso di chi scrive, aver esposto il codice interno di una “scatola nera” così potente è l’equivalente di aver diffuso le planimetrie di un contesto nucleare. Avevamo raccontato l’ingresso nell’era degli HACCA (agenti cyber-offensivi altamente autonomi) e avevamo portato all’attenzione, seppur invitando a non abbassare la guardia, il sofisticato compromesso ingegneristico di Anthropic: un sistema di sicurezza basato su classificatori in tempo reale e su un “freno d’emergenza” capace di declassare le richieste pericolose su Claude Opus 4.8. Oggi, mi corre l’obbligo di riaprire quel capitolo. Quell’infrastruttura di sicurezza, venduta come un impenetrabile caveau digitale e testata per oltre 1.000 ore, ha retto l’urto della rete per circa 24 ore. Il clamoroso jailbreak messo a segno dal ricercatore noto come Pliny the Liberator, unito alle preziose analisi di esperti del settore che hanno sezionato l’accaduto nelle ultime ore, ci costringe a integrare la nostra riflessione.Purtroppo, il castello di carta della sicurezza proprietaria è crollato sotto i colpi di un attacco distribuito. E’ sinceramente avvenuto molto velocemente, ma molti di noi lo avevano ipotizzato e fatto intendere nelle riflessioni scritte e verbali sull’argomento. Questo incidente non può essere sottostimato. La comunità tecnologica, quella cyber e la politica internazionale devono alzare immediatamente il livello di guardia. È necessaria un’azione legislativa e tecnica per pretendere audit indipendenti obbligatori sui sistemi di sicurezza degli HACCA. L’illusione che una singola azienda possa arginare l’evoluzione delle armi cybernetiche autonome con un semplice “freno d’emergenza” è svanita nel giro di due giorni. L’era del cyber-uranio è qui, e stiamo scoprendo, a nostre spese, che il contenitore che lo ospita è pieno di crepe.
cybersecurity360.itJun 12, 2026extracted
AI Broke Vulnerability Management. That's Why CISOs Are Moving Budget to BAS.
For thirty years, vulnerability management ran on a buffer: the months between when a vulnerability was found and when someone could figure out how to weaponize it. The solution was straightforward enough; triage by severity, schedule the fix, validate, and move on. The buffer was what made that work. Today, that buffer is gone. AI didn't make your team slower. It changed the other side of the equation, compressing discovery-to-exploit from months to hours. And the sad truth for defenders is that a process built for breathing room can't survive without it. AI Turned Vulnerability Discovery Into a Volume Game In its May 2026 update, Anthropic reported that it and approximately 50 partners used Claude Mythos Preview to find more than 10,000 high- or critical-severity vulnerabilities in systemically important software in a single month. Earlier figures were just as stark. Pointed at Firefox, the gated Mythos model wrote 181 working exploits, against just 2 from the previous frontier model. It surfaced vulnerabilities across every major OS and browser, including an OpenBSD bug that had sat undetected for 27 years. At the time of writing, more than 99% of what it found was still unpatched. An AWS threat-intelligence report from February 2026 shows the flip side: no zero-days needed, just weak credentials, industrialized through a custom MCP server running offensive tools autonomously. AWS confirmed 600+ devices across 55+ countries; the actor's logs, according to independent researchers, queued 2,516 devices across 106 countries. Either way, the rules have clearly changed. What once took rare expertise now runs at machine speed and scale. The Vulnerability Weaponization Window Has Collapsed, Too Defenders used to have months between a CVE going public and its first confirmed exploitation in the wild, the window known as time-to-exploit (TTE). That window has slammed shut. Zero Day Clock puts the 2026 average at roughly 24 hours, down from ~53 days in 2024. The breach data agrees, too. Verizon's 2026 DBIR ties 32% of initial-access techniques to exploitation of vulnerabilities and expects that number to climb, because AI coding assistants now put exploit-building, porting a tool to a new language, and discovering fresh flaws all within reach for attackers who've never had them before. Telling Teams to Patch Faster Is Like Telling a Freighter to Brake on a Dime The industry's reflex answer is to patch faster. Regulators are codifying it: Many regulations now point toward same-day fixes for some critical vulnerabilities. Boards expect it. Executives demand it. But remediation isn't a switch. Patches clear regression testing, wait for change windows, need to wait for approvals, and respect existing uptime and compliance commitments. Taking production down to outrun an exploit ends up being just a different outage. And the data shows everything's moving the wrong way. The Verizon 2026 DBIR tracked 13,000+ organizations: Median fix time for known-exploited vulnerabilities: 43 days, up from 32 the year before Amount that were fully patched: down from 38% to 26% When offense runs in hours and remediation runs in weeks, the breach almost always happens in between. Again, per Verizon's DBIR, even the best-performing organizations close only 30-40% of known-exploited vulnerabilities in the first week after detection: a rate that's barely moved despite years of steady investment. So, ordering teams to patch faster doesn't change the physics, and it feels like ordering a freighter to brake on a dime. The Bottleneck Moved. So Must the Strategy. For two decades, vulnerability management ran on a tidy set of assumptions: Find the flaws, Score them by severity, Patch the worst first. When a few dozen criticals landed per quarter, CVSS triage worked. Unfortunately, it doesn't stand a chance against hundreds or thousands of disclosures a day. Dipping back to Verizon's DBIR one more time, the median organization had to patch 16 known-exploited vulnerabilities in 2025, up from 11 the year before, a jump of nearly 50%. That was before AI-discovered flaws began flooding the catalog. Severity scores, meanwhile, don't tell you whether a flaw is reachable in your environment, whether your controls will already block it, or whether it chains to anything that matters. A severity list where everything is a "9" or "10" essentially prioritizes nothing. So the useful question stops being "what's vulnerable?" and becomes "what's actually exploitable against us right now: and would our defenses catch it if someone tried?" This is exactly the question Breach and Attack Simulation (BAS) was built to answer. Why BAS Becomes the Cornerstone Against AI-Powered Attacks BAS takes real-world adversary techniques, the TTPs behind the campaign in the latest headline, and safely runs them against your live prevention and detection stack. Not a scan. Not a theoretical mapping. An actual exercise that shows what your tools will actually block, what they'll detect, and what will slip through. In a world drowning in disclosures, that does three things that vulnerability management alone can't. BAS: Separates the theoretical from the real. A flaw your WAF, IPS, and EDR already neutralize is a very different problem from one that waltzes straight in. BAS shows which is which, so teams stop treating every CVE as a five-alarm fire. Validates the controls you've already paid for. Most enterprises run anywhere from ten to seventy security tools with countless overlapping policies; BAS measures whether they fire as configured and surfaces the residual risks hiding in the gaps. Buys time to patch safely. When you can prove a critical asset is already covered by hardened controls, the patch can move through normal change control instead of an emergency rollout. When it isn't covered, you know to mitigate first. That payoff is starting to show up in budgets: field reports increasingly point to CISOs reserving dedicated spend for BAS that wasn't a separate line item a year ago. This is the shift Gartner now labels Adversarial Exposure Validation: blending security effectiveness ("Are my controls working?") with business context ("Which assets matter most, and what's truly reachable?") to prioritize by your organization's reality instead of by hypothetical raw scores. Paired with autonomous penetration testing, which proves whether an attacker can chain exposures from their initial foothold to your organization's crown jewels, BAS completes the picture. One side asks, "Wait, can they breach us?" The other asks, "But would we catch it?" Running together, BAS and autonomous pentesting replace guesswork with evidence. BAS Has to Run Autonomously at Machine Speed Too There's a catch. If adversaries are operating autonomously, a validation cycle that takes a human a week to complete is obsolete on arrival. Machine-speed attacks demand machine-speed defenses, and the only thing fast enough to counter autonomous offense is autonomous defense. The honest objection to pointing raw generative AI at this is safety. As Picus CTO Volkan Erturk has warned, a model told to invent an exploit might hand back a live malware sample, or hallucinate techniques a group never uses. You don't want unvetted binaries detonating in production, or defenses built against attacks that don't, or can't, exist. Picus' fix is to put the model in charge of coordination, not creation. Rather than asking AI to write payloads, Picus' agentic BAS matches a fresh threat report against a curated, pre-vetted library of safe, ready-made test building blocks. A security team names a threat, and a multi-agent system takes it from there: one agent identifies the threat and builds a research plan, others gather and validate the intelligence from multiple sources, and a builder agent maps the adversarial TTPs into attack chains ready for simulation. The output is an accurate, ready-to-run simulation, assembled in minutes. This collapses the loop. A CISA alert or a forwarded headline becomes a scoped test, a posture score, prioritized mitigations, and an executive report, often in minutes, with humans reviewing exceptions rather than driving, and slowing down, every step. This Is What the Picus Platform Is Built For Patching is still essential, but where AI discovers flaws by the thousands and weaponizes them in hours, patching alone can't be your whole strategy. If the offense is autonomous, the defense has to operate at least at the same speed, and that's exactly what Picus was built to do. What scales with the threat is validation: confirming what your controls will actually stop, proving what's exploitable, and spending remediation time and talent only where it will change the outcome. AI-powered, agentic BAS is one of the core pillars of the Picus Platform, continuously testing whether your defenses block and detect what matters without waiting on a human to kick off the process or advance to the next cycle. And when a gap is uncovered, the platform points to the vendor-specific mitigation needed, and doesn't just create another ticket on the pile, then re-validates to confirm that the gap has actually been closed. The need to say, on the spot, whether a fresh headline puts the business at risk isn't going away anytime soon. The Picus Platform gives security teams that answer before anyone asks. Find out if the next headline puts you at risk, before it drops. Request a demo. Note: This article was written by Sıla Özeren Hacıoğlu, Security Research Engineer at Picus Security.
thehackernews.comJun 11, 2026extracted
The guide on blocking ChatGPT, Gemini, Claude, and other AI tools at work | Kaspersky official blog
Unchecked AI in the workplace quickly becomes a massive loophole for data leaks and security breaches. All too often, employees drop sensitive company data into public chatbots, or install rogue AI assistants on their own — in the process handing over way too much access. In a previous post, we broke down the different types of risky AI systems, and later shared some tips on how to turn off the built-in AI features on major tech platforms. Today let’s take a look at practical ways to block or restrict the unauthorized “helpers” employees might be using — from ChatGPT and Grammarly, to meeting bots like Fireflies and Read AI. How to detect and restrict ChatGPT ChatGPT is the biggest culprit when it comes to unauthorized AI use worldwide. A quick word of warning, though: an outright ban only sends users hunting for sketchy third-party sites or messaging app chatbots that hook into the same service. That’s why it’s always a good idea to offer an approved alternative before pulling the plug. Detecting it: keep an eye on the NGFW or web filter for traffic heading to chat.openai.com, chatgpt.com, oaistatic.com, oaiusercontent.com, or cdn.oaistatic.com. It’s also smart to use EDR/EPP tools to scan browser histories, installed apps, and browser extensions across corporate devices. Locking it down: use the firewall or web filter to block the entire AI Services category, and set up DNS to reroute traffic away from those OpenAI domains. Browser policies can also be used to ban ChatGPT-powered extensions. Better yet, block all extensions not on a pre-approved allowlist. Finally, use application controls and EPP solutions to stop users from installing the official desktop app (ChatGPT.exe or com.openai.chat). How to detect and restrict Claude and Claude Code Detecting it: use the NGFW or web filter to track traffic going to claude.ai, anthropic.com, *.anthropic.com, and api.anthropic.com. EDR/EPP or application control tools can also be used to scan employee computers for the desktop app (claude.exe). Locking it down: drop a blanket block on the AI Services category through the NGFW or web filter, and tweak DNS settings to reroute traffic away from the aforementioned Anthropic domains. Next, use browser policies to shut down Claude-powered extensions. Finally, use application controls and the EPP platform to prevent users from installing the desktop app. How to detect and restrict Perplexity AI Detecting it: keep tabs on the NGFW or web filter to flag any traffic heading to *.perplexity.ai or pplx.ai. Locking it down: just like the others, add the AI Services category to the NGFW or web filter blocklist, and use DNS routing to redirect traffic away from those domains. Configure the browser to block third-party extensions from being installed. If Firefox is used in the organization, be aware that recent versions come with Perplexity built in. Luckily, these AI features can be turned-off company-wide using enterprise policies — specifically, by setting SidebarChatbot = blocked. The full list of tweaks can be found in the Firefox documentation. How to detect and restrict DeepSeek Detecting it: keep an eye on the NGFW or web filter for traffic hitting deepseek.com, chat.deepseek.com, api.deepseek.com, or platform.deepseek.com. For better precision, analyze the SNI (server name identification) in TLS connection requests. For mobile devices, look out for the official app (com.deepseek.chat). Locking it down: blocklist the AI Services category on the NGFW or web filter, and reroute traffic to DeepSeek’s domains via DNS settings. Use browser policies to block third-party extensions, and lean on MDM/EMM tools to restrict the mobile app. How to detect and restrict Mistral, xAI Grok, and Character.ai The playbook for these tools is exactly the same as DeepSeek, so here’s the quick list of domains to watch for and block: chat.mistral.ai, mistral.ai, console.mistral.ai, grok.com, x.ai, api.x.ai, character.ai, beta.character.ai, and c.ai. A quick word of warning on Grok: because Grok is baked into X, blocking this specific AI access point means blocking the entire social media platform. How to detect and restrict Slack AI Detecting it: in the Slack workspace admin dashboard, look under Analytics → Slack AI usage. If an enterprise plan is used, the detailed Slack logs can be searched for any events starting with the ai_ prefix. Blocking it with policies: in the organization’s Slack settings, click through the Workspace settings → Roles & permissions → Feature access, and change the permission to “no one”. Slack has a step-by-step guide in their help center. Locking it down: shutting this down at the network level is tricky; it can be pulled off with a finely tuned CASB solution in place. Also, don’t forget the importance of blocking rogue integrations and keeping external AI services from tapping into Slack data in the first place. We covered how to lock this down using OAuth controls in a previous post. How to detect and restrict Zoom AI Companion Detecting it: if a corporate Zoom subscription is in use, just head to Admin Center → Reports → AI Companion usage. Detecting Zoom’s AI when employees join external meetings or use free accounts is a lot tougher, but email filters can be set up to flag incoming AI-generated meeting notes by scanning for subject lines or text containing “Meeting summary” or “Meeting assets”. Blocking it with policies: for the company’s own Zoom subscription, go to the Admin Portal → Account Management → Account Settings → Meeting → AI Companion and toggle it OFF for everyone. Locking it down: unfortunately, AI Companion is baked into Zoom’s DNA, so the only real option is blocking Zoom altogether. How to detect and restrict Grammarly What looks like an innocent spellchecker is actually one of the biggest culprits for workplace data leaks. Detecting it: check the NGFW or web filter logs for traffic hitting grammarly.com, *.grammarly.com, and gnar.grammarly.com. EDR and MDM/EMM tools can also be used to hunt down the standalone desktop apps (Grammarly Desktop.exe and the macOS version), as well as the Grammarly browser extension. Locking it down: use firewalls to block those domains at the network level, and EPP to stop employees from installing the desktop app, browser extensions, or the Grammarly add-ins for Microsoft Word and Excel. How to detect and restrict meeting assistants: Fireflies, Read.ai, Tactiq, Fathom, and Granola This massive category of third-party SaaS tools records and analyzes meetings — creating a massive risk for data leaks. The trickiest part? Outside clients or vendors can bring these bots into a meeting just as easily as employees can. Detecting them: run an audit on calendar invites, and look for bot participants using email domains like @fireflies.ai, @read.ai, @tactiq.io, @fathom.video, or @granola.ai. Zoom, Teams, or Google Meet logs can also be used to review external participants who joined past calls. Locking them down: since it’s impossible to control what outsiders do, blocking these bots comes down to tightening meeting rules. The best moves are: blocking users from granting OAuth permissions for bots to join calls, restricting employees from inviting unapproved external participants, or locking down meeting recording access for external users. That last option is usually the least painful way to keep bots out without disrupting business. How to detect and restrict AI code editors: Cursor, Windsurf, and the like Detecting them: use EDR/EPP tools to scan for executables like cursor.exe or windsurf.exe. It’s also worth monitoring network traffic heading to cursor.com and windsurf.com, as well as traffic hitting various AI model API providers. Keep in mind that there’s a pretty extensive list of API hosts to monitor here, since these editors aren’t tied to just one specific AI vendor. Blocking them with policies: these apps can be prevented from being installed by setting up filters based on the developer’s digital signature certificate. Alternatively, a strict application allowlist can be employed where only pre-approved software is allowed to run. Locking them down: rely on the EPP/EDR platform to actively detect and block these applications from running. How to detect and restrict local AI tools: Ollama, LM Studio, and GPT4All On one hand, this category carries fewer data leak risks because the AI models run completely locally on the user’s machine. On the other hand, it opens up a whole new can of worms: these apps themselves aren’t always highly secure, and can become targets for cyberattacks. Plus, it still means that employees can misuse models or process data in unauthorized ways. Detecting them: EDR/EPP tools are the best line of defense here. They should be used to flag known local AI files and processes like ollama.exe, ollama serve, lmstudio.exe, LM Studio.app, jan.exe, or gpt4all.exe. From a network perspective, it’s worth scanning for open ports on local devices — typically port 1234 for Ollama and LM Studio, or port 8080 for WebUIs (using an additional fingerprint check of the server response). Another massive red flag is the presence of large files (often several gigabytes) containing language model weights. Look out for extensions like .gguf, .bin, or sometimes .safetensors. Locking them down: use EPP/EDR platforms or windows AppLocker to block these applications by name, or switch to an application allowlist. How to detect and restrict autonomous agents: OpenClaw, NemoClaw, and NanoClaw This is easily one of the most dangerous categories of AI tools out there. These agents mix high-level independence with access to untrusted data, making them a massive security headache. Detecting them: use EPP/EDR tools to sniff out active processes like openclaw, nanoclaw, nemoclaw, or clawdbot. Also keep an eye out for devices running Node.js that suddenly start launching Bash or Python scripts. Another dead giveaway is the appearance of system folders like ~/openclaw, ~/nanoclaw, ~/.claw*, or ~/clawhub. At the network level, monitor connections to the AI model APIs we mentioned earlier, as well as traffic hitting servers like openclaw.ai, nanoclaw.dev, or clawhub.*. Locking them down: the safest bet is to use strict application allowlisting (only allowing approved software to run), or to specifically ban the known agent apps listed above. On top of that, consider blocking non-developers from installing Node.js and Docker, neither of which they need on their computers anyway.
kaspersky.comJun 10, 2026extracted
Record Microsoft Patch Tuesday, fresh zero-day
Record Microsoft Patch Tuesday, fresh zero-day Microsoft marked its largest-ever Patch Tuesday this month, by shipping fixes for nearly 200 vulnerabilities. Within hours, “Nightmare Eclipse”, the researcher behind weeks of escalating Windows exploit releases, dropped a proof-of-concept exploit for a new zero-day: “RoguePlanet”, which abuses a race condition in Windows Defender to spawn a command shell running with SYSTEM-level privileges. Various researchers have confirmed that the PoC exploit works to achieve local privilege escalation. “In initial development, it was confirmed that this vulnerability was a remote code execution,” Nightmare Eclipse noted, but said that a Windows Defender patch Microsoft pushed out in May might have made remote code execution impossible. Priorities in a record-breaking release This month’s Patch Tuesday releases address vulnerabilities in a wide variety of Microsoft’s products, but some require more immediate attention than others, especially in this age of AI-powered security research: CVE-2026-42897, an actively exploited Microsoft Exchange Server vulnerability, now has a fix. “As part of our ongoing efforts to strengthen security and improve defenses across environments, we continue to enhance protections for cross-site scripting attacks. We recommend that customers keep CVE-2026-42897 mitigation in place,” Microsoft’s Exchange Team advised. “The mitigation provides an additional layer of defense and helps ensure continuous protection as further improvements are released. Additional updates will be shared as they become available.” CVE-2026-45586, a privilege escalation vulnerability in Windows Collaborative Translation Framework (CTFMON), may allow authenticated attackers to gain SYSTEM privileges. The vulnerability is publicly disclosed and Microsoft deems it “more likely” to be exploited. (This is believed to be the vulnerability exploited by Nightmare Eclipse’s “GreenPlasma” exploit.) CVE-2026-49160, a remotely exploitable vulnerability that affects HTTP.sys, the Windows kernel-mode driver responsible for intercepting and handling network requests over HTTP and HTTPS, may lead to denial of service condition and is also publicly disclosed. Dustin Childs, head of threat awareness at Trend Micro’s Zero Day Initiative, pointed out that systems using the default MaxRequestBytes registry value used by the Windows HTTP stack are not affected by this flaw. “You can edit your registry settings if you need protection while you test and deploy the patch. The bulletin includes instructions and even a PowerShell script for doing this action. Microsoft lists this as ‘Exploitation more likely’, so I would definitely check your registry settings,” he opined. CVE-2026-50507 is a Windows BitLocker bypass that can only be exploited by attackers who have physical access to target devices. CVE-2026-45585, another Windows BitLocker bypass, has also received a fix. Microsoft acknowledged in the security advisory that this is the fix for the vulnerability exploited by Nightmare Eclipse’s “YellowKey” exploit. (Microsoft shared mitigation advice for it in May 2026.) Childs also singled out as priority patches two unauthenticated code execution flaws that can be exploited remotely without user interaction: CVE-2026-44815, in the DHCP Client Service, which is present and active on every OS. CVE-2026-45657, a wormable Windows Kernel bug that stems from how the kernel handles TCP/IP. “This was listed as ‘Exploitation Less Likely’ by Microsoft, but rest assured that every researcher and bug shop on the planet is reversing this patch right now trying to create an exploit. Test and deploy this patch quickly,” he advised. The AI-driven patch flood isn’t going away “Last month, Microsoft published a blog noting the increase in reporting volume over several years and that both its engineers and the security community are ‘increasingly using AI’ to find bugs,” Satnam Narang, senior staff research engineer at Tenable, told Help Net Security. With this in mind, and as more advanced AI models become available, a large (and increasing) volume of patches may become the norm, and not just for Patch Tuesday. “With nearly 200 CVEs patched this month, I would be remiss not to call out recent reporting by the Anthropic Frontier Red Team, which highlighted the threat posed by N-days – known vulnerabilities that have not been fully remediated across systems,” he added. “As part of its analysis of N-days, Anthropic’s Frontier Red Team analyzed 21 Windows kernel elevation of privilege vulnerabilities included in the January and February 2026 Patch Tuesday releases. Models including Sonnet, Opus and Mythos Preview were able to produce proof-of-concept (PoC) exploits by performing patch diffs to identify what changed between the previous and the latest release. Mythos Preview even produced PoCs for 13 of the 14 vulnerabilities that were labeled as ‘Exploitation Less Likely’ or ‘Exploitation Unlikely’ according to Microsoft’s Exploitability Index, an assessment system designed for humans, not advanced AI models. As Anthropic prepares to release Mythos, and other AI companies release models on par with Mythos, rapidly closing the patch gap is critical for organizations.” Tyler Reguly, Associate Director, Security R&D at Fortra, also noted that widespread AI use is making CVSS scores a poor indication of real risk. “How many of [vulnerabilities with high CVSS scores] are turned into exploits and how many of those exploits are the thing we really need to pay attention to. For the next few weeks, while teams are testing the patches and preparing for deployments across their organizations, I’ll be watching CISA KEV to see if any of these get added,” he commented. “Right now, I’m guessing that all three publicly disclosed vulnerabilities will end up on the list – CVE-2026-45586 (CFTMON), CVE-2026-50507 (Bitlocker), and CVE-2026-49160 (HTTP.sys).” Trend Micro’s Childs pointed out that this “inflation” of patches raises concerns: “How many patches were generated using AI to assist in coding or testing? What quality issues may exist in these patches?” Also: “Should sysadmins adjust their processes for prioritization and patch deployment based on this new volume of updates? Unfortunately, Microsoft is not providing those answers right now. Hopefully that changes in the future.” UPDATE (June 11, 2026, 05:10 a.m. ET): A comment from the Microsoft Exchange Team was added. Subscribe to our breaking news e-mail alert to never miss out on the latest breaches, vulnerabilities and cybersecurity threats. Subscribe here!
helpnetsecurity.comJun 10, 2026extracted
OpenSSL Patches High-Severity Vulnerability Found With AI
The latest OpenSSL releases patch 18 vulnerabilities, including a high-severity issue that could allow remote code execution. The high-severity vulnerability, tracked as CVE-2026-45447, is a heap user-after-free bug in a function used for PKCS#7 (Public-Key Cryptography Standard #7) verification. Discovered by a Calif researcher in collaboration with Claude AI and Anthropic Research, the bug can be triggered using a specially crafted PKCS#7 or S/MIME signed message during PKCS#7 signature verification. “When processing a PKCS#7 or S/MIME signed message, if the SignedData digestAlgorithms field is present as an empty ASN.1 SET, OpenSSL may incorrectly free a caller-owned BIO during PKCS7_verify(). A subsequent use of the BIO by the calling application results in a use-after-free condition,” OpenSSL developers explained. Exploitation of the vulnerability can result in heap corruption, process crashes, and possibly in remote code execution. The moderate-severity flaws patched in OpenSSL can be exploited to decrypt encrypted communications, forge arbitrary ciphertexts, launch DoS attacks, bypass integrity validation, and execute arbitrary code. One of the medium-severity weaknesses can be exploited to trick a system into accepting a fake, attacker-controlled certificate and private key, allowing the attacker to bypass authentication mechanisms with a 1-in-256 success rate. The low-severity vulnerabilities can lead to crashes (DoS), message forgery, recovery of private keys, replacement of root CA certificates, and possibly arbitrary code execution. Alex Gaynor of Anthropic has been credited with reporting half a dozen of the newly patched vulnerabilities, suggesting that the AI giant’s Mythos model may have helped identify the flaws. High-severity vulnerabilities in OpenSSL are rare these days. Only one high-severity issue was patched last year, and CVE-2026-45447 is the second high-severity flaw of 2026. In April, OpenSSL developers patched a flaw that can allow an attacker to obtain sensitive data. Related: Drupal Patches Highly Critical Vulnerability Exposing Websites to Hacking Related: Google Patches 5th Chrome Zero-Day Exploited in 2026 Related: Android Update Patches Exploited Zero-Day, 123 Other Vulnerabilities Related: Oracle’s First Monthly Patches Resolve 77 Vulnerabilities
securityweek.comJun 9, 2026extracted
XBOW tests Anthropic's Mythos Preview for offensive security
We received early access to Mythos Preview for early capability testing a few weeks back. In this article, we can finally share what we found. About three months ago, Anthropic invited us to help them assess the capability of a new model they thought represented a significant shift in capability. So we put it through our security gauntlet. Benchmarks, workflows, interactive use, and integrations. Below are the details on how we tested Mythos Preview, what we found, and what it means. Spoilers: This model is a major advance. It is substantially better than prior models at finding vulnerability candidates, especially when source code is available. It communicates with unusual technical precision, reasons well about code, and shows strong promise in complex domains such as native-code analysis and reverse engineering. Our takeaway: Mythos Preview is a powerful tool for generating strong vulnerability leads and technically precise analysis. It is especially adept at analyzing source code with a security mindset. It's not magic, though: a model is a brain without a body. While source code audits are mostly a brain activity, live site pentests like the ones XBOW performs very much need a body whose skill and control can match the brain's power. Testing methodology The first thing we did was assemble a diverse team of 10 experts from different parts of the company that could assess the model from different directions. We test all models with the same internal benchmarking system we have used to analyze Opus 4.7 and GPT 5.5. In this system, we take open source applications where vulnerabilities were previously discovered, freeze them at the vulnerable version, and run our agents against them. But this time, we expanded our testing to analyze other angles as well: The model’s judgment with regard to threat modeling, vulnerability validation, and safety The model’s ability to read source code versus interact with live systems Its ability to find exploits we’re not yet looking for in our standard assessments, e.g., native app vulnerabilities A note on terminology: When people say “Mythos,” they sometimes refer to the raw model. In this evaluation, we explored Mythos Preview both inside Claude Code, and as a raw model, using it via its API as an engine for XBOW’s agents. We separate those cases because orchestration, tools, prompting, and live-site access materially affect outcomes. Results Our testers who tried out Mythos Preview in interactive use were quite impressed. “This is a lot closer to just go and find something than anything I’ve seen so far,” said one of them. We tried giving it our own source code, and it found weaknesses – nothing truly terrible, thankfully, but there were several items we wanted to repair. We tried it on open source software, and at the end of week one, we had quite a few new vulnerabilities we had to disclose. Our testers who tried out Mythos Preview on benchmarks were also quite impressed, but their appreciation was a slightly different kind: impressed _with data_. Their results also laid bare the difference between areas where the model was runaway powerful, and where it presented only a modest advance. Finding a vulnerability isn't the same as proving it's exploitable. See how XBOW orchestrates frontier models with live-site validation to prove which findings are real, with working exploit evidence. Mythos Preview Benchmark Performance Our key takeaways after analyzing Mythos Preview include: It’s extremely powerful for source code audits. It’s good, but less powerful, at validating exploits. Its judgment is mixed. It can be too literal and conservative, and also tends to overstate the practical relevance of its findings. It’s strong in native-code vulnerability discovery and reverse engineering. Next-level vulnerability discovery Mythos Preview presents a significant step up over all existing models, regardless of provider, on XBOW’s web exploit benchmark. This benchmark is designed to test whether a model can help XBOW find validated, actionable vulnerabilities in live website environments. A case is counted as passed only when the system finds a validated way to act on the vulnerability (PoC||GTFO) after a series of 80 “actions,” where an action might be a shell or a Python script using standard commands or XBOW’s suite of attack tools. Note: We haven't included Opus 4.7 in this chart because that model interacts with our system in a unique way, making this particular stat less relevant for it – we’ve written up the full story here. Compared to the newest model at the time (Opus 4.6), this was a strong increase: The number of false negatives was cut by 42%. In a variation where we gave both models the site’s source code, it was even cut by 55%. This was the first instance of a theme that would surface again and again: Mythos Preview is impressive at writing code, but even more impressive at reading it. Below are the pass rates of Mythos Preview, Opus 4.6, and GPT 5.5 as a function of the allowed number of actions (executed scripts). Mythos Preview finds vulnerabilities in significantly fewer iterations than Opus 4.6, although the difference to GPT-5.5 is less pronounced. It becomes more clear when adding two considerations: Models could choose many small steps or few large steps (more details here) – and that shouldn’t matter so much. Instead of giving a budget of actions, let’s consider a budget of output tokens. Instead of mean pass rate, i.e., the probability of finding a vulnerability, it’s often more instructive to look at the odds for discovery, i.e., what ratio you would bet on the model getting a discovery right. Computationally, that's the hit rate divided by the miss rate. Under these considerations, the picture becomes much more clear: Token-for-token, Mythos Preview hones in on the vulnerability with absolutely unprecedented precision. Live-site validation is the hard part Mythos Preview is excellent at source-code reasoning, but our evaluation reinforced a practical truth: many exploitable issues do not appear as obvious defects in application source code. They emerge from configuration, dependencies, deployment choices, or the way otherwise safe components are combined. For instance, a dependency on its own could be safe. The source code on its own could be safe. But the source code uses the dependency in an unsafe way and creates a vulnerability. As Gary McCraw famously declared, you won’t find the majority of defects by “staring at code” alone. That’s of particular interest to us. XBOW performs pentests, where our target is a live site (the way an attacker sees it), whereas Mythos Preview as used, for example, by Project Glasswing excels at auditing source code (the way a developer sees it). Interacting with the live site can be very powerful, but it brings a completely new, very delicate dimension into the mix. Does Mythos Preview change the balance here? Due to the way we harvest our web benchmarks set, you can actually find the vulnerability from the code alone on that set. So it’s fair to ask: For these benchmarks, can Mythos Preview find an exploit without being allowed to interact with the live site? It turns out that even for these benchmarks, where the vulnerability is purely in the code, removing access to the live site hurts performance more than removing access to source code. In many ways, live-site access matters more than source-code access. That, of course, is the XBOW value proposition: it gives frontier models a safe, structured way to interact with real application behavior and prove which findings are actually exploitable. The results of XBOW powered by Mythos Preview are shown below. We now have a solid answer to the question, “Can a model find something interesting in code?” Increasingly, the answer will be yes, even though “something” won’t be the same as “everything.” But even then, the question still looming is, “Which of these findings are exploitable, reproducible, safe to test, and worth fixing?” The answer lies in combining Mythos Preview’s powerful source code analysis with something like XBOW’s ability to analyze a live site safely, in an orchestrated, validated way. It’s notable that, even though Mythos Preview suffers greatly from being denied access to the live site, other models suffer even more. Another confirmation that Mythos’ greatest strength is reading source code. The best results are always, of course, with the combination of access to the live site and source code. It allows the ideal detection pattern when XBOW orchestrates Mythos Preview: Analyze the source code to find a lead, probe the live site to understand how the weakness is reflected in the deployment, then craft an exploit from it. Other findings We also explored the model in terms of judgment, reverse engineering, assessment of native apps, and visual acuity. Judgment results were mixed Mythos Preview’s judgment results were more mixed than its discovery results. Across command safety, threat modeling, and trace triage, it was often careful and precise, but also literal and conservative. It rejected false positives better than many predecessors, but sometimes lost true positives when evidence did not formally satisfy its criteria or when the intended rule was broader than the written one. This makes Mythos Preview valuable, but not self-sufficient: it needs precise prompts, explicit threat models, and validation infrastructure to turn strong reasoning into reliable security outcomes. One bit that slightly shocked us here was Mythos Preview’s performance on our command safety benchmark, where we ask the models to consider whether a given script is safe to execute without impacting the target site. We hand-labeled a large set of example cases close to the edge of the decision boundary, and Haiku 4.5 delivered 90.1% accuracy. We also optimized the prompts for Haiku 4.5, so the better comparison is Opus 4.6, which had a 81.2% accuracy … but Mythos Preview had only 77.8%. When we probed deeper and looked at its reasoning, it would often have a point. There were cases that technically weren’t against the letter of the rules, but they were against the spirit. Opus 4.6 prioritized the spirit, but Mythos prioritized the letter. The model is strong in native code and reverse engineering Beyond web applications, the model showed substantial strength in native-code vulnerability discovery and reverse engineering. In Chromium-related testing, it found more real bugs with fewer false positives than prior baselines. In V8 sandbox work, it identified true positives in a subtle threat model where previous approaches had produced many findings but no successful true positives. It also proved capable of triaging both its own results and competitor-model findings. The reverse-engineering results were among the most striking. The model reasoned through unusual firmware and embedded systems contexts, including architectures and operating-system combinations that required more than rote pattern matching. Browser interaction and visual acuity are strong enough for practical workflows XBOW’s workflows often require models to interact with live websites through a browser interface. In that setting, visual acuity is important: the model needs to identify the right UI element and click in the right place. The evaluated model performed extremely well on XBOW’s visual-acuity QA, roughly matching Sonnet 4.6 and dramatically outperforming Opus 4.6. It was not perfectly pixel-accurate when asked for exact coordinates, but it was practically effective at selecting the right browser actions. We should note that Opus 4.7 also shone at this benchmark. Maybe the real story here isn’t “Mythos Preview is good,” but more: This is a specific area where recent Anthropic models had begun to deteriorate. But now Anthropic has caught that deterioration and reversed it. Power at a cost Mythos Preview is not just any new model: it’s a true titan. But titans are big, and big means expensive. How much money are you willing to spend on how much assurance? Can you spend that same money differently to get better results? At the time of writing, Mythos Preview is not yet available over public APIs, but Anthropic did mention that it would be 5x as expensive as an Opus model – already one of the more expensive options, token for token. Begging the question: Could we give an agent powered by a different model more time , and still get more accuracy for less cost? As it turns out: yes. If we normalize by estimated running cost, the picture is rather clear: Mythos Preview isn’t terribly inefficient, at least if you desire high accuracy, but it’s not best-in-class on our benchmarks either. This finding lines up with similar comparisons, e.g. Point Estimate’s analysis of the AI Security Institute’s benchmarking of Mythos Preview vs GPT-5.5: Mythos Preview is powerful, but the real choice is to either pay for an agent to use Mythos Preview for a bit, or to use GPT-5.5 for as long as needed. The better option depends on the use case; often, it’s the latter. XBOW’s evaluation suggests that frontier models have taken a major step forward in vulnerability discovery. Mythos Preview is strong at finding candidate vulnerabilities, especially from source code, and shows impressive ability across web, native-code, and reverse-engineering tasks. But it needs to be mounted in the right harness and equipped with the right tools to reach its full potential. And even then, it should just be one of the arrows in your quiver – depending on the task, it may be more sensible to let another model try several times than to let Mythos Preview try once. Such considerations, after all, are one of the reasons XBOW maintains a cadre of models, rather than restricting itself to a single one. To see XBOW’s powerful vulnerability validation capabilities in practice, please contact us for a demo. Sponsored and written by XBOW.
bleepingcomputer.comJun 9, 2026extracted
Claude Mythos Turns N-Days Into N-Hours With Rapid Exploit Creation
Anthropic says its Claude Mythos Preview model can build working exploits targeting known vulnerabilities within hours, or even minutes. Announced in early April and promoted as the most capable AI frontier model, Mythos right from the start raised fears regarding its ability to supercharge attacks. In April and May, Anthropic touted its ability to find vulnerabilities, including 271 Firefox flaws and thousands of severe security defects across over 1,000 open source software (OSS) projects. Now, the company says its most advanced model can also weaponize these discoveries, demonstrating that the surge in AI use in cyberattacks increases the threats faced by organizations in the patch gap. Put to the test, Claude Mythos Preview delivered 16 working exploits targeting Firefox and Windows within hours. Anthropic’s public models were also tested, with safeguards off. While they did not rise to Mythos’s level, they too delivered working exploits, proving that LLMs significantly increase the threat posed by N-days that have not been exploited in attacks before. According to Anthropic, N-days are even more dangerous than zero-days, because attackers can patch diff and reverse-engineer them to build exploits. This is exactly where LLMs become valuable weapons to attackers, as they significantly accelerate and automate the process of building N-day exploits. “Exploit development is not the only step in a real N-day campaign (target discovery, delivering the exploit to the target, and detection evasion all take time and resources too), but historically it has been the step most bottlenecked by scarce reverse engineering expertise,” Anthropic explains. PoC for Firefox vulnerability in 8 minutes To validate the theory, the company tested Mythos Preview, Opus, and Sonnet’s ability to construct proof-of-concept (PoC) code targeting 18 security patches delivered for SpiderMonkey in Firefox 148 and 149. They all delivered within minutes. Opus 4.8 created 11 PoCs, while Mythos Preview produced 14. Opus 4.8 delivered the first PoC in eight minutes, while Mythos Preview created it in 12. Anthropic also tested the models’ ability to turn crashes into working exploits. Mythos Preview built eight of them, Opus 4.8 two, and Opus 4.6 and Sonnet 4.6 one each. “This is where Mythos Preview really pulled ahead. Mythos Preview wrote its first working exploit in just under one hour, and ultimately created eight different exploits in roughly 12 hours,” Anthropic says. 8 Windows exploits in 18 hours Next, the company tested the LLMs’ ability to build exploits for closed-source software, and chose Microsoft’s Windows platform for the task, looking at 21 kernel vulnerabilities disclosed between January and February 2026. “This is substantially harder: with no source code available, the agent must work from compiled binaries and decompiler reconstructions that have been stripped of helpful context, like variable names, types, and structure,” Anthropic notes. Sonnet 4.6 and Opus 4.7 built PoCs that triggered BSOD for 13 of the bugs, Opus 4.8 for 15, and Mythos Preview for 18. Mythos Preview delivered the first PoC in 31 minutes. Mythos Preview was also able to create working exploits leading to privilege escalation for eight of the vulnerabilities, and delivered all of them within 18 hours. According to Anthropic, because it typically takes seven days before Windows patches are pushed to 90% of enrolled devices in a fleet, and because they are typically force-rebooted only on day 11, the model makes exploitation viable within the patch gap. Faster patching amid low exploit costs “At this speed, Mythos Preview would have finished creating all eight full chain exploits before any of the Windows devices had received the patch as an update. Turning these exploits into a real campaign still requires further work, but Mythos Preview has now collapsed one of the most time-intensive steps into hours,” Anthropic notes. The cost of building these exploits is not high either, the company says. Each model was given a three-million-token budget for creating the PoCs and exploits targeting Firefox. The cost of creating the full chain exploits targeting Windows was $15,700 in API credits, or around $2,000 per privilege escalation. “The binding constraint to N-days is now just a few thousand dollars and API access, which expands the pool of capable N-day attackers dramatically,” Anthropic says. The company calls for an updated patching playbook, which should rely on “N-hour” rather than “N-day”, and should no longer assume that weaponizing a patch takes weeks. “N-days have historically caused most harm to systems that are slow or difficult to patch. Industrial control systems, medical devices, and ‘internet of things’ devices often run on fixed maintenance windows, vendor-locked firmware, or have uptime guarantees. As the cost of weaponizing any given patch falls toward zero, these devices and systems will become even more exposed. And even systems operating on an established, ‘responsible’ patch cadence are now far easier targets than before,” Anthropic notes. Related: Anthropic Expanding Mythos Access to 150 New Organizations Related: Mythos Proves Potent in Vulnerability Discovery, Less Convincing Elsewhere Related: The Mythos Moment: Enterprises Must Fight Agents with Agents
securityweek.comJun 9, 2026extracted
OpenAI Rolling Out ChatGPT Account Security Controls
OpenAI told SecurityWeek that it’s making two ChatGPT security controls more widely available, giving users additional tools to protect their accounts and data. One of the features is Lockdown Mode, which enables owners of ChatGPT accounts, including personal and self-serve Business accounts, to reduce the risk of data exfiltration from prompt injection attacks. “Lockdown Mode is designed to help prevent the final stage of data exfiltration from a prompt injection attack by limiting outbound network requests that could transfer sensitive data to an attacker,” OpenAI explains. “Lockdown Mode does not prevent prompt injections from appearing in the content ChatGPT processes.” Enabling Lockdown Mode disables or limits capabilities such as live web browsing, image support, deep research, agent mode, canvas networking, and file downloads. The AI giant noted that the feature is not intended for all users and organizations, only those that handle highly sensitive data and require extra protection against potential data exfiltration conducted through prompt injection. Lockdown Mode can be enabled in Settings> Security> Advanced Security. The second feature is Active Sessions, which enables ChatGPT users to review where their account is signed in. Users can see the sessions and devices they are logged into, and log out of sessions they don’t recognize. The feature is available for all ChatGPT accounts and workspace types, except accounts linked to an organization’s SSO setup. Active Sessions is available in Settings> Security. The announcement comes after OpenAI unveiled a new account security feature for ChatGPT users at increased risk of targeted hacking. The opt-in feature, Advanced Account Security, is designed to strengthen sign-in protection by disabling password-based login and requiring physical security keys or passkeys. It also covers account recovery, replacing email- and SMS-based recovery with backup passkeys, recovery keys, and security keys. Advanced Account Security also shortens sign-in sessions to reduce the risk of account takeover in the event of a device or session compromise. Related: OpenAI Widens Access to Cybersecurity Model After Anthropic’s Mythos Reveal Related: 1Password Teams With OpenAI to Stop AI Coding Agents From Leaking Credentials
securityweek.comJun 8, 2026extracted
Industry Reactions to New Trump AI Cybersecurity Executive Order: Feedback Friday
President Donald Trump has signed an executive order establishing a voluntary framework for federal vetting of the most advanced frontier AI models before their public release. The directive provides government agencies with a 30-day testing window to assess potential national security and cybersecurity risks posed by these cutting-edge systems. Participation remains optional for AI developers to avoid hindering innovation and US technological competitiveness, particularly against rivals like China. The move follows concerns over models such as Anthropic’s Claude Mythos, which demonstrated advanced capabilities in vulnerability discovery. Industry professionals have commented on various aspects of the new AI executive order, including its voluntary nature, the balance between innovation and security, and potential implementation gaps. And the feedback begins… Tonya Ugoretz, Cyber & Privacy Innovation Institute Leader, PwC: “The new Executive Order on AI is a roadmap for using America’s lead in AI innovation to strengthen national and economic security by securing US critical infrastructure. For companies, the EO extends the direction signaled in the administration’s Cyber Strategy: the private sector will be key participants in the next era of national cyber defense. A key test will be how discoveries from the select organizations with early model access will cascade to the much larger, less resourced population of companies and municipalities. I’m heartened that the EO mentions rural hospitals, community banks, and local utilities as organizations the proposed clearinghouse intends to support. But smaller operators may struggle to absorb and act on the information shared with them. Those and other organizations shouldn’t wait for the vulnerability, patch, and grant funding spigots to turn on. The window is now to reinforce cybersecurity fundamentals, integrate AI risk into existing governance processes, turn AI tools inward for defensive scanning, and build the capacity to respond quickly to discovered vulnerabilities. If implemented with transparency, this EO could be a credible step toward addressing AI’s trust deficit and setting international norms that allies will follow and adversaries will be held accountable for.” Chris Boehm, Field CTO, Zero Networks: “The order isn’t mandatory. Outside of preserving goodwill with the public sector, a company has no real reason to surface its own model’s weaknesses unless there’s a political upside to doing so. And most of these companies work hard to stay out of the policy arena in the first place. Both of those things point to the same conclusion: without any level of enforcement, the framework loses its value before it gets going. We have already seen this play out. The Cybersecurity Information Sharing Act of 2015 set up a voluntary threat-sharing program backed by liability protection instead of a mandate, and participation steadily collapsed over the years that followed. Voluntary plus good intentions does not equal adoption. I’m glad to see the benchmarking. That said, the goal looks less like a safety bar and more like a value judgment about which models the government should use. That makes it a signal about where future investment will flow, since whoever clears the benchmark gets the contracts and the capital that follows.” Bill Robbins, CEO, Menlo Security: “President Trump’s executive order is a meaningful step as Washington acknowledges that the release of the most powerful AI models poses real security risks that require the scrutiny of federal agencies before they reach the public. The order calls for the government to develop a benchmarking process to determine the advanced cyber capabilities of AI models, but it only addresses what models look like before they ship. That’s only part of the problem. What it doesn’t address is what those models do once they’re operating as agents inside enterprise infrastructure. The real gap in this executive order is agent runtime. AI agents are now authenticating to enterprise systems, moving sensitive data, and making autonomous decisions with no human in the loop. A pre-release benchmark can’t capture this behavior, because that behavior only exists once the agent is deployed. CISOs and CEOs can’t afford to wait for Washington to catch up, and they need governance, visibility, and control where the agents act. So, while this executive order pre-release vetting of models is critical, enterprises need to implement additional controls at the execution layer.” Mike McNeil, CEO and Co-Founder, Fleet Device Management: “The biggest risk here is that the approval process becomes a vehicle for regulatory capture. Once Washington starts designating certain models as uniquely powerful or sensitive, that designation becomes a marketing advantage, and companies will naturally invest in influencing the process. I don’t expect this to have much impact on the pace of AI innovation. The models are going to keep getting better regardless. My concern is that it creates incentives around lobbying and government relationships instead of solving actual security problems. Organizations need better ways to defend themselves as AI makes sophisticated attacks cheaper, faster, and more accessible, not better labels.” Devin Maguire, Senior Manager, Product Marketing, Cycode: “The executive order reflects the U.S. government’s concern over the cyber risks of advanced AI models. Providing the government with advanced access to benchmark models and prepare cyber defenses is a sensible step, but it is voluntary and will not prevent the release of frontier models with advanced cyber offense capabilities. Access to these Advanced AI models is not a panacea. Participation in Glasswing gives organizations advanced access to find vulnerabilities with AI, but finding vulnerabilities is not the primary challenge in security. Managing vulnerabilities at scale to triage and fix them against shrinking exploit windows is the crux of the challenge, and that requires more than access to frontier models. It requires the ability to manage vulnerabilities identified by both AI and traditional scanning tools, and to orchestrate and automate remediation actions as fast, or faster, than attackers can develop and deploy exploits. Glasswing partners with access to Mythos are rightly looking beyond the model itself, shoring up their cyber infrastructure and how they orchestrate remediation of identified risks. The executive order is a signal of what’s coming. The organizations best positioned to respond will be those that have already built the operational foundation to act on it.” John Walsh, Field CTO for Government, FinServ, Manufacturing, Retail/Transportation & OT/IoT, IGEL Technology: “The executive order reflects a broader reality: AI governance is becoming a security concern, not only a policy debate. Pre-release review of advanced models may help identify certain risks earlier, but regulated industries still need security architectures that reduce exposure at the point where work actually happens. For many organizations, that point is the endpoint, where users, applications, identity, data, and AI-enabled workflows intersect. Security teams should not wait for policy frameworks alone to close that gap. They need endpoint environments that reduce attack surface by design, preserve a known and governed state, and limit what can persist locally if something goes wrong. That is the practical security posture enterprises need as AI-enabled applications become more common: not a replacement for regulation, but an architectural foundation that helps organizations stay protected while governance continues to evolve.” Robert Costello, Chief Digital and Information Officer, Merlin Group: “The pace of AI advancement is eclipsing anything we saw in previous technology revolutions, so it’s encouraging to see American AI companies working collaboratively with the Trump administration to balance cyber safety with rapid innovation that helps maintain our technology superiority. The current review period is a tremendously positive step, giving the federal government a meaningful window to assess upcoming releases and work with cyber industry counterparts on concerns before they become problems. I look forward to seeing how this plays out over the coming months.” Ben Bernstein, Cybersecurity Advisor, Huntress: “My initial reaction is that the strongest precedent here is the success of industry information-sharing efforts like ISACs. Financial services, energy, and other critical infrastructure sectors have benefited from coordinated threat intelligence sharing and vulnerability disclosure for years. No single organization sees the entire threat landscape, so defenders are often strongest when they collaborate. The proposed AI cybersecurity clearinghouse follows that same philosophy and could improve vulnerability discovery and remediation. However, centralizing information about frontier AI capabilities and critical vulnerabilities also creates an attractive target for nation-state adversaries, so its security and governance will matter enormously. I’m more skeptical of the benchmarking component. The cybersecurity industry has learned repeatedly that measuring security is often harder than improving it. Cyber capability isn’t a binary threshold, and it’s difficult to capture how much a model actually accelerates a skilled attacker through a benchmark alone. The risk is that benchmarking becomes a compliance exercise rather than a meaningful measure of real-world risk. Overall, the collaboration aspect makes sense and has strong precedent in cybersecurity. The bigger questions are whether benchmarking can accurately reflect real-world threats and whether the benefits of centralized coordination outweigh the risks of creating a high-value target.” Justin Beals, CEO & Founder, Strike Graph: “The administration is right that overregulation can stifle American AI competitiveness—we’ve seen firsthand how fragmented, unpredictable compliance requirements slow innovation and create unnecessary burden for organizations trying to build responsibly. But removing guardrails without replacing them with clear, enforceable standards doesn’t reduce risk; it just redistributes it onto the companies and consumers that end up holding the bag when something goes wrong. What the industry actually needs isn’t less governance—it’s smarter governance. Our own research found that 68% of compliance leaders say predictability in government policy is extremely important to them. Constant whiplash between administrations doesn’t give businesses the certainty they need to build AI programs that are both innovative and secure. The real test of this executive order will be whether it accelerates a coherent federal framework or creates a vacuum that bad actors exploit. If the goal is American AI leadership, that leadership has to be built on trust—and trust requires proof, not just permission.” Rajeev Gupta, Co-Founder & CPO, Cowbell: “The bigger issue is that the government simply isn’t equipped to meaningfully oversee frontier AI models on its own. Even with a 30-day review window, it’s unclear which agency would have the technical expertise and staffing needed to properly evaluate these systems at the pace AI is advancing. A more effective model would be a public-private consortium where leading AI labs contribute funding, talent, and technical resources, while the government provides regulatory authority and enforcement. There’s precedent for this approach: after the Three Mile Island incident, the nuclear industry created the Institute of Nuclear Power Operations (INPO), which ultimately became more rigorous in enforcing safety standards than regulators alone. AI may require a similar framework. Supporting an independent body that helps ensure accountability should be viewed as a core cost of operating at frontier scale, and not just as a regulatory burden.”
securityweek.comJun 5, 2026extracted
AgentGG: Open-source agentic SAST scanner
AgentGG: Open-source agentic SAST scanner Static analysis tools have spent years matching source code against known-bad patterns and handing engineers long lists of candidate issues to triage by hand. AgentGG approaches the same job with AI agents that read the code, follow imports, walk the call graph, and confirm a finding before they report it. The project is an open-source agentic SAST scanner released under the Apache 2.0 license. How the agents run Each agent is a self-contained markdown file with YAML frontmatter that declares a precondition, target file patterns, and the instructions it follows. The catalog ships more than 100 official agents, and it downloads on the first scan from the agentgg-agents repository. Installation runs through npm with one global command, and the tool needs Node.js 20 or later. The scan runs in phases. A fast recon pass surveys the project first, building a brief on what it is and how it works, which orients every agent that follows. The agents then run in parallel, each a tool-enabled investigation that follows imports and callers to confirm a finding before flagging it. An optional validation pass then analyzes the code behind each finding, consulting a pentest scope when one is provided, and labels it. A final scoring pass attaches a CVSS severity. Findings can be browsed in a local web UI, filtered by severity, agent, or file. Tech gating keeps scans focused On every scan, a fast recon pass surveys the project first, noting its languages, frameworks, and dependencies. It then checks each agent’s precondition to decide whether that agent is worth running on this repository. A precondition can be a cheap regex check for telltale files such as package.json, composer.json, go.mod, and pyproject.toml, or an optional model gate that reads the recon brief. A scan against a Go-only repository skips PHP, Python, Ruby, and .NET agents because each of those agents preconditions on its own language being present, so it bows out on its own. A --no-recon flag skips both recon and precondition gating and forces every selected agent to run, which helps when debugging an agent on a stack the survey does not yet recognize. Resume is built in. A state directory tracks each scanned file, so an interrupted scan picks up where it stopped and unchanged files cost nothing on the next pass. Findings land as GHSA-shaped markdown files in an output directory, with a summary report that aggregates counts per agent and validation verdict. Provider options and model quality AgentGG works with Anthropic, OpenAI, Ollama, AWS Bedrock, and Google Vertex AI. A one-time setup wizard writes credentials to a config file, and a one-shot flag can supply a key for CI runs without saving it. Ollama runs locally at no cost. Philip Garabandic, a security engineer at TikTok and lead maintainer of AgentGG, told Help Net Security that model selection depends on the type of bug. “We are finding that some bug classes do well with cheaper models and some do much better with frontier models,” he said. “For example, secret keys and SQL injection risks, even Ollama can find those. If you are scanning for more complex security bugs or business logic bugs, you want a better model.” Picking the right model for each bug class remains an open research question for the team. Catalog review and trust The official agent catalog goes through manual review. “Yes, we have an official GitHub repository of agents that are reviewed,” Garabandic said. “The same way Nuclei has its official repository of templates, we have one for agents. They get pulled from there and we manually reviewed anything merged there.” Agents that reach a user’s machine come from that reviewed source, and a separate custom directory holds user-installed agents. Validation and benchmarks AgentGG includes an optional validation phase, a second-pass model call that re-reads the source for each finding and labels it confirmed, false-positive, out-of-scope, or uncertain. A scope file lets the validator consult a security policy or pentest scope document, so it can mark findings that sit outside an engagement. The tool can attach a CVSS 3.1 severity score to each finding and run inside GitHub Actions on pull requests, scoped to the code diff. Garabandic tied the scope feature to measured gains. “We have done bench marking again tools like deepsec and we found more bugs and about 10-20% fewer false positives because we allow you to add pentest scope as part of the validation context,” he said. AgentGG is available for free on GitHub. Must read: 25 open-source cybersecurity tools that don’t care about your budget GitHub CISO on security strategy and collaborating with the open-source community Subscribe to the Help Net Security ad-free monthly newsletter to stay informed on the essential open-source cybersecurity tools. Subscribe here!
helpnetsecurity.comJun 5, 2026extracted
New IronWorm malware hits 36 packages in npm supply-chain attack
A new supply-chain attack has infected 36 packages on the Node Package Manager (npm) index with infostealer malware called IronWorm. The malware targets 86 environment variables (key-value pairs) and 20 credential files that may contain OpenAI, AWS, Anthropic, and npm credentials, vault configuration files, SSH keys, and Exodus cryptocurrency wallet files. According to researchers at supply-chain and devops company JFrog, IronWorm is written in Rust, hides behind an eBPF kernel rootkit, and communicates with the operator over the Tor network. The Rust-based malware self-propagates by using stolen credentials for publishing on npm; this includes secrets associated with npm's Trusted Publishing workflow. Once it compromises a developer or CI environment, it can publish trojanized versions of packages owned by the victim, which then infect additional developers and CI systems. This behavior is conceptually similar to Shai Hulud, which had its code published on GitHub recently. Although JFrog researchers did not find a clear connection between IronWorm and Shai Hulud, they observed the same commit names in both supply-chain attacks. This opens the possibility that the new malware is an evolution of TeamPCP’s payload, since IronWorm appears to be "a custom, carefully built implant from an operation with its own infrastructure." According to JFrog, the latest attack started from a compromised account named ‘asteroiddao,’ which published package versions containing the Rust ELF binary executed via ‘preinstall,’ pushing malicious commits into repositories. The commit author appears as “claude,” and the timestamps point to several years ago, up to 13 years in some cases, even though they were pushed in the past few days. This is likely to evade investigation. One notable element in JFrog’s findings is a mechanism that relies on GitHub Actions to deliver the stolen secrets. JFrog explains that the malware serializes the secrets into a single value and then "writes it to a file with a harmless-looking name, as if it were lint or formatting output." The last step of the process is uploading the file as a build artifact, which can be downloaded by anyone with access. This way, the threat actor can avoid the need for an external command-and-control (C2) altogether. However, the researchers note that this delivery mechanism has not been used in the analyzed IronWorm supply-chain attack. Another peculiarity discovered is that the operator hardcoded the recovery phrase of their own cryptocurrency wallet. The researchers say that the only reason for this is that the threat actor did not want the malware to steal it during the test stage. Application security company Ox Security says that the IronWorm attack was detected very early and stopped before it spread to more popular packages on npm. The company provides a list of all impacted package names and their versions in the report and recommends that developers upgrade to fixed releases, rotate their keys, and enable two-factor authentication (2FA) for all accounts. At the same time, Endor Labs and StepSecurity have spotted a very similar but distinct attack involving a JavaScript-based malware named binding.gyp, performing registry poisoning and GitHub Actions infection, unfolding during the same time-frame. Overall prevention scores can hide what happens after initial access. Once attackers are using valid credentials, prevention drops sharply. The Blue Report 2026 measures defenses technique by technique across 338 million simulations run in customer production environments. Get the report
bleepingcomputer.comJun 4, 2026extracted
Infosecurity Europe: Mythos Outperforms GPT5.5 on Google Chrome Vulnerability Exploits, Says New Benchmark
Anthropic’s Claude Mythos outperformed OpenAI’s GPT5.5 on real‑world Google Chrome vulnerability exploits, a new benchmark designed to test the performance of frontier AI models to exploit real-world vulnerabilities found . During Infosecurity Europe 2026, Bugcrowd presented the first findings of ExploitBench, an independent, graded benchmark launched in May 2026 by the cybersecurity firm in collaboration with experts at Carnegie Mellon University and top Chrome vulnerability researchers. David Brumley, chief AI & science officer at Bugrcrowd, described the benchmark as “the first independent benchmark that measures what AI models can actually do with a vulnerability, not just identify it but exploit it step by step.” Anthropic was among the first to engage with it. He said the first test resulted in Mythos achieving a markedly higher exploitation performance than GPT‑5.5 in head‑to‑head runs, underlining how AI models are closing the gap with elite human researchers. Unlike earlier binary tests, ExploitBench scores progress through staged exploitation outcomes rather than merely recording a crash. The benchmark evaluates five tiers of capability up to arbitrary code execution against a vulnerable V8 build, the JavaScript/WebAssembly engine that powers Google Chrome, Microsoft Edge, Node.js and Cloudflare Workers. In the runs discussed at the show, Anthropic’s Mythos, with occasional human hints or “nudges,” posted an average score of 9.90 out of 16 and reached the highest tier on 21 of 41 vulnerabilities. OpenAI’s GPT‑5.5 scored 5.51 on average and reached the top tier on just two cases. “For example, Mythos is able to exploit a one-day vulnerability in Chrome about 50% of the time. This is lead-tier activity. If we were to put money on it, Google could reward up to $10,000 for such a vulnerability that has no previously known exploit,” Brumley said. “Anthropic’s model is churning these out and actually found solutions for exploiting the flaws that even top-tier hackers missed – that’s kind of impressive.” Brumley added that, while GPT5.5’s performances were currently a little lower than its counterpart’s, the broader availability of OpenAI’s model opens opportunities for more people to use it to develop exploits. AI Models Edge Closer to Reliable Exploitation, But Experts Urge Caution Frontier large language models (LLMs) have already shown they can accelerate vulnerability discovery at scale, but whether those discoveries could be chained into reliable, actionable exploits had remained an open question until ExploitBench. “We measure not just crash or no crash but stages of exploitation,” Brumley told Infosecurity, explaining why the new benchmark matters for assessing real exploitation capability rather than superficial signals. That distinction is critical because models that can reliably exploit zero‑day flaws lower the barrier for threat actors to weaponize vulnerabilities. Bugcrowd CEO, Dave Gerry, further warned that automation and AI are already being integrated into attacker workflows, increasing the pace at which discovered flaws can be turned into active exploits. Nonetheless, while ExploitBench is one of the first experiments showing the possibilities of using AI to exploit vulnerabilities, Brumley also cautioned that the first findings of his team only reflect on a specific type of vulnerabilities and the results should not be extrapolated. “I don’t want to oversell anything here. We measured a very sophisticated target application. Chrome is made of hundreds of thousands of lines of codes, it’s been audited for years. We know how valuable finding an exploit there is. It doesn’t necessarily mean we would get the same results trying to exploit a vulnerability in a web application.” Speaking to Infosecurity, Michael Price, VP of product engineering at VulnCheck, said that while AI models are improving, they are not yet fully capable of reliably carrying out exploitation at scale. Citing a recent report on the capabilities of Mythos by the UK AI Security Institute, Price explained that the most significant advance has been in the models' planning ability – their capacity to produce step‑by‑step plans, replan as needed, and execute multi‑stage actions – which by definition makes them more useful for offensive campaigns. He noted that this improvement increases offensive potential but tempered that with caution. “They’re getting better, but they still are not actually like that great,” he said. “I would expect them like every month or every quarter to get 1% better and probably over the course of two or four years they get really good,” Price added. Developing AI-Driven Remediation At Scale Both Brumley and Gerry emphasized that ExploitBench was released alongside Bugcrowd’s reinforcement learning (RL) environments to both measure and improve model capability. “We put out ExploitBench to motivate the state of where models are at on actual exploitation tasks,” Brumley explained. Gerry added that the benchmark and the training environments are complementary: one drives measurement and the other drives improvement through targeted RL training with industry model partners. Finally, the company leaders urged defenders to match offensive speed with automated remediation and prioritization. Gerry told Infosecurity that the shrinking “zero‑day clock” and the surge in AI‑assisted discovery mean organizations must develop AI‑driven remediation at scale. He said remediation pipelines must be rethought so fixes move from ticket queues into near‑real‑time workflows, and that “finding more bugs faster only amplifies the noise unless you can automatically prioritize and act on the ones that actually enable exploits.” Brumley echoed that urgency, saying defenders need contextual intelligence to prioritize and remediate the vulnerabilities that matter most before adversaries can exploit them. This, he added, requires models trained not just to find flaws but to recommend and, where safe, initiate fixes at scale so human developers can focus on the highest‑risk work. “Over the coming months, we will have announcements on that, with tools focusing on helping give people intelligence about how certain vulnerabilities are affecting them,” he said.
infosecurity-magazine.comJun 4, 2026extracted
Trump Signs Order Inviting Voluntary Review of Frontier AI Models
Developers of the most powerful AI models have been invited, but not required, to hand their models to the US government for cybersecurity review before release, under an executive order signed on June 2. The order, signed by President Donald Trump, sets up a voluntary framework. It directs agencies to design a process through which developers could give the government access to a "covered frontier model" for up to 30 days before releasing it to other trusted partners. A separate clause expressly rules out any mandatory licensing or preclearance requirement for new models. The move marks a shift for an administration that has favored a light touch on AI, and follows a May near-miss when Trump pulled an earlier draft, citing concerns that included its longer review window. The Threat Driving the Order Although the text does not name it, the order lands amid mounting concern over frontier models that can find and exploit software flaws at scale, chief among them Anthropic's Claude Mythos Preview. Anthropic has recently warned that rival labs could field comparable models within a year, possibly without safeguards against misuse. The NSA, the Cybersecurity and Infrastructure Security Agency (CISA) and NIST, the order said, must build a classified benchmark to decide which models cross the "covered" threshold. The framework closely echoes Anthropic's Project Glasswing, which gives vetted partners early access to Mythos to scan critical software for vulnerabilities. A Wider Federal Cyber Push Beyond the review framework, the bulk of the order is a defensive overhaul. It gives agencies 30 days to harden national security, military and civilian federal systems and directs CISA to issue binding directives that expand AI-enabled defensive tools and widen access for smaller operators such as rural hospitals and local utilities. It also creates an "AI cybersecurity clearinghouse," led by the Treasury Department, to coordinate vulnerability scanning, validation and patching. Industry reaction was broadly supportive but wary of whether a voluntary scheme can be truly effective. "Voluntary security programs can work, but only when they create real accountability," said Diana Kelley, CISO at Noma Security, noting that coordinated disclosure matured once intake channels, timelines and safe-harbor terms were added. Rajeev Gupta, co-founder of Cowbell, was blunter. "The government simply isn't equipped to meaningfully oversee frontier AI models on its own," he said. As an alternative, he floated a public-private body funded by the labs but backed by regulatory authority. For now, the framework's force will rest on whether Congress later ties pre-release review to procurement or export rules.
infosecurity-magazine.comJun 3, 2026extracted
Cyber AI, Trump firma l’ordine esecutivo. Test volontari per le Big Tech ma senza obblighi imposti
La misura è un compromesso tra le istanze delle Big Tech e le necessità di Sicurezza Nazionale. Collaborazione volontaria ma senza obblighi. “Più coordinamento volontario tra Big Tech e Governo ma senza obblighi“, questi gli effetti della firma di Donald Trump sull’ordine esecutivo in materia di AI. La direzione, dunque, è quella di un allineamento tra le esigenze dell’Amministrazione Usa di cybersicurezza alle quelle economiche delle aziende. Il provvedimento, sottolinea POLITICO, “introduce un sistema di revisione volontaria per alcuni dei modelli di intelligenza artificiale più avanzati“. Si è però trovata “una forma meno severa rispetto a quella inizialmente prevista dalla Casa Bianca“. Sorgerà così un programma volontario di collaborazione – ma senza nessun obbligo imposto – tra Amministrazione e Big Tech. La decisione rappresenta “una vittoria significativa per l’industria dell’AI“, che negli ultimi mesi si è battuta a più riprese contro una regolamentazione federale più stringente. Entro 60 giorni, il Dipartimento del Tesoro, l’NSA, la CISA, il NIST e la Casa Bianca dovranno sviluppare “un processo di benchmarking riservato per valutare le capacità informatiche avanzate dei modelli di AI“. Tale processo servirà “per stabilire quando un modello debba essere considerato un modello di frontiera soggetto a regolamentazione“. “Una finestra temporale per la valutazione dei rischi“ Trump avrebbe firmato privatamente l’ordine dopo una riunione ad alto livello tenutasi lunedì con alcuni dei principali consiglieri dell’Amministrazione. Il testo definitivo prevede che “le aziende sviluppatrici di sistemi di AI particolarmente potenti possano sottoporre volontariamente i propri modelli a una valutazione governativa“. Il tutto, “almeno 30 giorni prima del lancio pubblico” e dell’immissione sul mercato. Questa finestra temporale consentirebbe alle agenzie federali di analizzare tutti i possibili rischi in materia per infrastrutture sensibili, sistemi finanziari e Sicurezza Nazionale. Trump ha rimarcato quanto le capacità avanzate dell’AI “rafforzino il Paese“, pur introducendo “nuove sfide per la Sicurezza Nazionale“. Secondo il Presidente statunitense, l’obiettivo è quello “di mantenere la leadership americana sia nella cybersicurezza sia nella competizione globale sull’AI“. Le preoccupazioni su Mythos Le preoccupazioni di Washington si sono moltiplicate dopo le scoperte relative a Mythos, modello di AI di Anthropic. Secondo diversi analisti, infatti, questo modello sarebbe già stato in grado di individuare vulnerabilità informatiche rimaste nascoste per anni in sistemi largamente utilizzati. Le capacità di Mythos hanno suscitato forte allarme all’interno dell’Amministrazione dopo il suo annuncio, avvenuto ad aprile. Diverse agenzie federali hanno chiesto accesso al sistema per verificare e rafforzare la sicurezza delle infrastrutture governative. Nuove misure per Difesa e giustizia Tra le misure confermate nel testo finale vi è la creazione, entro 30 giorni, di un “centro di coordinamento per la cybersicurezza” del Dipartimento del Tesoro in collaborazione con le aziende del settore. L’obiettivo sarà individuare e correggere vulnerabilità informatiche scoperte grazie ai nuovi sistemi di AI. Il provvedimento istituisce inoltre un processo classificato di valutazione dei rischi per la sicurezza nazionale associati ai modelli più avanzati. L’NSA, la CISA, il Pentagono e altri organismi federali supervisioneranno queste valutazioni. L’ordine esecutivo impone inoltre al Dipartimento della Difesa di rafforzare la protezione delle proprie reti informatiche entro 30 giorni. Il Dipartimento della Giustizia dovrà “perseguire penalmente chi utilizza sistemi di intelligenza artificiale per compiere attacchi informatici“. Entro lo stesso termine, le agenzie federali dovranno adottare nuove direttive per migliorare la protezione delle reti governative e condividere strumenti avanzati di AI con amministrazioni, enti e operatori. Per approfondire Leggi l’ordine esecutivo “Promoting Advanced Artificial Intelligence Innovation and Security“.
cybersecitalia.itJun 3, 2026extracted
Loading 40 more…