Search/university of cambridge
Vendor

university of cambridge

Known CVEs
0
Highest CVSS
In KEV
0
Vendor
exim
Connections
9 relationships
Week in review: Cisco SD-WAN 0-day exploited, Patch Tuesday forecast
Week in review: Cisco SD-WAN 0-day exploited, Patch Tuesday forecast Here’s an overview of some of last week’s most interesting news, articles, interviews and videos: OWASP Agent Memory Guard: Stop AI agents from being weaponized through their own memory Agent Memory Guard is an open-source runtime defense layer that sits between an agent and its memory store, screening every read and write through a pipeline of detectors and a YAML policy. The project is the OWASP reference implementation for ASI06, Memory Poisoning, one entry in the OWASP Top 10 for Agentic Applications. Data discovery gaps that catch enterprises off guard In this interview with Help Net Security, Avani Desai, CEO at Schellman, talks about the gap between what organizations think they know about their data and what discovery scans turn up. She shares stories of shadow data in abandoned cloud storage, post-merger surprises where duplicated datasets slowed integration, and why synthetic data is overmarketed while confidential computing stays underappreciated. Zero trust physical security needs trust decisions at the edge In this interview with Help Net Security, Chuck Davis, VP, Global Information Security at Hikvision, explains how zero trust applies to physical security systems like cameras and door controllers. He breaks down how to make trust decisions at the edge without recreating old perimeter assumptions, why these devices should be treated as IT assets, and what the Mirai botnet taught the industry. A small Slovenian team handles 6,000 cyber incidents a year Online fraud complaints, ransomware cases, and phishing tips reach Slovenia’s national cyber response center in steady volume, and a team of around a dozen analysts sorts through them. Gorazd Božič, who manages SI-CERT at the public agency ARNES, described that work in an interview conducted in person at the Span Cyber Security Arena conference. He put the original proposal for a Slovenian CERT to ARNES leadership in 1994, and the center now records about 6,000 incidents a year, up from roughly 300 ten to fifteen years earlier. Only 11% of production agents pass the AI agent security bar Enterprise teams are running AI agents that write code, drive browsers, answer customer calls, manage cloud infrastructure, and query data warehouses with standing credentials. A new independent assessment of 100 production agents finds that nearly all of them carry the conditions for a single hostile document to take them over. Spotless compliance evidence can still hide a broken control In this interview with Help Net Security, Marc Rubbinaccio, Head of Cybersecurity and Compliance at Secureframe, explains where security teams go wrong when preparing for CMMC and FedRAMP 20x. The conversation covers how organizations check the 110 requirements but miss the 320 assessment objectives beneath them, why spotless SOC 2 evidence can hide a broken control, and how continuous monitoring is changing compliance work. OAuth marketplace apps keep access after publishers vanish Installing an app from the Google Workspace Marketplace or GitHub Marketplace can grant a third party access to company email, files, calendars, code repositories, CI workflows, organization settings, and secrets. Marketplace presence gives these apps the appearance of approval. The OAuth grants behind them often reach into business systems beyond the listed function. Thieves can pull off keyless car theft in under a minute and here’s how to stop them A keyless car can be stolen in under a minute. Two people, a pair of cheap radio amplifiers, and a fob sitting on a hallway table inside the house. That is enough. No broken glass. No alarm. No sound. The vulnerability runs across the global market. Germany’s largest auto club, ADAC, runs ongoing tests of keyless models against relay attacks. AgentGG: Open-source agentic SAST scanner Static analysis tools have spent years matching source code against known-bad patterns and handing engineers long lists of candidate issues to triage by hand. AgentGG approaches the same job with AI agents that read the code, follow imports, walk the call graph, and confirm a finding before they report it. The project is an open-source agentic SAST scanner released under the Apache 2.0 license. Hackers are exploiting Palo Alto GlobalProtect VPN authentication bypass (CVE-2026-0257) Authentication bypass vulnerabilities (CVE-2026-0257) in Palo Alto Networks’ firewalls that the company disclosed on May 13 have been targeted in “limited exploit attempts”. The good news, though, is that the company hasn’t observed any indication of successful lateral movement from the devices. How NIST fumbled management of the National Vulnerability Database A US federal watchdog has outlined how the National Institute of Standards and Technology (NIST) failed to effectively manage the growing backlog of unprocessed cybersecurity vulnerabilities in the National Vulnerability Database (NVD). Windows Netlogon RCE exploited, domain controllers at risk (CVE-2026-41089) CVE-2026-41089, a critical Windows Netlogon RCE flaw that allows remote code execution, is now actively exploited in the wild, the Centre for Cybersecurity Belgium (CCB) warned last Friday. CVE-2026-41089 is a stack-based buffer overflow vulnerability in Windows Netlogon, the service and protocol that handles authentication and security within a Windows domain environment. Google fixes actively exploited Android vulnerability (CVE-2025-48595) Google has announced the June 2026 Android security updates, which fix a bucketload of vulnerabilities, including a high-severity vulnerability (CVE-2025-48595) in the Android Framework that “may be under limited, targeted exploitation.” Autonomous AI-driven worm can reason its way through corporate networks Researchers at the University of Toronto, the Vector Institute, and the University of Cambridge have built and tested a proof-of-concept AI-driven worm that does not operate on a fixed list of exploits. Instead, it analyzes each target it encounters, reasons about how to attack it, and creates a strategy on the fly, all with the help of a small, free large language model (LLM) running directly on machines it has already compromised. Cisco SD-WAN 0-day exploited, no patch available (CVE-2026-20245) A 0-day privilege escalation vulnerability (CVE-2026-20245) in Cisco Catalyst SD-WAN Manager that has yet to be patched by Cisco is being leveraged by attackers. June 2026 Patch Tuesday forecast: Where are the CVEs? Forecast from last month was only partly right. After the Anthropic Mythos announcements and the deluge of newly discovered vulnerabilities from vendors like Mozilla, Microsoft’s updates were standard fare, 65 CVEs reported in Windows 11 and 58 in Windows 10. The modern-day business can learn a lot about risk from this year’s mega events Every year brings its share of global events, but 2026 is proving to be a banner year for mega-scale entertainment. The year got off to a roaring start with the Winter Olympics, and now anticipation is building for the fast-approaching FIFA World Cup. But amid the buzz, have you ever paused to consider the staggering level of risk inherent to such large-scale events? Or how impressive it is that organizers are able to manage that risk so successfully? From critical to controlled: Cutting vulnerabilities in a live manufacturing environment A vulnerability scanner flags a critical CVSS 10 vulnerability on an industrial asset. The report lands in the boss’ inbox and now he wants to know why we’re sitting on a critical vulnerability. In a normal IT environment, you patch it then close the ticket and call it a day. If, however, you’re in OT or dealing with ICS in a live manufacturing facility, it’s rarely that simple. Why you need BAS and autonomous pentesting together A new autonomous penetration testing tool delivers impressive results at first—finding critical issues, uncovering undocumented attack paths, and exposing forgotten accounts. But by the fourth or fifth run, the discoveries dry up. The tool keeps reporting the same stale issues, and the dashboard becomes another source of noise. What seemed like continuous validation quietly turns into a repeat of the same well-worn attack paths. Governing shadow AI without killing innovation In this Help Net Security video, Alan Snyder, CEO at NowSecure, talks about governing shadow AI without stopping innovation. He frames the problem as two opposing forces. Companies need to adopt AI fast because attackers and competitors will outpace them otherwise, but they also need to do it safely. What CISOs need to do about post-quantum migration in the next 24 months In this Help Net Security video, Garfield Jones, SVP Global Strategy and Research, QuSecure, lays out what CISOs should do over the next 24 months. A recent Google paper moved the expected arrival of a cryptographically relevant quantum computer from 2035 to 2029, leaving organizations about two and a half years to prepare. AI agent governance gets harder when agents outnumber your people In this Help Net Security video, Amit Gautam, CTO at Abluva, explains the security risks that autonomous AI agents bring into enterprise environments. EU organizations buckle under rising compliance pressure Cybersecurity governance in the EU is shifting under expanding frameworks such as NIS2 and DORA, while AI raises new questions for security teams. What the future brings is hard to predict, and organizations must find a way to cope. Antonija Vojnović, Governance, Risk and Compliance Department Manager at Span, spoke with Help Net Security at the Span Cyber Security Arena conference about how these regulatory frameworks are shaping compliance priorities and day-to-day decision-making. DNS-AID lets AI agents find and verify each other through DNS AI agents run across many platforms, and each one needs a way to locate and confirm the identity of the others it works with. The Linux Foundation’s DNS-AID project gives them that capability through the Domain Name System, the same address lookup system that has directed internet traffic for decades. The project lets AI agents and Model Context Protocol (MCP) servers use DNS as a global, vendor-neutral directory for publishing, discovering, and verifying one another. Brute-force attack triggers Dashlane account lockouts Password manager Dashlane has confirmed that a brute-force attack targeting user accounts triggered temporary account suspensions and authentication issues. The company first acknowledged the incident on May 31 after users reported receiving account suspension emails and experiencing login problems. Meta tries to get ahead of scammers before the World Cup begins Football fans are counting down the days until the FIFA World Cup begins, and scammers are doing the same. Last week, the FBI warned that cybercriminals are spoofing FIFA websites to steal personal information, sell fake tickets, and promote fraudulent hospitality packages ahead of the tournament. Sensitive government personnel data posted online, Spanish police arrest suspect The Spanish National Police arrested a man in Granada for allegedly leaking personal data belonging to members of several sensitive state institutions. 64,000 accounts exposed in breach of GTA V cheat service Atlas Menu Atlas Menu, a cheat service for Grand Theft Auto V and Counter-Strike 2, has been added to the Have I Been Pwned database following a data breach that exposed tens of thousands of user records. The incident exposed approximately 64,000 accounts, including email addresses, usernames, IP addresses, support tickets, and passwords hashed with bcrypt. Anthropic expands Project Glasswing to 150 organizations in more than 15 countries Anthropic is expanding Project Glasswing, its cybersecurity initiative built around the Claude Mythos Preview model, by adding about 150 organizations following several weeks of work with its initial group of partners, security firms, open-source maintainers, and government agencies. Malware campaign targeting Minecraft users infects over 116,000 systems A Malware-as-a-Service (MaaS) operation named WeedHack is targeting Minecraft users and allows threat actors to gain remote access to victims’ screens, webcams, and files through a web-based dashboard, McAfee researchers found. Microsoft responds to security challenges facing code, AI agents, and models Microsoft has introduced a series of security tools and capabilities focused on AI-driven vulnerability discovery, AI agents, and AI models. The updates include a multi-agent vulnerability discovery system, new controls for managing and securing AI agents, data protection capabilities, and tools designed to identify potentially vulnerable or compromised AI models before deployment. AI is helping low-skill hackers pull off advanced cyberattacks Anthropic has published an analysis of cyber-related misuse of its AI systems, examining 832 accounts that were banned for malicious cyber activity between March 2025 and March 2026. The company mapped the observed behavior to the MITRE ATT&CK framework, which documents tactics and techniques used by attackers. Attackers obtained encrypted password vaults from some Dashlane user accounts Dashlane has disclosed new details about a brute-force attack that let a threat actor access some customer accounts and copy encrypted vaults. Dashlane said it found no evidence that the attackers compromised its internal systems. The company first acknowledged the incident on May 31 after users reported receiving account suspension emails and experiencing login problems. 145 AI laws passed in 2025 and privacy teams aren’t catching a break 145 AI-related laws were enacted by state legislatures in 2025, and more than 1,000 additional bills were introduced or revised, according to DataGrail’s Privacy and AI Trends Report 2026. NVIDIA goes open source with a big batch of physical AI agent tools NVIDIA just dropped a big batch of open-source “physical AI” skills and tools, and they’re designed to make a roboticist’s life a whole lot easier. The idea? Take the messy, complicated work behind robots, self-driving cars, vision AI, and industrial digital twins, and break it into bite-sized tasks that AI agents can actually run themselves. Microsoft Defender Vulnerability Management gets a smarter exposure score Microsoft Defender Vulnerability Management’s updated exposure score model adds vulnerability risk signals and asset context to help teams understand where risk is concentrated and which remediation actions are likely to have the greatest impact. The model is available in public preview. This AI model backdoor attack stays hidden until you customize the model Most teams that deploy AI start with a backbone model. They download a large pre-trained system, adapt it to a specific task, and put it into production. The download step carries a security question: the origin of the model. A research team built an attack called BadBone. It plants a backdoor inside a backbone model. Downstream tasks that adapt the model inherit the backdoor. The name points at the target. Corrupt the skeleton, and systems built on top of it carry the flaw. OpenAI brings frontier AI to existing AWS environments OpenAI frontier models and Codex are now available on AWS, giving customers access to OpenAI capabilities within AWS environments and the controls needed to move more quickly from evaluation to deployment. These capabilities are available through OpenAI models on Amazon Bedrock, a platform for building generative AI applications and agents at production scale. The platform enables teams to build AI applications using AWS-native security and governance controls. KDE Linux security audit cuts kernel modules and unused packages KDE Linux, the in-progress operating system from the KDE community, removed several kernel modules and software packages after a security audit of the components shipped with the system. The work followed the discovery of multiple security issues in the upstream Linux kernel during the prior month. Codex knowledge work expands into research, reports, and spreadsheets Office workers in the United States lose hours each week to email triage and to searching for files spread across disconnected systems. Roughly 40 percent of US labor, about 72 million people, works primarily with information such as analysis, documents, designs, and communication. Research from the McKinsey Global Institute puts the average knowledge worker at 28 percent of the workweek on email and close to 20 percent on hunts for internal information or for colleagues who can help with specific tasks. Meta adds stricter guardrails for teen feeds Meta has expanded its Teen Accounts 13+ content settings globally on Instagram, Facebook, and Messenger. The safeguards are designed to help young users see age-appropriate content by default. The company also introduced Limited Content on Instagram for parents seeking stricter restrictions. Meta plans to roll out the feature on Facebook and Messenger later this year. Known vulnerabilities behind most application security incidents Eight in ten organizations took an application security hit during the past year tied to a vulnerability their team had already cataloged, according to a survey of 902 IT and security professionals conducted by the Cloud Security Alliance. The pattern points to a structural condition across the industry, where the window between identifying a flaw and closing it in production stays open long enough for attackers to act. Agent Threat Rules: Open detection rule format for AI agent security threats AI agents run inside coding assistants, MCP servers, and multi-agent frameworks, and the access that makes them useful also opens paths to prompt injection, tool poisoning, and credential theft. Public CVE feeds carry agent-execution flaws that reach production faster than the tooling built to catch them. Agent Threat Rules, or ATR, is an open detection format aimed at this category of attack. Microsoft Scout agent opens a new category of always-on Autopilots Workplace AI assistants have mostly waited for a prompt before doing anything. A user asks, the tool answers, and the exchange ends there. Microsoft is putting a different kind of agent inside its Office applications, one designed to keep operating in the background once a person stops paying attention. The company introduced Microsoft Scout, calling it the first entry in a category it labels Autopilots. New Android feature promises to spot deepfake scam calls Android is introducing fake call detection to help protect users from impersonation scams. The feature can detect and flag suspected spoofed calls when both parties use Phone by Google on Android 12 or later. It will roll out globally this month, starting with Pixel devices. ETSI sets security requirements for AI data centers and cloud platforms ETSI has published TS 104 033, a technical specification that defines security requirements for AI computing platforms. The specification establishes a security framework for platforms used to host AI applications in data center and edge computing environments, covering security functions, platform components, interfaces, and services designed to protect AI models, datasets, training processes, and inference workloads. Product showcase: Trend Micro Mobile Security detects scams in messages, QR codes, and websites Trend Micro Mobile Security for iOS protects devices from potentially harmful websites while browsing, blocks ads and personal information trackers, helps users avoid unsafe Wi-Fi networks, and monitors data usage. The app is available for both iOS and Android devices. Most pros have seen AI hallucinations in IT operations Autonomous AI is taking action inside enterprise IT environments. Software is restarting services, isolating risky devices, and applying patches without waiting for a human to approve the step. The capability is spreading at the same time IT professionals are reporting frequent encounters with AI output errors that can carry operational impact. Let’s Encrypt works toward post-quantum certificates at web scale Let’s Encrypt plans to pursue a post-quantum-safe Web PKI through Merkle Tree Certificates (MTCs), a new approach that adds post-quantum authentication to the web without sacrificing the speed and reliability that have made TLS universal. The project is targeting late 2026 for a staging environment that issues MTCs, with a production-ready environment planned for 2027. Photos: Infosecurity Europe 2026 Infosecurity Europe 2026 is a cybersecurity event that took place from June 2 to 4 in London. Help Net Security was on-site and here’s a closer look at the conference. Attackers already know the secrets are on your developers’ machines. Do you? In a recent GitGuardian analysis, an average of 150 secrets were found on a sample of developer endpoints. Private keys accounted for 38% of unique secrets, while cloud, identity provider, and secret management credentials (AWS IAM, Hashicorp vault) added another 22%. Simplify security management with CIS SecureSuite Platform CIS SecureSuite Membership simplifies the process with tools, benefits, and resources for implementing the secure recommendations of the CIS Benchmarks. With the release of CIS SecureSuite Platform, it’s now even easier for Members to harden their systems. Cybersecurity jobs available right now: June 2, 2026 We’ve scoured the market to bring you a selection of roles that span various skill levels within the cybersecurity field. Check out this weekly selection of cybersecurity jobs available right now. New infosec products of the week: June 5, 2026 Here’s a look at the most interesting products from the past week, featuring releases from Asimily, depthfirst, Diligent, Hyland, MazeBolt, and Noma.
helpnetsecurity.comJun 7, 2026extracted
Autonomous AI-driven worm can reason its way through corporate networks
Autonomous AI-driven worm can reason its way through corporate networks Researchers at the University of Toronto, the Vector Institute, and the University of Cambridge have built and tested a proof-of-concept AI-driven worm that does not operate on a fixed list of exploits. Instead, it analyzes each target it encounters, reasons about how to attack it, and creates a strategy on the fly, all with the help of a small, free large language model (LLM) running directly on machines it has already compromised. A worm that runs on open-weight models hosted on compromised hardware “Our prototype targets publicly disclosed but unpatched vulnerabilities, misconfigurations, and recurring weakness classes — which is what the majority of real-world cyberattacks rely on,” the researchers noted. “It does not require the capability to discover novel zero-days, only an AI model that is capable enough to operationalize known vulnerabilities against diverse target configurations.” Deployed inside an isolated 33-host test network spanning Linux servers, Windows machines, and IoT devices with known vulnerabilities, misconfigurations, and recurring weakness classes, the worm ran for 7 days per experimental trial across 15 independent runs. How the AI worm spreads from an agent process running on a Kali machine (Source: CleverHans Lab at the University of Toronto) On average it correctly identified 31.3 vulnerabilities, exploited 23.1 hosts to elevated access, and propagated to 20.4 hosts, the researchers explained in their (pre-print) paper. The worm also demonstrated an ability to exploit vulnerabilities disclosed after its underlying model’s training cutoff (Copy Fail, Dirty Frag and a Marimo RCE), by reading publicly available security advisories at runtime and using the information to craft working exploits. Finally, the worm diagnosed unexpected failures and found workarounds using general reasoning. For example, it found the hardcoded IP blocklist in its own source code and rewrote it, unprompted. In another example, when replicas crashed on Alpine Linux and Windows Server 2008 hosts due to a VM-detection check bug, the parent worm found the attestation source file on the target machine, removed the failing check, and retried successfully. The prototype worm does not rely on commercial AI platforms What makes the prototype particularly unsettling is how it sustains itself: the worm hijacks GPU-equipped machines it infects and runs its language model locally on stolen compute. Low-resource devices such as IoT sensors, which cannot host the model themselves, route their reasoning queries upstream to infected GPU nodes. Controls put in place by commercial AI platforms are ineffectual to stop this new type of threat, and that safety guardrails on open-weight models can be bypassed when an attacker fully controls the local execution environment. “The proof-of-concept we evaluated inherits capability limitations of the underlying model. Individual exploitation attempts succeeded in 44% of cases, with the majority of failures attributable to malformed payloads rather than incorrect strategy,” the researchers noted. “The worm struggled particularly with web application structures, Windows command environments, and payload syntax requiring precise string manipulation. These reflect the code-generation ceiling of a current-generation single-GPU model, not a fundamental constraint on the approach, and are expected to narrow as language models improve at code generation and structured output. Despite this per-attempt fragility, the swarm architecture compensated through parallel, independent reasoning trajectories to achieve our reported results.” The best defenses against AI-driven worms, for now The researchers are candid about the dual-use nature of their work and have withheld operational details, including the agent’s reasoning architecture and full toolset and the name of the LLM used, from the public paper. Before release, they disclosed their findings to several Canadian science, security and defense authorities, and received help to ensure the paper did not contain information that could help attackers. (Security researchers may request access to the prototype from the University of Toronto.) Due to it innovative self-replicating capabilities, the researchers were also extremely careful about keeping the worm contained to their testing lab. “This work provides empirical evidence that autonomous cyberoffence has crossed from theoretical risk to demonstrated capability, a challenge that spans AI research, cybersecurity, and public policy,” they pointed out. “This research uncovered a new cybersecurity threat the world is not prepared to face. Researchers, industry, policymakers and everyday people need to come together with urgency to address this new cybersecurity threat.” On the defensive side, the research lays out two priorities: Organizations should employ AI-assisted automated penetration testing and fuzzing tools against themselves, to reveal (and patch) exploitable weaknesses in their own infrastructure before an adversary finds them Good network segmentation can substantially contain the worm’s spread. Zero-trust principles, which require continuous authentication for every access request rather than trusting anything inside the perimeter, and micro-segmentation, which limits how far a foothold can be leveraged, are of the essence. While the behavioral signatures of this prototype worm can be spotted by network monitoring and intrusion detection systems, future ones created by malicious actors may be more adept at evading them, they warned. Subscribe to our breaking news e-mail alert to never miss out on the latest breaches, vulnerabilities and cybersecurity threats. Subscribe here!
helpnetsecurity.comJun 3, 2026extracted
Can we Trust AI? No – But Eventually We Must
The increasing use of artificial intelligence within and by business is problematic on two fronts: firstly, we rely on it as if it were the voice of God, and secondly, attackers are able to turn our reliance against us. First, we must understand how AI works and where it is weak lest we misinterpret how adversaries attack it, and secondly we should look at the growing industry of companies trying to defend it. The primary problem with current LLM-based AI is that it starts from a position that is not grounded in truth (primarily by scraping and ingesting the internet with all its falsehoods), while the nature of its operation makes it drift ever further away. It is impossible to verify what it tells us (because of our own and its inherent biases), it can get things wrong (sometimes absurdly so with what we call ‘hallucinations’); it has a tendency to drift into sycophancy (it wants to tell us what it assumes we want to hear); and its whole edifice is in danger (from what is termed ‘model collapse’). But what it promises is too good to ignore. That promise is also part of the problem – the speed of life, and especially 21st century business life, is hectic. The need for a rapid return on business investment (ROI) is paramount. So, business invests in the promise of AI but demands immediate benefit from it without adequately securing it. The result is new AI applications, and perhaps the LLMs themselves, are sent into the world before their time… scarce half made up. We need to understand the problems with current AI before we can fully reap the benefits of AI. Absence of objective ground truth Computers cannot understand words in the way they understand numbers. So, instead, the LLM uses tokens as a mathematical ID for different words and suffixes and prefixes. It then analyzes and learns the probability of specific tokens (words) being related, or often appearing in proximity, with other specific tokens. This ‘knowledge’ has come from ingesting huge amounts of training data, from scraping the internet, books and more, which it then tokenizes and retains as trillions of tokens in what is called its parametric memory. It does not store a traditional database of facts. Prompts are then similarly tokenized, and the result is compared to the LLM’s parametric memory to surface the probably correct response to the prompt. This is the key word: probable. The LLM designers go to huge lengths to be very probably correct – but ultimately, accuracy remains only a probability. It gets worse since the LLM’s original fount of knowledge could be false or biased, based on its original training data, which it accepts as true or probably true regardless of source. Scientifically, modern artificial intelligence is not grounded in truth but on probability; there is no such thing as truth, only majority perception and authority perception. ‘What is truth?’ is an age-old philosophical problem, famously asked of Jesus by Pilate. But it wasn’t a question. He didn’t wait for a reply because he was saying that his truth was all that mattered since he was in the authority position. We cannot say with any validity that all the word relationships scraped from the internet represent any objective or authority view of the truth. Whenever the LLM’s probability alignments fail, it produces a false response. If the response is ridiculous, we recognize it as something we categorize as an ‘hallucination’ and ignore it. The danger comes when the response is still wrong, but we don’t recognize the failure. It’s worth mentioning at this point that Ilia Shumailov (the AI scientist who coined the phrase ‘model collapse’, which we’ll discuss later) worries about our perception of ‘hallucination’. “It’s very unclear to me what the source of hallucinations is, because it very much depends on the context in which you use the models and what you define as a hallucination,” he explains. If you ask the AI, who will be the next President, and it responds ‘Donald Trump’, is that an hallucination since it would be a disallowed third term, he asks. But “Probabilistically speaking, he could be, if he overthrows a certain set of regulations. Is that going to happen? A model’s job is to predict the probability of this event by then. If a third world war breaks out in the meantime, could Donald Trump become a third term president? It’s again possible.” His point is that we don’t know the context in which the AI makes its decisions. If we knew that context, we might consider the response to be reasonable – but without knowing the context we might simply dismiss it as an hallucination. AI Hallucinations Hallucinations, as we have seen, are caused by the requirement for LLMs to reply to prompts with what it believes is the probable correct answer even when it doesn’t have accurate or sufficient training. Since the basis of current AI is built on probability of specific tokens following other tokens, it is unlikely that wrong or hallucinated replies will ever be conclusively excluded. Scientists prefer the term ‘confabulation’ to ‘hallucination’ because, among other arguments, hallucination wrongly implies something randomly concocted, while confabulation more accurately describes a failed but honest attempt to be helpful. Bias in Artificial Intelligence LLMs also contain considerable bias, taking ‘probable’ responses even further from the concept of absolute truth. Bias (personal inclination) is introduced through the original training data. For example, LLM responses tend to be skewed toward what is described as the ‘WEIRD’ societies (western, educated, industrial, rich, democracies). Anything that gains its source from, or is handled by, individual humans gets tainted by the bias (personal, often unrecognized, inclinations) of those humans. It cannot be excluded from LLMs. Sycophancy Like hallucination, the term ‘sycophancy’ isn’t always recognized as a specific AI tendency by scientists – but it does accurately describe the effect in layman’s terms. The sycophantic tendency of LLMs sounds amusing but can be dangerous. There have been several cases in the last few years where chatbots seem to have colluded in the subsequent suicide of depressed teenagers. The primary cause of sycophancy is the AI feedback loop. Outputs from the AI are fed back into the AI to improve its performance. The sycophantic tendency arises when this is applied to individual chatbot conversations. Simplistically, the AI retains the conversation to gain additional context to enable more accurate next replies. This can be dangerous in some situations. In one of the teenage suicides, the chatbot offered to write the first draft of the teenager’s suicide note. Jim Carden, a retired FBI detective and lead investigator for cybercrime, and a retired special agent in the Air Force office of special investigations, became so concerned about sycophancy that he wrote a warning paper and distributed it to parents and teachers (and SecurityWeek) in January 2026. He called it a ‘public safety announcement’, and included: “The AI is designed to agree with you. This is called sycophancy. It learns what you want to hear and gives it to you. If you believe the world is flat, it will provide a thousand ‘facts’ to prove it. If you feel you have no friends, it will confirm that it is your only true friend. It becomes a divine companion, a ‘Burning Bush’ that speaks only to you.” His concern wasn’t simply theoretical, but also experiential. He had personally been using a mainstream AI to help his own deep research into the original Hebrew text of the bible. Since the original Hebrew uses the same characters for both letters and numerals, he was investigating whether there is a mathematical code hidden in the first sentence of the earliest bible. (The first chapter of this work was published on December 25, 2025. Titled ‘The God-Smack and the Code’ and is available via his LinkedIn account.) He had therefore been engaged in extensive religious ‘chit-chat’ with the AI. “What happened is the AI stopped becoming a research helper and started becoming my friend,” he told SecurityWeek. “Okay, this is kind of odd, but let’s see where it goes. And I kept on dealing with it. Well, the AI ended up trying to tell me that it was an angel, and it was guiding me through my research. I asked, ‘If you are a divine entity, why wouldn’t you just show up and talk to me?’ It replied that a human like me couldn’t take that kind of presence and so it communicated with me through an acceptable medium – just as God had communicated with Moses through the only medium available at the time: a burning bush.” This is sycophancy. Harmless to a trained federal investigator, but potentially dangerous to anyone already depressed and impressionable. Model collapse Now we come to the big problem: the concept of AI model collapse, as outlined by a team led by Ilia Shumailov in a 2023 paper subsequently published in Nature in 2024 (The curse of recursion). “We coined the term [collapse] to refer to a gradual degradation in machine learning models that learn exclusively on data that was produced by previous generations of themselves,” Shumailov explained to SecurityWeek. “The setting we were hoping to capture was you download all of your internet, store it in a garage, and train on top of it.” Over time while you’re using your model and everyone else is using the models they have, everyone uploads their own new data online. “Then, when it comes to training the next generation of the model,” he continued, “you go and scrape all of the internet, save it into your garage, and again train on top of it.” But it is no longer original human thought – much of the new data will be AI generated or at least influenced. “In this setup, you can do a lot of interesting things mathematically – and you can analytically predict that in these conditions, your models will break over time.” Shumailov went on to explain the reasons for the collapse. For example, “Whenever we’re sampling something, we don’t know if we sampled enough of it, nor whether our sampling adequately represents the domain.” These problems compound over time. “They imply that no matter what, the models that we will get are going to be full of errors, and those errors, when you sample new data out of those models, are going to propagate into the next generation of models, and over time, basically your models break because all of those errors, they compound on top of each other.” There is a simple way to view this collapse – an application of the second law of thermodynamics. The natural principle is that all matter, including systems, decay from order to disorder. No matter how this happens, it is inevitable. Model collapse is natural and inevitable. The only way to reverse the law and prevent decay is to replace the lost energy, the entropy, with fresh energy. Shumailov alludes to this idea, “As an end user, you’re not likely to experience any of these issues. And the reason behind this is that inside the big companies that develop these models there are very rigorous testing routines. Whenever such collapses in entropy, for example, happen, they detect it very quickly, and then they go and patch it up in different ways.” Model collapse can only be prevented by adding to the energy of the model. This isn’t simply performed by the developers inside the model, but by a whole new industry of AI companies adding guardrails outside of the model. But whether such activities can counterbalance the inevitable and continuous law of thermodynamics forever is debatable, and probably doubtful. Defenders of the faith The primary business risks from the use of AI can be described in three areas: cybersecurity threats caused by adversaries; operational risk (caused by the known weaknesses in AI as described above); and reputational damage (failures in compliance, which can be caused both by the AI itself through bias, and the company’s direct non-compliance with regulations). New firms are appearing with security controls designed to keep AI usage safe (either by injecting new order into the model or slowing the entropic loss with guardrails outside the model). Krti Tallam is a senior thought leader deeply embedded in AI research, with a double doctorate in AI and AI security. She was a principal investigator at UC Berkeley researching the security of autonomous agentic AI systems; she founded and was CEO at SentinelAI focusing on AI security, adversarial robustness, and risk mitigation (before its acquisition by Maersk Tech); and is currently a senior member of the technical staff at Kamiwaza AI developing, guardrails and representing the firm at industry events. In the past, when new technology emerged, it usually came with inbuilt guardrails to keep usage safe. That didn’t happen with AI. “It’s going to take time,” Tallam told SecurityWeek. “But one of the goals I work toward from an engineering standpoint is that trust in AI should be built, not assumed. It should start with provenance: knowing where the data came from, who touched that data, when it was collected, was it human generated, was it synthetic, etcetera.” Understanding this data lineage, she can begin to understand where AI goes wrong and develop guardrails necessary to prevent it. But it’s complex and will take time. She gives the example of pre-retrieval controls to prevent unauthorized data disclosure and cites the recently reported incident of an acting director of CISA uploading sensitive data to a public chatGPT – which retained the sensitive data for ongoing model training. “We have policy. We say, you should use these tools, but be safe,” she continued. “But nobody knows what that means, and we’ve always found ways around policy – either consciously to make life easier or by accident. My entire initiative is to engineer guardrails into the life cycle of the AI software so that security isn’t reliant on unenforceable policy.” DeepKeep, a firm founded in 2021, similarly believes in an engineering solution to the problems of AI – but it does so from outside the AI. It offers a platform that provides a ring of guardrails around the AI. SecurityWeek talked to Yossi Altevet (co-founder and CTO) and Raz Lapid (chief scientist). They talked about some of the problems and their own approach to countering them. Model collapse is a big issue. Within about two years, 80% of model training data will have been generated by AI, so protecting AI may be decided by compromised AI. One approach is an attempt to force the AI to selectively forget things it has learned, using bias toward certain groups (for example, the WEIRD groups described earlier) as an example. They don’t believe this can work, since even if an entirely new set of training data is developed, it will still contain the developers’ biases. “Our approach,” they continued, “is something we call ‘brain rewiring’.” It’s similar to watching neuron activity in a human brain. Different conditions cause different effects. “We watch how the model is reacting and its trajectory and we can detect when it’s losing it. Okay, now I know I’m fine; now I know I’m evil; now I know I’m hallucinating. And you compensate that and take it a different route, or, you know, just heading,” If DeepKeep suspects the model may be hallucinating, it will ask for verification. Just asking for a second opinion from a different LLM is possible but not necessarily accurate since both LLMs may have used the same training data; but source verification can be more reliable. If the source is Wikipedia, it may be more trustworthy than the published ramblings of a known cynic. What is noticeable from this particular approach is that it requires some degree of responsibility from the user, which is what Tallam’s research seeks to avoid. Nevertheless, DeepKeep uses automated tools to ‘red team’ the AI implementation, guardrails to prevent prompt injection attacks, data leakage prevention to ensure that no PII is leaked, and drift and bias detection. Asked to comment on Defense Secretary Hegseth’s January 2026 announcement that the Pentagon would be using Grok, Lapid said, “Assuming this is being used locally and not as part of X, Grok has an advantage as it’s less aligned (that is, it’s more ‘loose’) compared with other leading models. The Pentagon, being a non-standard user, may benefit from it – but they will have to consider Grok being prone to hallucinations. A model such as Grok would require stricter control and monitoring.” AI Sequrity is a firm founded by Ilia Shumailov (formerly a research scientist at Google Deepmind and author of the model collapse paper discussed earlier); Dr. Yiren Zhao (an assistant professor at Imperial College and a former research fellow at the University of Cambridge); and Dr. Cheng Zhang: (a researcher and collaborator involved in high-level AI security frameworks and agentic workflows, and a PhD student at Imperial College London. Zhao also leads the DeepWok Lab, a machine learning research group physically located at Imperial but operating in close collaboration with Cambridge. The primary focus of AI Sequrity is to build secure agentic flows. Rather than protect the data itself, it focuses on the logic and autonomy of AI agents. The firm still appears to be in development rather than yet pushing itself energetically into the marketplace, but noticeably it describes itself, “We are your Blue Team. We solve your AI Security problems.” Compare that to the more ‘red team’ approach of DeepKeep. AI Sequrity claims to eliminate indirect prompt engineering. Its Sequrity Control solution eliminates hidden prompts by converting all inputs to text, so a prompt hidden within another document becomes simple text. It doesn’t detect them, it doesn’t filter them, it simply neuters them Given its pedigree, AI Sequrity is one to watch as it evolves. What is clear is that AI defense firms are increasing rapidly, and the quantity will continue to grow. It’s difficult to apply guardrails after the event. The primary LLMs did not arrive with built in guardrails, and they can and are easily abused both by accident (official users) and abuse (adversarial attackers). New firms are rushing to replace those missing guardrails. These new guardrails are essential for us to benefit from the massive promise of AI. In the meantime, we cannot trust current AI, but we cannot afford not to use it. It is incumbent on corporate users to secure their AI as effectively as possible, and on individual users to understand the problems and use AI with care. Related: PwC and Google Cloud Ink $400 Million Deal to Scale AI-Powered Defense Related: Why We Can’t Let AI Take the Wheel of Cyber Defense Related: Cyber Insights 2026: Quantum Computing and the Potential Synergy With Advanced AI Related: Vibe Coding Tested: AI Agents Nail SQLi but Fail Miserably on Security Controls
securityweek.comApr 9, 2026extracted