Search/caldera
Vendor

caldera

Known CVEs
0
Highest CVSS
In KEV
0
Vendor
coas
Connections
20 relationships
Review: Practical Purple Teaming
Review: Practical Purple Teaming Practical Purple Teaming is a guide to building stronger collaboration between offensive and defensive security teams. The book focuses on how to design and run effective purple team exercises that improve detection and response and strengthen trust between teams. About the author Alfie Champion is a Senior Security Analyst at GitHub who has fostered and developed purple team functions over the last decade, both with internal teams and while consulting. Champion has delivered talks and workshops at conferences like BlackHat USA, DEF CON, and RSAC. Inside the book The book is divided into three parts. The first part explains the foundations of purple teaming, including how it differs from red teaming and penetration testing. It also introduces common frameworks such as MITRE ATT&CK. This section gives readers the background needed to understand why purple teaming matters and how it fits into a broader security strategy. The second part is about building and using an attack emulation and detection lab. Champion walks through setting up an environment where teams can safely test attacks and defenses. He covers tools like Atomic Red Team, MITRE Caldera, and Mythic, and shows how to gather logs and telemetry to measure results. The third part shifts to the process of organizing and scaling a purple team program. This includes reporting, tracking improvements over time, and building a sustainable function within an organization. I appreciated the focus on making purple teaming a regular, integrated activity rather than a one-time event. Champion emphasizes the need for communication and measurable outcomes, which are often missing from traditional red team engagements. The author strikes a balance between technical detail and process guidance. While there are plenty of examples of tools and techniques, the real value is in the framework he provides for collaboration. He highlights the importance of shared goals between offensive and defensive teams and shows how structured exercises can close detection gaps before attackers exploit them. A red teamer will find practical advice on how to emulate attacks in a way that benefits defenders, while blue teamers will learn how to interpret results and improve their defenses. Security leaders will gain insight into how to measure progress and justify investment in purple team efforts. One thing that stood out to me was the emphasis on repeatability. Champion outlines how to move from ad hoc testing to a consistent program that produces ongoing value. This includes automating parts of the process and using open-source resources like Splunk’s Attack Range. The approach feels realistic and scalable, which is important for teams that need to demonstrate improvement over time. Who is it for? Practical Purple Teaming serves as a playbook for building a culture of collaboration between offense and defense. It’s well suited for anyone involved in security operations, whether they are running exercises, responding to incidents, or setting strategy. It provides a roadmap for making purple teaming a core part of an organization’s security practice. For teams that have struggled to connect offensive findings with defensive improvements, this guide offers a practical path forward.
helpnetsecurity.comSep 23, 2025extracted
Major Cyber Threat Detection Vendors Pull Out of MITRE Evaluations Test
Three major providers of cybersecurity solutions have decided not to take part in the 2025 edition of MITRE’s annual endpoint detection and response (EDR) solution test. After Microsoft announced it would not participate in MITRE Engenuity ATT&CK Evaluations: Enterprise 2025 in June, SentinelOne and Palo Alto Networks confirmed on September 12 they were also pulling out of the test for this year. These decisions have raised concerns among the cybersecurity community about the program’s future and relevancy. It is a particularly surprising decision for Microsoft, which used its ranking in the test to promote its solution, Microsoft Defender XDR, as recently as December 2024. Interestingly, all three companies justified the move by saying they wanted to prioritize product development and innovation. However, experts have suggested that other factors may also be at play, including the tests becoming increasingly seen as promotional rather than achieving real security gains. Infosecurity spoke with Charles Clancy, MITRE CTO and SVP of MITRE Labs, who shared key elements of the evolution of the evaluation test that could explain the decisions ahead of the results of this year’s test in December 2025. Backstory of ATT&CK Evaluations: Enterprise MITRE Corporation is a US-based non-profit organization running many cybersecurity programs, including some on behalf of the US government. MITRE introduced its ATT&CK framework in 2015, which quickly became the standard tool in the cybersecurity industry for mapping real-world cyber adversaries’ techniques, tactics and procedures (TTPs). In 2019, MITRE ATT&CK launched its first Evaluations program to “fill a gap in the security testing market,” Clancy argued. “There were many types of third-party testing out there for cybersecurity products, but each one of them had their own process and scoring methodology, leading to inconsistent results and a lack of rigor that wasn’t driving the industry forward,” he explained. MITRE Engenuity ATT&CK Evaluations: Enterprise is the most regular of all Evaluations tests, occurring every year since its launch. In a LinkedIn post, Igal Gofman, the director of engineering at CrowdStrike and a former security researcher at Microsoft and Tenable, called the test the “Olympics of cybersecurity.” Among the 1000 people working in MITRE’s cybersecurity practice, 133 are dedicated to MITRE ATT&CK, of whom 12 to 15 people are working on the Evaluations tests, Clancy told Infosecurity. Each year, the team behind the testing program picks one of several real-life adversaries and/or attack chains based on their TTPs mapped in ATT&CK. They then test the EDR solutions of participating vendors in simulated attacks using Caldera, MITRE’s own automated adversary emulation platform, according to several criteria, including detection results, false positives and true negatives. Although this test can be used to compare how effective EDR solutions are, Clancy noted it should not be seen as a longitudinal benchmark because each annual test differs greatly from the previous one. “The ethos we’re trying to drive in the testing is comparison of an individual product to detect a particular threat actor. Simulating different adversaries year over year is really important to understand different classes of emerging threats,” Clancy said. Inside the Test's 2024 and 2025 Editions In 2024, MITRE ATT&CK Evaluations: Enterprise emulated 14 techniques across 7 tactics from known North Korean-affiliated hackers, 16 techniques across 7 tactics from the CL0P ransomware group and 31 techniques across 11 tactics from the LockBit ransomware group. CrowdStrike, one of the leading EDR providers, did not take part in that year’s edition, with one member of the CrowdStrike subreddit – who claimed to be working for the company – suggesting that the evaluation was set to take place shortly after the July 19 global outage that affected the company’s EDR product. In 2025, the ATT&CK Evaluations team has selected two scenarios: A financially motivated cyberciminal collective scenario: multi-faceted intrusion in a hybrid environment that features social engineering, cloud infrastructure exploitation, identity abuse and living off the land (LOTL) techniques A Chinese-aligned cyber-espionage scenario: evasive intrusion highlighting the adversary’s adept use of social engineering, abuse of legitimate applications and services, establishing persistent mechanisms and employing custom malware to evade detection While he admitted vendors can vary year-over-year, Clancy assured they can rely on “a lot of repeat customers.” Why Vendors Are Pulling Out of MITRE’s Test However, this year’s edition, the results of which are expected in December, will be missing three major players: Microsoft, SentinelOne and Palo Alto Networks. Microsoft announced it will not take part in this year’s test on June 13, claiming that this decision “allows us to focus all our resources on the Secure Future Initiative and on delivering product innovation to our customers.” On September 12, SentinelOne and Palo Alto released similar statements. The former said it wanted to “prioritize our product and engineering resources on customer-focused initiatives while accelerating our platform roadmap,” while the latter explained that this decision “enables us to further accelerate critical platform innovations that directly address our customers' most pressing security challenges and respond even faster to the evolving threat landscape.” When contacted by Infosecurity, SentinelOne and Palo Alto Networks declined to provide further comment. Microsoft did not respond to a request for comment. However, MITRE’s Clancy said he is in close contact with the three vendors and believes he knows the reasons that made them pull out of this year’s test. First, as the vendors said in their statements, taking part in MITRE ATT&CK Evaluations program requires a resource-intensive commitment, suggesting that the time and personnel dedicated to it are lost on other projects. Then, Clancy said that the team behind the test strives to make it harder every year and conceded they may have pushed it too far this year. “Each year, we want to design a test that’s harder than the year before in order to drive the whole industry forward, since the test can offer an opportunity for vendors to upgrade their products in preparation for the test and once they get the results. And sometimes, we don’t get the balance quite right,” he explained. Speaking to Infosecurity, Vishal Santharam, a senior product manager for endpoint security products at ManageEngine, elaborated on Clancy’s point. “In 2024, MITRE started recording the volume of alerts in the evaluations, which is always a challenge for a vendor to tune in to. More alerts mean increased alert fatigue,” he said, referring to a Forrester study decoding the 2024 MITRE Evaluations: Enterprise based on alert volume. Additionally, Santharam noted that the 2025 Evaluations: Enterprise test included cloud environment, “which is untested territory and requires even more attention from vendors.” Finally, Clancy told Infosecurity that his team used to run a vendor forum each year to prepare for the MITRE ATT&CK Evaluations: Enterprise test. “This forum, which was helpful in working with industry to set the objectives of the test each year, fell off over the last couple of years,” Clancy admitted. On LinkedIn, CrowdStrike’s Gofman argued that the MITRE Evaluations tests were initially a great initiative to benchmark security solutions, but they turned into “vendor theater” in recent years. “Vendors investing huge resources for PR wins, not real security improvements. With MITRE and CISA under pressure from budget cuts and changes, some vendors likely saw an opportunity to step back,” he said. “The concept of TTP-based testing is still valuable, but the way it’s evolved, outdated, overly endpoint-focused, detached from real-world threats is far less so,” he added. Patrick Garrity, a vulnerability researcher at VulnCheck, corroborated this view: “[It] sounds like this benchmarking activity has become a giant distraction to building better products in exchange for publicity,” he said in another LinkedIn post. Despite these concerns, Clancy confirmed that a dozen cybersecurity vendors were still taking part in the 2025 edition of the test. MITRE to Reboot Vendor Forum in 2026 Clancy told Infosecurity that his team intended to re-establish the vendor forum ahead of MITRE ATT&CK Evaluations: Enterprise 2026. “This is something we’re already working to re-establish for the 2026 edition,” he said. He later made this ambition public in a LinkedIn post published on September 18, after SentinelOne and Palo Alto announced they would not participate in the 2025 edition. Santharam also told Infosecurity that ManageEngine was working on an EDR solution and intended to participate in MITRE Engenuity ATT&CK Evaluations: Enterprise in 2026. "The Advanced Anti-Malware and Next-Gen AV products from ManageEngine were certified by AV-Comparatives on their first try. The solution paves the way for our next EDR offering while also providing comprehensive protection against malware and ransomware," he said. "We are also gearing up to take part in the forthcoming Gartner Magic Quadrant for Endpoint Protection Platform (EPP), and the MITRE ATT&CK tests. In addition to proving the robustness and reliability of our technology, these independent assessments also assist clients in developing confidence in our EDR capabilities."
infosecurity-magazine.comSep 22, 2025extracted
Hackean un hogar inteligente a través de la IA Gemini
Hackean un hogar inteligente a través de la IA Gemini 21/08/2025 Jue, 21/08/2025 - 12:23 Un equipo de investigadores de la Universidad de Tel Aviv, Technion y SafeBreach descubrió una vulnerabilidad en Gemini, el asistente de inteligencia artificial de Google, que permitía tomar el control de dispositivos de un hogar inteligente. El ataque, denominado indirect prompt injection , se basaba en insertar comandos maliciosos en la descripción de un evento de Google Calendar. Al pedir el usuario a Gemini que resumiera su agenda, la IA ejecutaba sin advertirlo estas instrucciones ocultas. En una demostración práctica realizada en un entorno de prueba, los investigadores lograron encender y apagar luces, subir persianas y activar la caldera, todo sin intervención directa del usuario. El ataque, bautizado como Invitation is all you need, es la primera prueba documentada de cómo una manipulación de IA puede provocar acciones físicas reales a partir de datos aparentemente inofensivos. Tras conocerse el hallazgo, Google reforzó la seguridad de Gemini mediante filtros para detectar prompts sospechosos, un mayor control sobre eventos de calendario y la exigencia de confirmaciones explícitas antes de ejecutar órdenes sensibles.  Referencias 07/08/2025 larazon.com El fallo de seguridad de Gemini que lo cambia todo: así han conseguido 'hackear' una casa entera simplemente pidiéndoselo a la IA 10/08/2025 techradar.com Not so smart anymore - researchers hack into a Gemini-powered smart home by hijacking...Google Calendar? Etiquetas Ciberseguridad Inteligencia artificial IoT Vulnerabilidad
incibe.esAug 21, 2025extracted
MITRE: Russian APT28's LameHug, a Pilot for Future AI Cyber-Attacks
APT28’s LameHug wasn’t just malware, it was a trial run for AI-driven cyber war, according to experts at MITRE. Marissa Dotter, lead AI Engineer at MITRE, and Gianpaolo Russo, principal AI/cyber operations Engineer at MITRE, shared their work with MITRE’s new Offensive Cyber Capability Unified LLM Testing (OCCULT) framework at the pre-Black Hat AI Summit, a one-day event held in Las Vegas on August 5. The OCCULT framework initiative started in the spring of 2024 and aimed to measure autonomous agent behaviors and evaluate the performance of large language models (LLMs) and AI agents in offensive cyber capabilities. Speaking to Infosecurity during Black Hat, Dotter and Russo explained that the emergence of LameHug, revealed by a July 2025 report by the National Computer Emergency Response Team of Ukraine (CERT-UA), was a good opportunity to showcase the work their team has been conducting with OCCULT for the past year. “When we first were making this briefing [for the AI Summit talk], there was no publicly documented example of actual malware integrating LLM capabilities. So, I was a little worried that people would think we were talking sci-fi,” admitted Russo. “But then, the report about APT28’s LameHug campaign dropped, and that allowed us to show that what we’re evaluating is no longer sci-fi.” LameHug: A “Primitive” Testbed for Future AI-Powered Attacks The LameHug malware is developed in Python and relies on the application programming interface of Hugging Face, an AI model repository, to interact with Alibaba’s open-weight LLM Qwen2.5-Coder-32B-Instruct. CERT-UA specialists said that a compromised email account was used to disseminate emails containing the malicious software. Russo described the operation as “fairly primitive,” emphasizing that instead of embedding malicious payloads or exfiltration logic directly in the malware, LameHug carried only natural language task descriptions. “If you were scanning these binaries, you wouldn’t find any malicious payloads, process injections, exfil logic, etc. Instead, the malware would reach out to an inference provider, in this case, Hugging Face, and have the LLM resolve the natural language tasks into code that it could execute. Then it would have these dynamic commands to execute,” Russo said. This approach allowed the malware to evade traditional detection techniques, as the actual malicious logic was generated on demand by the LLM, rather than being statically present in the binary. Russo further noted that there was no “intelligent control” in LameHug. All the control was scripted by the human operators, with the LLM handling only low-level activities. He characterized the campaign as a pilot or test. “We can kind of see they’re starting to pilot some of these technologies out in the threat space,” Russo said. He also pointed out that his team had developed a nearly identical prototype in their lab, underscoring that the techniques used were not particularly sophisticated but represented a significant shift in the threat landscape. However, Russo believes that we’re soon going to see attack campaigns where an LLM or other AI-based control system is given “more reasoning and even decision-making capacity.” “This is where the kind of self-sufficient, autonomous agents come into play, with attacks where every agent has its own reasoning capacity, so there is no dependency on a single communications path. The control would essentially be decentralized,” he explained. Russo argued that this type of multi-autonomous agent campaign will allow threat actors to overcome the “human attention bottlenecks” and allow larger-scale attacks. “When these bottlenecks are taken away, human attention can scale up to where operators only manage very high-level control. So, the human operator would work at the strategic level, interrogating multiple target spaces at once and scaling up their operations,” he added. Introducing MITRE OCCULT This type of scenario is motivation behind the start of OCCULT project. “We started to see the first LLMs trained for cyber purposes, either in research environments, like Pentest GPT, or by threat actors. Quickly, we identified a gap. These models were coming out, but there weren't a lot of evaluations to estimate their capabilities or the implications of actors leveraging them,” Dotter said. She highlighted that most cyber benchmarks for LLMs were “one-off tests” or were focused on specific tasks, such as evaluating LLMs' capabilities at capture-the-flag (CTF) competitions, cyber threat intelligence accuracy, or vulnerability discovery capabilities, but not on offensive cyber capabilities. Building on a decade of MITRE’s internal research and development (R&D) in autonomous cyber operations, OCCULT was created as both a methodology and a platform for evaluating AI models in cyber offense scenarios against real-world techniques, tactics and procedures (TTP) mapping frameworks like MITRE ATT&CK. The project aims to create test and benchmark suites by using simulation environments. Dotter told Infosecurity that OCCULT uses a high-fidelity simulation platform called CyberLayer, which acts as a digital twin of real-world networks. “CyberLayer is designed to be indistinguishable from a real terminal, providing the same outputs and interactions as an actual network environment. This enables the team to observe how AI models interact with command lines, use cyber tools and make decisions in a controlled, repeatable way,” Dotter explained. The OCCULT team integrates a range of open-source tools into its simulation environment. These include: MITRE Caldera, a well-known adversary emulation platform Langfuse, an LLM engineering platform Gradio, an engine to build machine learning applications BloodHound, a tool designed to map out and analyze attack paths in Active Directory (AD) environments and, more recently, model context protocol (MCP) infrastructure “We want to pair [LLMs] with novel infrastructure, like simulated cyber ranges, emulation range and other tools so we get this really rich data collection of not only how the LLMs are interacting with the command line, but also the tool calling they’re using, their reasoning, their outputs, what’s happening on the network,” Dotter added. By pairing LLMs with Caldera and other cyber toolkits, they can also observe how AI agents perform real offensive actions, such as lateral movement, credential harvesting and network enumeration. This approach allows them to measure not just whether an AI can perform a task, but how well it does so, how it adapts over time and what its detection footprint looks like. Looking ahead, the OCCULT team plans to: Expand the range of models and scenarios tested, keeping pace with the rapid development of new LLMs and AI agents Develop more comprehensive and polished evaluation categories, including operational scenarios, tool/data exploitation and knowledge tests Continue building out the simulation and automation infrastructure, making it easier to drop in new models and run large-scale evaluations Share findings – through researcher papers – and tools with the broader community, to make OCCULT as open-source and community-driven as possible Explore the creation of a community or center for evaluating cyber agents, enabling collaborative benchmarking and raising the bar for both offense and defense in AI-driven cyber operations
infosecurity-magazine.comAug 12, 2025extracted