{"id":26178,"date":"2026-09-21T15:13:28","date_gmt":"2026-09-21T23:13:28","guid":{"rendered":"https:\/\/www.palada.net\/index.php\/2026\/09\/21\/owasp-top-10-for-llm-apps-2026-excessive-agency-risk-on-the-rise\/"},"modified":"2026-09-21T15:13:28","modified_gmt":"2026-09-21T23:13:28","slug":"owasp-top-10-for-llm-apps-2026-excessive-agency-risk-on-the-rise","status":"publish","type":"post","link":"http:\/\/www.palada.net\/index.php\/2026\/09\/21\/owasp-top-10-for-llm-apps-2026-excessive-agency-risk-on-the-rise\/","title":{"rendered":"OWASP Top 10 for LLM Apps 2026: Excessive agency risk on the rise"},"content":{"rendered":"<div class=\"rich-text_richText__UyrDZ\" data-anchor-headings=\"true\" data-component=\"rich-text\" data-reader-view=\"false\">\n<div class=\"payload-richtext\">\n<p>The Open Worldwide Application Security Project\u2019s newest OWASP Top 10 for LLM Applications includes a change that puts a spotlight on the rise of AI risk.\u00a0<\/p>\n<p>While prompt injection and the disclosure of sensitive information still hold the top two spots, excessive agency jumped from sixth place in 2025 to third in 2026 (pushing supply chain and data model poisoning risks down to the fourth and fifth slots).<\/p>\n<p>Now part of OWASP\u2019s GenAI Security Project and created\u00a0 by hundreds of AI security experts, the <a href=\"https:\/\/genai.owasp.org\/resource\/owasp-genai-llm-top-10-2026\/\" rel=\"noopener noreferrer\" target=\"_blank\"><span style=\"text-decoration:underline\">OWASP Top 10 for LLM Applications 2026<\/span><\/a>, provides updated rankings, expanded threat coverage, and new research based on real-world AI security incidents, all organized with the practitioner in mind.<\/p>\n<p style=\"padding-inline-start:40px\"><em>\u201cThe guide provides practical explanations, attack scenarios, and actionable mitigations for developers, architects, security teams, and CISOs, while mapping risks to leading industry frameworks including NIST, MITRE ATLAS, CWE, and the OWASP Top 10 for Agentic Applications.\u201d<\/em><br \/>\u2014OWASP GenAI Security Project<\/p>\n<p>Here are key takeaways from the OWASP Top 10 for LLM Applications 2026 \u2014 and why it&#8217;s time to think beyond security controls.<\/p>\n<p><strong>[ Join webinar: <\/strong><a href=\"https:\/\/www.reversinglabs.com\/events\/autonomy-not-autopilot-agentic-soc\"><strong>Autonomy, Not Autopilot: Get Real About the Agentic SOC<\/strong><\/a><strong> ]<\/strong><\/p>\n<h2 id=\"a-whole-new-attack-surface\"><strong>A whole new attack surface<\/strong><\/h2>\n<p>Jason Soroko, a senior fellow at Sectigo, said prompt injection remains the top risk because the flaw is fundamentally architectural.<\/p>\n<p style=\"padding-inline-start:40px\"><em>\u201c[With <\/em>prompt injection], i<em>nstructions and data share one channel, the context window, and no equivalent of the parameterized query exists to separate them, so every mitigation lowers probability without reaching zero.\u201d<\/em><br \/>\u2014<a href=\"https:\/\/www.linkedin.com\/in\/jason-soroko-19b41920?originalSubdomain=ca\"><span style=\"text-decoration:underline\">Jason Soroko<\/span><\/a><\/p>\n<p><a href=\"https:\/\/www.linkedin.com\/in\/pete-pickerill-2347835\"><span style=\"text-decoration:underline\">Pete Pickerill<\/span><\/a>, co-founder of Liquibase, said no one has solved the problem. He noted that attackers can now conceal instructions inside images, audio, and documents, adding that every tool that an AI system connects to represents another potential entry point. What\u2019s changed, he said, is the level of risk: Where a manipulated chatbot can produce an embarrassing response, a manipulated agent can take genuinely harmful action.<\/p>\n<p>Jeremy London, director of engineering, AI and threat analytics at Keeper Security, said excessive agency\u2019s rise reflects how organizations are deploying AI. When the list was first created, in 2024, most LLM applications were simple chat interfaces or single-step tools. By 2026, agents had become the norm \u2014 systems with persistent memory, tool access, file and API permissions, and the ability to carry out multistep tasks with varying degrees of autonomy, he said.<\/p>\n<p style=\"padding-inline-start:40px\"><em>\u201cThe model is not just returning a response. It is taking action, which presents a qualitatively different attack surface.\u201d<\/em><br \/>\u2014<a href=\"https:\/\/www.linkedin.com\/in\/jeremyclondon\"><span style=\"text-decoration:underline\">Jeremy London<\/span><\/a><\/p>\n<p>Larry Maccherone, founder and CTO of Lumenize, is skeptical about excessive agency\u2019s ranking, arguing that a self-repairing system can only exist if the AI is given enough agency to fix itself \u2014 a form of recursive self-improvement, with humans providing high-level direction while the machine handles the actual repairs. Clipping an AI system\u2019s agency in the name of safety will limit its ability to find and close its own security gaps, he said, leaving organizations with something less capable and less secure.\u00a0<\/p>\n<p style=\"padding-inline-start:40px\"><em>\u201cThe cure for excessive agency is more agency \u2014 pointed at the right thing.\u201d<\/em><br \/><em>\u2014<\/em><a href=\"https:\/\/www.linkedin.com\/in\/larrymaccherone\"><span style=\"text-decoration:underline\">Larry Maccherone<\/span><\/a><\/p>\n<h2 id=\"why-including-incident-data-matters\"><strong>Why including incident data matters<\/strong><\/h2>\n<p>New to the 2026 list is the addition of incident data in the ranking process. In the <a href=\"https:\/\/genai.owasp.org\/resource\/owasp-genai-llm-top-10-2026\/\"><span style=\"text-decoration:underline\">2026 Top 10 document<\/span><\/a>, the project leaders explained that this and every previous version of the list has been built on judgment, with hundreds of practitioners weighing in on what matters most.<\/p>\n<p>This year, however, the project tested that vote against a record of what has actually gone wrong. The project pulled together records on 7,714 incidents from public vulnerability databases and an AI-harm database, then built classifiers that read them and identified 6,639 that carried enough detail to sort, said Steve Wilson, founder and co-chair of the OWASP GenAI Security Project.<\/p>\n<p style=\"padding-inline-start:40px\"><em>\u201cThis was the first time in three years of doing this that we had enough data for it to really weigh into the decision making.\u201d<\/em><br \/><em>\u2014<\/em><a href=\"https:\/\/www.linkedin.com\/in\/wilsonsd\"><span style=\"text-decoration:underline\">Steve Wilson<\/span><\/a><\/p>\n<p>The first time the list was created, he noted, the risks had been purely theoretical.<\/p>\n<p>Still, the vote was given more weighting \u2014 three-quarters of the total \u2014 the project leaders said, because the list is a consensus product and a single noisy year of incident data could erroneously\u00a0 override the collective judgment of practitioners. The quarter-weight given to incident data, they explained, is enough to shift an entry by a tier when there\u2019s a wide gap between perception and evidence but not enough to let imperfect data rewrite the list unilaterally \u2014 and that balance determines every final ranking in the Top 10.<\/p>\n<p>Differences between belief and evidence can occur, said Advait Patel, a senior cloud security and site reliability engineer and creator of OWASP DockSec, because the two are measuring different things. The community vote reflects what practitioners are worried about and focused on, he said, while the incident data shows what has caused real-world harm. The gap between the two, he said, is where things get interesting.\u00a0<\/p>\n<p>Prompt injection, for instance, generates relatively few clean public incidents \u2014 not because the risk has disappeared, but because so many people are actively working to defend against it, making the low incident count a sign that those defenses are working. Misinformation runs in the opposite direction, he added; teams didn\u2019t rank it highly, yet it appears frequently in the incident record, because a confidently wrong AI answer can now feed directly into downstream tools and code rather than simply appearing in a chat window.<\/p>\n<h2 id=\"why-consumption-concerns-matter\"><strong>Why consumption concerns matter<\/strong><\/h2>\n<p>Two categories that had sat at the bottom of the 2025 list, misinformation and unbounded consumption, moved up to the sixth and seventh spots in 2026.<\/p>\n<p>Eilon Cohen, head of security research at Pillar Security, explained unbounded consumption by saying that a single agent request can consume far more than just model tokens. A request might trigger an extended reasoning process, repeated calls back to the model, queries to external APIs, code execution, or new cloud workloads \u2014 and a failed step might be retried repeatedly.\u00a0<\/p>\n<p style=\"padding-inline-start:40px\"><em>\u201cOne user request can turn into dozens or hundreds of operations behind the scenes. That creates both denial-of-service and denial-of-wallet risks.\u201d<\/em><br \/>\u2014<a href=\"https:\/\/www.linkedin.com\/in\/eilon-cohen?originalSubdomain=il\"><span style=\"text-decoration:underline\">Eilon Cohen<\/span><\/a><\/p>\n<p>Whether the problem is caused by an attacker or simply involves an agent caught in a loop, he said, the result can be an exhausted budget, drained SaaS quotas, occupied cloud infrastructure, and degraded service for everyone else. Rate-limiting the chat endpoint alone is no longer sufficient, he cautioned; organizations need per-user and per-session budgets, limits on agent steps and tool calls, timeouts, concurrency controls, and circuit breakers that can halt abnormal consumption before it spreads to connected systems.<\/p>\n<p>Meanwhile, improper output handling fell from fifth place in 2025 to the bottom of the list in 2026, and system prompt leakage \u2014 renamed and rescoped as hidden context exposure \u2014 and vector and embedding weaknesses each slipped a notch, from seventh and eighth, respectively, in 2025 to eighth and ninth in 2026.<\/p>\n<p>DockSec creator Patel cautioned against reading the drop of embedding weaknesses as a sign the risk itself had diminished. In fact, he said, the category\u2019s scope expanded this year to include things such as terminal escape sequences and renderers that automatically fetch links.<\/p>\n<p style=\"padding-inline-start:40px\"><em>\u201cIt dropped because everything above it got more urgent, not because output handling got safe. It is a relative move on a crowded list.\u201d<\/em><br \/><em>\u2014<\/em><a href=\"https:\/\/www.linkedin.com\/in\/advaitpatel93\"><span style=\"text-decoration:underline\">Advait Patel<\/span><\/a><\/p>\n<h2 id=\"prompt-injection-death-and-taxes\"><strong>Prompt injection, death, and taxes<\/strong><\/h2>\n<p>The project leaders\u2019 core recommendation is to stop trying to build a model that is immune to manipulation and instead improve the surrounding system so that when the model does get fooled \u2014 which it will \u2014 nothing critical breaks. That posture, they wrote, runs through all 10 entries on the list, and for the first time, the project can back it with evidence rather than asking practitioners to take it on faith.<\/p>\n<p>The OWASP GenAI Security Project\u2019s Wilson said that when the list debuted three years ago, the community had only a general sense of how the models worked, and at the time, prompt injection felt comparable to SQL injection \u2014 something that could largely be avoided with sufficient care. That assumption, he said, has proved false; the common wisdom now is that prompt injection is as unavoidable as death and taxes. While frontier model providers continue to improve their models\u2019 resistance to manipulation, he said, attackers are advancing their prompt-injection techniques even faster than the models are learning to detect them.<\/p>\n<p>Dave Hayes, vice president of product at FusionAuth, said the recommendation about the surrounding system is sound because it shifts the problem to a layer that security teams can control. There\u2019s no guaranteeing a model won\u2019t be fooled, he said, because accepting untrusted input is inherent to what it does. But what that model is permitted to do once it\u2019s fooled is a permissions question, and permissions can be made deterministic even when the model itself is probabilistic.\u00a0<\/p>\n<p>That\u2019s why the real work should happen upstream of the model, he said, ensuring that whatever credential it holds can reach only what it strictly needs, will expire quickly, and can be revoked.\u00a0<\/p>\n<p style=\"padding-inline-start:40px\"><em>\u201cIf a fooled agent can only reach one record instead of the whole database, you\u2019ve turned a breach into a log entry.\u201d<\/em><br \/><em>\u2014<\/em><a href=\"https:\/\/www.linkedin.com\/in\/daveeurica\"><span style=\"text-decoration:underline\">Dave Hayes<\/span><\/a><\/p>\n<p>\u201cEvery mitigation on the list is about limiting damage after it succeeds,\u201d Liquibase\u2019s Pickerill added. That\u2019s a major shift in mindset, he said, moving from asking how to prevent the model from being tricked to asking what the agent is allowed to change and who verifies its work. For any agent touching production systems, verification needs to happen outside the model itself, he said.<\/p>\n<h2 id=\"its-time-to-think-beyond-security-controls\"><strong>It\u2019s time to think beyond security controls<\/strong><\/h2>\n<p>Lumenize\u2019s Maccherone said he agrees with the project leaders\u2019 guidance in principle but worries that the industry will misread it. Most security professionals, upon hearing \u201cContain the damage,\u201d will just reach for another control: another gate, another approval step, another policy engine standing in front of the agent asking permission. But you cannot inspect your way to a secure agent, he said.<\/p>\n<p>Traditional software behavior is a fixed artifact that can be reviewed, scanned, and signed off on, he explained. An agent\u2019s behavior, by contrast, is generated fresh at runtime with every execution and is never quite identical twice, and you can\u2019t pre-approve behavior that doesn\u2019t yet exist.<\/p>\n<p>The real answer, Macherrone said, is getting better feedback rather than improving control: sensing failures quickly, pinpointing their cause precisely, and letting the system adapt so the same failure doesn\u2019t recur. Security for agents will end up looking a lot more like site reliability engineering than traditional application security. His advice: Stop building guardrails and build a nervous system instead.<\/p>\n<blockquote class=\"blockquote-block_blockquote__2i0LW\" data-auto-quote=\"true\" data-component=\"block-quote-block\">\n<div class=\"rich-text_richText__UyrDZ\" data-component=\"rich-text\" data-reader-view=\"false\">\n<div class=\"payload-richtext\">\n<p><em><strong>The real answer, Macherrone said, is getting better feedback rather than improving control: sensing failures quickly, pinpointing their cause precisely, and letting the system adapt so the same failure doesn\u2019t recur.<\/strong><\/em><\/p>\n<\/div>\n<\/div>\n<\/blockquote>\n<p>Hayes said the OWASP Top 10 for LLM Applications 2026 reads like something compiled for a field that is maturing. Last year\u2019s edition, he said, still held out some hope that a model could be built that would resist being fooled. This year\u2019s update opens by assuming all models will be fooled and so advises teams to build systems where that doesn\u2019t matter much.\u00a0<\/p>\n<p>That containment mindset is fundamentally an identity and authorization question, he said \u2014 asking, \u201cWhat is this thing, and what is it allowed to do?\u201d Hayes said that if he were a developer working from this list, he\u2019d spend less budget hardening the model\u2019s inputs and more on limiting what the model and its agents can actually reach when they get things wrong \u2014 because they will get things wrong, and the list is now built around that reality.<\/p>\n<\/p>\n<\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>The Open Worldwide Application Security Project\u2019s newest OWASP Top 10 for LLM Applications includes a change that puts a spotlight on the rise of AI risk.While prompt injection and the disclosure of sensitive information still hold the top two spots, excessive agency jumped from sixth place in 2025 to third in 2026 (pushing supply chain and data model poisoning risks down to the fourth and fifth slots).Now part of OWASP\u2019s GenAI Security Project and created\u00a0 by hundreds of AI security experts, th<\/p>\n","protected":false},"author":4,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"colormag_page_container_layout":"default_layout","colormag_page_sidebar_layout":"default_layout","footnotes":""},"categories":[32775],"tags":[],"class_list":["post-26178","post","type-post","status-publish","format-standard","hentry","category-reversinglabs"],"_links":{"self":[{"href":"http:\/\/www.palada.net\/index.php\/wp-json\/wp\/v2\/posts\/26178","targetHints":{"allow":["GET"]}}],"collection":[{"href":"http:\/\/www.palada.net\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"http:\/\/www.palada.net\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"http:\/\/www.palada.net\/index.php\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"http:\/\/www.palada.net\/index.php\/wp-json\/wp\/v2\/comments?post=26178"}],"version-history":[{"count":0,"href":"http:\/\/www.palada.net\/index.php\/wp-json\/wp\/v2\/posts\/26178\/revisions"}],"wp:attachment":[{"href":"http:\/\/www.palada.net\/index.php\/wp-json\/wp\/v2\/media?parent=26178"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"http:\/\/www.palada.net\/index.php\/wp-json\/wp\/v2\/categories?post=26178"},{"taxonomy":"post_tag","embeddable":true,"href":"http:\/\/www.palada.net\/index.php\/wp-json\/wp\/v2\/tags?post=26178"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}