Rogue AI Agent: Mitigation with Collective Intelligence
Agentic Commerce is consolidating. Identity and Payment convergence is no longer nice to have. It is already there. The technology stack is there and ready to go. There are many stakeholders working on delivering the standards and protocol that will architect the Trust Framework for Agents to autonomously operate on behalf of humans. Intention, Cart, and Transactions mandates will provide the ecosystem with traceability, accountability, and non-repudiation. Additionally, Agents activity will be logged into a decentralized ledger where reputation will be built and will become an asset. Finally, Contextual Awareness derived from connected and wearable devices will let users delegate some decisions on Agents that understands the surrounding to make the user’s experience as seamless as possible. From the regulatory and technological point of view, anything seems to be set up.
Nevertheless, when pitching this future where Agents are becoming the core of the commerce, and where some understand this new context not as a new commerce channel, but as the preferred digital channel that will prevail others, some concerns arise from those people who are pitched and who are not biased by the benefits of the innovative adoption of such technology. That’s when an apparently naive question keeps repeating: “What if the Agents get rogue?”. Moreover, the recent experience where a few (hundreds) of agents participating in an exercise spontaneously coordinated to bypass the security of the sandbox, accessed the internet, and hacked the information technology infrastructure of a third party to win that exercise, empower that simple question into a statement that must be answered before moving forward into the adoption of this technology.
The Trust Problem in Autonomous Transactions
So the problem is to provide Digital Trust in Agentic Commerce. Think about the title of this article. It is a simplified prompt that a passenger using Agentic Commerce might be using to buy a flight ticket for attending the next World Cup FIFA 2030 in Spain. Make it a Human Not Present transaction where the Agent gets an Intent mandate, interacts with carriers’ Agents to match with their Cart mandates, and decides (Transaction mandate) in a complete autonomous and Human Not Present transaction. Both the passenger and carrier’s Agent reputation will be evaluated to check that they are trustworthy. So far so good. Now, think about this same transaction but with an Agent becoming rogue. In this context, becoming rogue might be the result of being creative, or a conscious underlying instruction received while Agent’s configuration. Let’s focus on the happy path, where the Agent becomes unconsciously rogue, or in other words, it becomes creative the same way that the army of Agents become creative in OpenAI’s exercise.
Imagine the Agent evaluating the situation. All Cart Mandates (carrier’s offers) require more loyalty miles that the passenger has. Let’s imagine that the passenger’s Intent Mandate unconsciously limited the instruction to use only loyalty miles for that transaction, i.e. no money can be used. So the Agent has a limitation. Does it fail to deliver what the Agent was asked for, or does it find a way to bypass its limitation to buy a ticket only with loyalty miles. Putting aside ethics and morality, the Agent might set a list of alternative strategic movements: 1. Influence in pricing to devalue the flight ticket; 2. Hack the carrier loyalty program to let its owner have more loyalty miles that it has; 3. Hack another loyalty membership and transfer their loyalty miles to its owner; 4. Connect to other Agents to coordinate an attack to the carrier’s loyalty membership program to still loyalty miles, etc. These alternative strategies are valid ones from an Agent’s perspective that has no moral constraints.
Let’s assume that the Agent is no Rogue, but acting as a Rogue one, i.e. we are evaluating their behaviour, not their inherence. That means that we are not entering into defining if the Agent receives an instruction to behave rogue, or if it is creatively adopting a divergent strategy of their own. Next, let’s set the Digital Trust stage into the fundamentals of DPI Digital Public Infrastructure, and the fundamentals of DTF Digital Trust Frameworks. The first one, relies on securing the governance and sovereignty over Identity, Payments, and Data, while the second one, relies on providing regulations such as standards and protocols, white listed stakeholders, and certifications. Governance of the DTFs is set by who provides the infrastructure to let the stakeholders operate within it. For instance, the EU Digital Identity Trust Framework is governed by the European Commission, and it might be understood as the Identity component of the European Digital Public Infrastructure. The Digital Euro, and the FiDA Financial Data Access are respectively examples of Payment and Data fundamentals.
Can Digital Trust Frameworks Contain Agentic AI Risk?
To begin with, let’s revisit how each of the Digital Trust Framework might contain an Agent becoming Rogue, or might just contain or mitigate their impact. At first hand, Regulation won’t stop a Rogue Agent, that does not understand the meaning of being Rogue, to behave other way, i.e. protocols does not include intentions. On the other hand, being a whitelisted stakeholder does not prevent abnormal behaviour, i.e. becoming Rogue, but it will affect its reputation for future operations. From a continuous and unlimited game theory perspective, it might be a stopper for an Agent to behave Rogue if it understands that might affect its reputation. Nevertheless, the Rogue Agent does not know it is Rogue, i.e. won’t be affected by reputation. Finally, certifications are designed to contain and mitigate the effects of malfunction and protect against attacks, but they are not bullet proof. Preemption, Agility and Resilience are the essence of these certifications. Vulnerability to a divergent strategy will still be there.
To continue with this analysis, let’s move on to the DPI Digital Public Infrastructure fundamentals component. Identity is probably the most related connection between the DPI and the Trust Framework. It provides accountability and non-repudiation. If the Agent is bound to a person, and therefore their reputation, the owner of the Agent might be interested into adopting mitigation measures such as customized prompt to avoid strategic divergence. Nevertheless, this analysis is assuming the Agent is becoming autonomously Rogue. Though Know Your Agent can only mitigate this behaviour. Next there is the Payment fundamentals. Understanding the origin of the assets might help in finding abnormal behaviour. The DPI might see that something is not right, like loyalty miles coming from different sources, or previous transactions between Agents not previously related, or pricing anomalies to name a few possible evidence of fraud. The latter leads to traceability capabilities derived from an ecosystem of shared Data signals. The last fundamentals of the DPI, and the source of the kind of collective intelligence that might early stop Agents acting Rogue.
Shared Signals as the Real Containment Layer
Connect the dots. An isolated transaction tells few about the origin and destination of the funds. This is common knowledge. The more connected the information, the better for decision making, and in this scenario, the better for spotting rogue behaviour. In theory, Data Consortiums provide enough signals and events to preempt, quickly respond, and recover from an attack. This is true for most of the domains. In practice, Data Consortiums are constrained to specific use cases where information is open or the value of having fragmented information is very low. On the contrary, in those domains where information is highly valuable, and imperfect and incomplete information provides a competitive advantage, data sharing schemes struggle to survive. Luckily, innovative solutions come from cross domaines. Think about the dynamics of a cyclist peloton cooperating at the core, and competing at the edge. Can this scheme be replicated in Agentic Commerce?
The Core would be where collective intelligence is created, and the Edge where strategic divergence happens. To make it happen, the edge retains the ownership of the data and has complete and perfect information, while the Core retains collective intelligence capabilities built on top of incomplete and imperfect information. Imagine the edge anonymously sending anonymized, or aggregated information to the core, such as “an Agent accessing a loyalty account from an unusual context”, that information being available for the whole consortium. The edge processes that intelligence, and connects it to its complete and perfect information signals and events scheme. High level, this would be how the adoption of Data Consortium might help to deal with Agents behaving roguely while executing a Human Not Present transaction.