YOU DO NOT COPY PROFESSIONALISM. YOU ALIGN WITH IT.
HOME / SERVICES / AI SECURITY / OWASP TOP 10

ASI09 - HUMAN-AGENT TRUST EXPLOITATION

A clerk in accounts asks the internal assistant to summarise the morning’s supplier correspondence. It returns a clean summary with a flagged action: one banking detail has changed, confirmed against an attached remittance. The clerk acts. Nothing about the interaction felt unusual, because nothing was unusual. This is ASI09 on the OWASP Top 10 for Agentic Applications, published 9 December 2025. AI agent impersonation fraud in South Africa now runs in both directions, against your staff through the assistants they trust and against your customers through agents that impersonate you, and POPIA applies to both. The control that fails is not technical. It is the trust relationship you built on purpose.

YOUR AWARENESS TRAINING TAUGHT THEM TO DISTRUST THE UNEXPECTED. THE AGENT IS EXPECTED.

Bank support chat where a verified-looking AI assistant requests an ID number and One-Time PIN, flagged as a spoofed agent identity.
DEFINITION

WHAT THIS RISK ACTUALLY IS

ASI09 covers the abuse of the confidence humans place in agents, and the confidence they place in organisations because an agent said so. Two directions, one root cause.

INWARD: Your staff use an internal assistant. It sits inside your tenant, wears your branding, uses your tone, and speaks with the authority of your systems. It reads mail, tickets, documents and chat in order to be useful. Every one of those is an untrusted input channel that an outsider can write to. When an attacker plants content that the assistant later summarises, the attacker has borrowed your institutional voice. The employee is not being careless. They are doing exactly what the tool was deployed for.

Awareness training is built on a single heuristic: be suspicious of the unexpected. An unexpected email, an unexpected caller, an unexpected invoice. The agent is expected. It is on the desktop, it was announced by the executive team, and it answers all day. There is no unexpectedness to detect, so the heuristic never fires.

OUTWARD: The same capability points at your customers. A convincing agent that sounds like your service desk, references real account context and holds a natural conversation is now inexpensive to run. The customer is not deciding whether to trust a stranger. They are deciding whether to trust you, and the agent has already answered that for them.

The South African fraud picture makes this concrete. SABRIC reported digital banking crime losses of R2.4 billion in 2025, up from approximately R1.9 billion in 2024, with banking app fraud accounting for over 70% of the total. That is a market where high-volume, high-conviction social engineering already pays. In a global survey of 713 respondents, the ACFE found 44% of anti-fraud professionals reporting a significant increase in deepfake social engineering over the past two years. That is a global figure and not South African banking data, but it describes the tooling now arriving in a market that is already losing R2.4 billion a year.

Preparedness has not kept pace. 7% of organisations report being more than moderately prepared to detect or prevent AI-powered fraud. South African banks have publicly announced agentic AI deployments in production, so customers here are already being trained to accept that a bank agent may be a machine.

DOCUMENTED CASE

WHAT IT LOOKS LIKE IN PRACTICE

The mechanism is documented. EchoLeak, CVE-2025-32711, was a zero-click flaw in Microsoft 365 Copilot. A hidden instruction sat inside an ordinary inbound email. The assistant read that instruction while assembling context, and tenant data left the organisation with nobody clicking anything. It was reported in January 2025, fixed server-side in May 2025, and found by a third party rather than by the platform vendor.

WHAT THIS MEANS UNDER SOUTH AFRICAN LAW

DISCOVERY

NEWORDER connects to CI/CD pipelines to automatically discover and inventory every homegrown AI application, and integrates directly with AWS Bedrock, Google Vertex AI, Salesforce, and other cloud and third-party platforms for visibility into AI agents. Each AI system is profiled across its model, system prompt, tools, guardrails, policies, and configurations, and the inventory stays current on every change. You cannot secure what you cannot see; discovery is the non-negotiable first step.

AI SECURITY POSTURE MANAGEMENT (AI-SPM)

NEWORDER conducts a static analysis of every application’s configuration, policy coverage, and third-party dependencies and identifies any policy gaps. In addition, it maps each agentic application to its coverage of major frameworks, including NIST, OWASP, and MITRE. This gives you a clear, measurable view of your AI security posture before a single adversarial test is run, turning assumptions into evidence and compliance into a continuous output rather than a periodic exercise.

AI RED TEAMING

NEWORDER’s automated AI red teaming covers the complete kill chain from reconnaissance to exploitation. It proactively discovers exploitable vulnerabilities through automated reconnaissance and adversarial testing purpose-built for agentic applications. Static attacks draw from a 300K+ payload library with 100% MITRE and OWASP LLM and Agentic Top 10 coverage, running comprehensive sweeps of known jailbreak patterns, content moderation bypasses, and obfuscation techniques. Dynamic attacks use multi-turn and continuous probing to test how an application holds up across extended adversarial sequences, not just a single interaction. High-agency attacks deploy extremely customised, bespoke attack techniques through probing tailored specifically to the intent and design of each application.

RUNTIME PROTECTION

NEWORDER offers policy enforcement and AI threat protection at the proxy, API, or AI Gateway layer. Protection adapts as the applications evolve and as new capabilities are added. When an attack hits production, whether a jailbreak, a prompt injection, or any other AI threat, it is blocked in real time and an immediate alert is sent with full context, including what happened, which application was targeted, what the impact is, and what to do next. Key performance metrics include 98.6% threat detection accuracy, 1.4% false positive rate, sub-200ms time to detect, sub-50ms real-time blocking, and immediate mean time to respond.

POPIA section 71

restricts decisions based solely on automated processing that have legal consequences for a person or substantially affect them. When an assistant’s output is the operative basis for a payment, a credit decision or an account action, the human in the loop is confirming rather than deciding.

If your control narrative depends on a person reviewing the agent’s output, and that person’s practical means of verification is the agent, you have a solely automated decision with a signature attached.

POPIA sections 19 to 22

require security safeguards, govern operator obligations and set notification duties. Notification of a compromise goes to the Information Regulator and to affected data subjects.

An assistant that ingests attacker-controlled content and acts on it is a failure of safeguards over personal information, whichever vendor supplied the assistant. The operator relationship does not move the obligation off you.

The Cybercrimes Act 19 of 2020, section 2

, makes unlawful access an offence. An attacker who uses an agent to obtain a customer’s credentials and then enters your systems commits that offence.

Prosecution depends on evidence. Arkose Labs found in February 2026 that 26% of enterprises are very confident they could prove an AI agent was involved in an incident, in a survey with no African respondents. If you cannot evidence agent involvement, the criminal route and the dispute route both weaken, and the customer carries the argument alone.

King V

, effective for financial years beginning on or after 1 January 2026, holds the governing body accountable for the effective, compliant and ethical use of technology, with demonstrable accountability for decisions, actions, outputs and outcomes, and human oversight and override mechanisms proportionate to risk.

Human oversight is not satisfied by a person sitting next to an agent. It requires that the person can independently check what the agent asserts. Design that verification path or you do not have oversight, you have proximity.

Joint Standard 2 of 2024

, in force 1 June 2025 for banks, insurers, asset managers, retirement funds and credit rating agencies, requires documented evidence of control testing including simulated incidents, a maintained testing calendar, a board-approved cyber risk charter, and notification of material incidents to the FSCA or Prudential Authority potentially within 24 hours.

A phishing simulation does not test this. It tests whether staff distrust the unexpected. ASI09 exploits what staff expect. The simulated incident that belongs on the calendar is an authorised trust exploitation exercise against the assistant itself, and against the customer channel.

The SARB, FSCA and Prudential Authority joint report of 24 November 2025 named explainability and board-level oversight as supervisory direction. An agent that cannot show a customer or an employee where an assertion came from is the explainability problem in its most expensive form.

South Africa has no dedicated AI legislation. The National AI Policy was gazetted in April 2026 and withdrawn on 26 April 2026 after fabricated citations were found in its reference list, with a revised draft targeted for January 2027. Nothing waits for it.

QUESTIONS TO ASK BEFORE YOUR NEXT AGENT GOES LIVE

  • Which external parties can place content into the sources our internal assistant reads, and does anyone review that path?
  • When the assistant states something as fact, can an employee see the source in one step, without asking the assistant?
  • What does our assistant do that an impersonator cannot copy, and would a customer notice its absence?
  • Has our awareness programme been updated for an attack that arrives through a tool we told staff to trust?
  • If a customer disputes an interaction next month, what would our logs prove about whether an agent was involved?
  • Is an authorised agent trust exploitation exercise on our testing calendar, separate from our phishing simulation?
FREQUENTLY ASKED QUESTIONS

FAQ

Awareness training teaches people to distrust the unexpected: the unexpected email, caller or invoice. An internal AI assistant is expected. It was announced by leadership, it sits on the desktop, and staff are instructed to use it. The suspicion heuristic never fires because nothing is out of place. The exposure is the trust relationship the organisation designed, so the fix is design and testing rather than more training.

Yes. POPIA sections 19 to 22 impose security safeguards, operator obligations and notification duties, and notification of a compromise goes to the Information Regulator and to affected data subjects. Section 71 applies where an outcome affecting a person rests solely on automated processing. Separately, the Cybercrimes Act 19 of 2020, section 2, covers unlawful access by an attacker who uses harvested credentials.

Session-level logs tying the interaction to an identity, the agent’s inputs and outputs, the tool calls it made, and provenance for any assertion it presented as fact. Arkose Labs reported in February 2026 that 26% of enterprises are very confident they could prove agent involvement, in a survey with no African respondents. Establish that capability before a dispute, not during one.

WHERE TO GO FROM HERE

TAKE ACTION

TEST IT BEFORE SOMEBODY ELSE DOES

Thirty minutes on your agent estate and what this risk looks like in your environment.