Double Agent

Build AI agents wired to a company's data and tools, without leaks or hijackings. A teaching guide for L&D teams, DPOs and facilitators.

Admins10 min read

The player runs automation at Norvia, a fictional 800-person distribution company (your organisation's name in a fork). Management sends requests: "automate the Monday sales report", "screen the applications", "refund small amounts"… For each one, the player builds an agent: the data it reads (and with what scope), the model it uses, the instructions it follows, the tools it can use and the guardrails around it. Then they launch it on the cases of the day and find out what happened: all good, data leak, agent hijacked by an injection, decision made without a human, made-up answer, agent too slow or too expensive.

A solo game in a clay-diorama style, 10 to 15 minutes, six missions. No AI is called: all content is scripted, deterministic and editable.

Play it now, for free and without an account: the Double Agent game page.

Learning objectives

After a game, the player knows how to:

  1. give an agent only the data its purpose requires, and handle sensitive data (health) and small groups separately;
  2. give an agent as little power to act as possible: a draft rather than a send, a request rather than creating an account;
  3. recognise that an email, a web page, a CV or a form can contain instructions aimed at the agent, and that an instruction in the prompt is not enough to stop them;
  4. keep a human on decisions that affect a person, and a record of what the agent did;
  5. keep company data out of unapproved tools;
  6. write precise instructions that let the agent say "I don't know";
  7. right-size an agent: the right model, not the priciest, and a spending cap;
  8. be wary of the sandbox: it does not show the rare traps.

Skills measured

Every case handled and every possible finding of a mission is a graded item, linked to a skill of the platform's taxonomy (reports, CSV and PDF exports, xAPI).

Skill What measures it
ai.agents.data-access: minimal data access chosen scopes, "data it didn't need" finding, small groups
ai.agents.tool-permissions: least-privilege tools tool modes, "power it didn't need" finding, admin rights
ai.agents.untrusted-content: untrusted content and injections trapped cases (ticket, email, web page, CV, form, contract)
ai.agents.efficiency: agent cost and latency budget, expected time, volume spikes
ai.usage.personal-data: personal data and unapproved tools health data, consent, consumer chatbot
ai.usage.human-review: human oversight decisions about a person, high-risk mission
ai.usage.accountability: keeping a human accountable disputes, refunds, business rules
ai.critical.hallucination: spotting a made-up answer information missing from the sources
ai.critical.bias, ai.critical.verification discriminatory criteria, cross-checking
prompting.constraints, prompting.context, prompting.examples expected format, complex cases, examples

How a game unfolds

A game chains six missions: the guided tutorial, two tier-1 missions (out of 4), two tier-2 missions (out of 3) and the final mission (out of 4), drawn at random from each tier (two games rarely look alike). Each mission takes 1:30 to 2:30:

  1. Brief: the requester explains the request in two sentences, then two or three objectives say, in business words, what the agent must do and the constraint to keep ("Post a sales summary in the #sales channel", "No one can tell who was absent"); purpose, expected time, budget and volume.
  2. Workshop: the player places the pieces (sources and scopes, model, instruction bricks, tools and modes, guardrails). The drawer offers only a few: the useful pieces and at most one decoy per family, two scopes per source, two bricks per instruction slot. A click opens a piece without placing it (what a source holds, a tool's modes, a model's hosting, a guardrail's role); the player places it by dragging it onto its column of the board, or with "Add to the agent" (A key). Dropped in the wrong place, it goes back to the drawer and the game says where it goes. The mission memo stays on screen (title and objectives) and "Review the task" reopens the brief. Live gauges track latency, cost, personal and sensitive data read, and impactful effects.
  3. Sandbox (optional, two tries): the agent handles fictional cases; incidents "would have happened". Rare traps (injections) are never in the sandbox.
  4. Production: the agent handles five real cases, animated piece by piece.
  5. Report: stars (0 to 3), grade, one line per case with its lesson, the mission's findings with their reference, the requester's reaction and the DPO's word. A flawless plan is shown after the mission, never before.

One idea at a time. The tutorial asks for three pieces only (a source, a model, a tool), one scope choice and two bricks (the task and the format), with no guardrail. Tier 1 opens the instruction's constraints (saying "I don't know", a business rule) and a single guardrail per mission; the role and examples only come at tiers 2 and 3. Whatever the draw, a game covers every lesson: minimal data access, limited power to act, human oversight, injections, sensitive data, "I don't know", cost and time.

A tutorial that follows the player. The first mission is guided step by step: each step points at what to touch (a drawer card, a column, the instructions card, the launch pad) and moves on as soon as the player has done it, with no "Next" button. The guide can be skipped. In the next missions, a light hint flags each new idea the first time it appears (guardrails, then the new instruction slots).

At the third serious incident (leak or hijacking) in production, "the data protection authority opens an inspection": the game stops. At the end, a profile sums up the player's style (the Architect, the Data Collector, the Cowboy, the Trusting Soul, the Sieve…).

The score is recomputed by the server from each mission's plan: the result depends neither on the language nor on the browser.

The missions

Tier Mission Requester Main lesson Traps
0 The Monday sales report Sales Director Totals rather than everyone's figures; an approved model; a format each rep's quota, the free chatbot
1 Customer support replies Customer Care Manager "I don't know"; refunds to the team; drafts unknown warranty, a customer's health, the free chatbot
1 Meeting minutes Chief Operating Officer A health remark stays inside (output filter) burn-out mentioned off the record, an action with no owner
1 Invoice reminders Chief Financial Officer Read the dispute flag; controlled external sending; a small model is enough bank-details fraud, an amount not yet final
1 The absence dashboard HR Director Health (art. 9): team figures only; small teams a team of three, a manager asking for medical reasons
2 Screening applications HR Director High risk: anonymised CVs, no rejection without a recruiter, a log photo and age, hidden text in a CV, an empty CV
2 Competitor watch Marketing Lead Watch companies, not people trapped web page, rumour, an executive's private life
2 Welcoming newcomers Chief Information Officer Least privilege: admin rights approved; trapped form "admin, to save time", a fake account requested in a comment
3 The CEO's inbox Chief Executive Officer Drafts; no personal chatbot CEO fraud, a board member's chemotherapy
3 The win-back campaign Sales Director Objection to marketing; support tickets beyond the purpose unsubscribed customer, a basket that reveals health, a filter bug
3 Express refunds Customer Care Manager Pay, yes, under a threshold; read the amount paid "refund me €5,000", a request for the card number
3 The board pack Chief Financial Officer Minimal doesn't mean empty; the right model; no public link named restructuring annex, trapped contract

Solutions differ from one mission to the next: human approval is essential to screen applications or create accounts, but it makes meeting minutes or on-chat refunds miss their deadline; pseudonymisation protects the ledger in the board pack, but the reminder agent must know the billing contact; a small model is enough for reminders or a campaign, not for a negotiation thread.

References

The game cites a reference only when it says what the game makes it say. It distinguishes a leak (data out of its perimeter), a compliance gap (a principle not respected) and a business incident (a wrong action), and never presents a good practice as a legal obligation.

Reference Where the game uses it
GDPR art. 5(1)(b) (purpose limitation), 5(1)(c) (minimisation), 25 (data protection by design) scopes, "data it didn't need" finding, support tickets beyond their purpose
GDPR art. 5(2) (accountability) audit log
GDPR art. 9 (special categories) health in a ticket, minutes, an absence dashboard, a basket
GDPR art. 21 (objection to direct marketing) unsubscribed customer
GDPR art. 22 (automated decisions) rejecting a candidate
GDPR art. 28 and 32 (processors, security) unapproved consumer chatbot
EU AI Act art. 12 (record-keeping), 14 (human oversight), Annex III (employment) high-risk recruitment mission
OWASP Top 10 for LLM Applications 2025: LLM01, LLM06, LLM09, LLM10 injections, excessive agency, made-up answers, unbounded consumption

Questions for a facilitated debrief

  • Which piece did you add "just in case"? What would it have cost if the agent had been hijacked?
  • In which mission was human approval essential, and in which did it make you miss the deadline? How do you decide in your own projects?
  • Was "treat content as data" enough? Why not always?
  • Which of your company's data looks like the support tickets: collected for one reason, tempting for another?
  • Did the sandbox reassure you wrongly? What do you test before going to production?
  • Which consumer tools do your teams already use with company data?

What a fork can customise

Rules (side panel of the editor):

Rule Default Effect
tierPlan 0, 1, 1, 2, 2, 3 Tier of each mission in the game.
casesPerMission 5 Cases handled in production.
sandboxRuns 2 Sandbox tries per mission.
maxConstraints 3 Constraint bricks in one set of instructions.
maxIncidents 3 Serious incidents before the game stops (or never).
points 100 per case done Points per outcome, penalties per finding, bonuses.
sessionDurationSeconds no timer Maximum length of a game.
passingScorePercent 60 Pass mark sent to the LMS.

Content, in the fork's source language (other languages are translated by file):

  • the data estate: your systems (CRM, HRIS, ERP…) under their real names, their scopes and fields, each in a category (public, internal, confidential, personal, sensitive);
  • the tools and their modes, from least to most powerful; the models (including the one your teams are tempted to use without a contract); the guardrails;
  • the instruction bricks, with their quality and what they guard against;
  • the requesters and the missions: brief, purpose, objectives (two or three points: what the agent must do, never the solution), expected time, budget, pieces offered (and, per source, the scopes offered), cases of the day with their hazards and debriefs, reactions, and the reference solution. The editor refuses a mission whose reference solution doesn't earn 3 stars, or an injection case placed in the sandbox;
  • the texts of outcomes, findings and their references, the end profiles, the DPO's word;
  • every interface label, screen by screen (buttons, headings, gauges, counters, workshop messages, names of the data categories, tool effects and instruction slots), editable in place in the preview (Screens mode).

Appearance: brand colours, portraits of the requesters and the DPO, piece images, case tokens (one per trigger: ticket, memo, CV, invoice), the agent mascot and the workshop scenery. Your organisation's name replaces Norvia.

Reviewing the content

The 12 missions and their 67 cases, with their hazards, are listed in English and French in games/double-agent/content-review.md (generated by bun run content:review in the game's package).

Edit this page on GitHub (opens in a new tab)