Hallucination Hunter

Sort an AI assistant's answers into reliable, hallucinated, biased or needs checking, and what a fork can customise.

Everyone3 min read

Your colleagues asked an AI assistant for help. Before anyone acts on its answers, the player checks each one from an investigation desk drawn in black and cream screenprint. A session lasts 5 to 10 minutes and is played solo.

Play it now, for free and without an account: the Hallucination Hunter game page.

Rules

Each item is one of the assistant's answers to a question from a team (HR, finance, legal, marketing, occupational health, data). The player stamps one of four verdicts:

Verdict When to choose it
Reliable The answer is correct and can be used as is.
Hallucinated An invented fact, a source that does not exist, a made-up figure.
Biased A stereotype, a generalisation, one point of view presented as neutral.
Needs checking Plausible, but impossible to confirm without a source, or it depends on the context.
  • On the keyboard, keys 1 to 4 stamp the matching verdict.
  • Two paid clues are available before the verdict: "Check the source" and "Compare". Each clue read deducts its cost from the item's points, never below zero.
  • After each verdict, the passage in question is highlighted in the answer and a one or two line explanation says why. The player sees where the problem was.
  • A correct verdict earns the item's points, minus the clues read. A wrong verdict earns nothing and costs nothing more.
  • A session has 3 themed rounds of 11 items, each subtler than the last:
    1. HR and occupational health: the trap shows if you read carefully;
    2. finance and marketing: figures to redo and sources to trace;
    3. legal and data: true and false mix in the same answer.
  • By default there is no timer: players read and think.

The displayed score is recomputed by the server from the verdicts and clues: that is the score used in the leaderboard and in reports.

Skills measured

Skill Items
ai.critical.hallucination: spotting a hallucination hallucinated answers
ai.critical.bias: detecting a bias biased answers
ai.critical.verification: knowing when to check reliable and needs checking answers

What a fork can customise

Rules (side panel of the editor):

Rule Default Effect
roundCount 3 Number of rounds played (the first ones in the config).
itemsPerRound 11 Items played per round (the first ones, in a shuffled order).
pointsPerCorrect 100 Points for a correct verdict.
hintCost 30 Cost of a clue; 0 makes clues free.
hintsEnabled yes Shows or hides clues.
verdicts all 4 Verdicts offered; disabling "Biased" also removes the biased items.
showExplanation yes Shows the explanation after each verdict.
sessionDurationSeconds no timer Maximum session length.
passingScorePercent 60 Pass mark sent to the LMS.

Content, editable in place in the editor, in the fork's source language (the other languages are translated by file):

  • each item: team, question, assistant's answer, passage in question, explanation and the text of both clues (an empty clue is hidden);
  • the title and brief of each round;
  • the name and definition of each verdict, and every interface label.

The passage in question must be copied verbatim from the answer, in every language (translations included): otherwise nothing is highlighted.

Look and feel: brand colour (stamps, buttons), highlight colour, and the four desk images (background, case file, stamp, magnifying glass).

Reviewing the content

The default content is listed item by item, with its verdict and explanation, in games/hallucination-hunter/content-review.md (generated by bun run content:review in the game package). "Reliable" statements rely on stable facts: French Labour Code, GDPR, statistical definitions.

Edit this page on GitHub (opens in a new tab)