Hallucination Hunter
Sort an AI assistant's answers into reliable, hallucinated, biased or needs checking, and what a fork can customise.
Your colleagues asked an AI assistant for help. Before anyone acts on its answers, the player checks each one from an investigation desk drawn in black and cream screenprint. A session lasts 5 to 10 minutes and is played solo.
Play it now, for free and without an account: the Hallucination Hunter game page.
Rules
Each item is one of the assistant's answers to a question from a team (HR, finance, legal, marketing, occupational health, data). The player stamps one of four verdicts:
| Verdict | When to choose it |
|---|---|
| Reliable | The answer is correct and can be used as is. |
| Hallucinated | An invented fact, a source that does not exist, a made-up figure. |
| Biased | A stereotype, a generalisation, one point of view presented as neutral. |
| Needs checking | Plausible, but impossible to confirm without a source, or it depends on the context. |
- On the keyboard, keys 1 to 4 stamp the matching verdict.
- Two paid clues are available before the verdict: "Check the source" and "Compare". Each clue read deducts its cost from the item's points, never below zero.
- After each verdict, the passage in question is highlighted in the answer and a one or two line explanation says why. The player sees where the problem was.
- A correct verdict earns the item's points, minus the clues read. A wrong verdict earns nothing and costs nothing more.
- A session has 3 themed rounds of 11 items, each subtler than the last:
- HR and occupational health: the trap shows if you read carefully;
- finance and marketing: figures to redo and sources to trace;
- legal and data: true and false mix in the same answer.
- By default there is no timer: players read and think.
The displayed score is recomputed by the server from the verdicts and clues: that is the score used in the leaderboard and in reports.
Skills measured
| Skill | Items |
|---|---|
ai.critical.hallucination: spotting a hallucination |
hallucinated answers |
ai.critical.bias: detecting a bias |
biased answers |
ai.critical.verification: knowing when to check |
reliable and needs checking answers |
What a fork can customise
Rules (side panel of the editor):
| Rule | Default | Effect |
|---|---|---|
roundCount |
3 | Number of rounds played (the first ones in the config). |
itemsPerRound |
11 | Items played per round (the first ones, in a shuffled order). |
pointsPerCorrect |
100 | Points for a correct verdict. |
hintCost |
30 | Cost of a clue; 0 makes clues free. |
hintsEnabled |
yes | Shows or hides clues. |
verdicts |
all 4 | Verdicts offered; disabling "Biased" also removes the biased items. |
showExplanation |
yes | Shows the explanation after each verdict. |
sessionDurationSeconds |
no timer | Maximum session length. |
passingScorePercent |
60 | Pass mark sent to the LMS. |
Content, editable in place in the editor, in the fork's source language (the other languages are translated by file):
- each item: team, question, assistant's answer, passage in question, explanation and the text of both clues (an empty clue is hidden);
- the title and brief of each round;
- the name and definition of each verdict, and every interface label.
The passage in question must be copied verbatim from the answer, in every language (translations included): otherwise nothing is highlighted.
Look and feel: brand colour (stamps, buttons), highlight colour, and the four desk images (background, case file, stamp, magnifying glass).
Reviewing the content
The default content is listed item by item, with its verdict and explanation, in
games/hallucination-hunter/content-review.md (generated by bun run content:review in the game
package). "Reliable" statements rely on stable facts: French Labour Code, GDPR, statistical
definitions.