Data Cleaning Rush
The arcade game where players clean data rows before they reach the model. Rules, skills and fork customization.
A data factory at night. Data rows (customer records, orders, sensor readings, job applications) move along a conveyor belt toward "the model". The player picks the right action for each row before it gets there. Dirty data teaches the model the wrong things: every mistake shows, in one line, what the model learned.
Solo session of 5 to 10 minutes, available in English and French.
Play it now, for free and without an account: the Data Cleaning Rush game page.
How it plays
- 4 rounds, 37 items by default. One item is one data row.
- Each round opens with a short brief (expected formats, context).
- For each row, the player:
- lets it through when it is clean, including a rare but legitimate value;
- deletes the duplicate of a row already sent to the model;
- handles a missing value (fill it in or set it aside);
- fixes a format (date, unit, case, separator);
- flags an outlier (impossible value).
- Fixes add up: a row can carry several issues. "Send" goes with the chosen fixes; deleting a duplicate decides at once.
- A row is right when the chosen actions are exactly the expected ones: nothing missed, nothing overdone.
- Difficulty grows: one issue per row at first, then formats, outliers and rare legitimate values, and finally several issues per row and near-duplicates.
The "Already in the model" table keeps the last rows sent: that is where duplicates show up.
Feedback
- The model quality gauge rises with each well-handled row and drops faster with each dirty one.
- After each decision, the row shows the faulty cells, the fixed value and the right action. A mistake shows its concrete consequence, for example "Blank age read as 0: the model thinks babies are shopping".
Accessibility
- Keyboard:
2to5to pick an action,1or Enter to send and continue. - Without a timer (
sessionDurationSeconds: null) or with the system's reduced motion, the belt waits for the player, one row at a time. - All text is readable HTML, feedback is announced to screen readers, and a refused action plays a sound.
Skills measured
| Skill | What is measured |
|---|---|
data.quality.duplicates |
Spot a duplicate, even with different casing, without deleting a unique row |
data.quality.missing |
See a missing value and decide to fill it in or set it aside |
data.quality.formats |
Harmonize dates, units, casing and separators |
data.quality.outliers |
Flag an impossible value, and let a rare but legitimate one through |
Each row declares its skill. A clean row measures a skill too: letting a freezer at -21 °C
through measures data.quality.outliers.
Score
- A right row earns
pointsPerCorrectRowpoints, a wrong row loseswrongRowPenalty. The score stays between 0 and the maximum. - A row that reaches the model undecided goes with the fixes already chosen.
- The server recomputes the score from the raw answers and the release config.
Customize a fork
Rules
| Rule | Default | Effect |
|---|---|---|
sessionDurationSeconds |
600 | Session length; null = no timer, belt at the player's pace |
baseRowSeconds |
16 | Seconds for a row to cross the belt in round 1 |
speedUpPercentPerRound |
15 | Belt speed-up at each round, in % |
rowsPerRound |
12 | Maximum rows played per round (the first ones) |
pointsPerCorrectRow |
100 | Points for a right row |
wrongRowPenalty |
30 | Points lost for a wrong row |
enabledIssues |
all 4 | Issue kinds in play: duplicate, missing, format, outlier |
Disabling an issue kind removes the rows that test it, and its button disappears.
Content
Every displayed text is editable, in the fork's source language (the other languages are translated by file):
- the labels (title, introduction, actions and their hints, feedback, end screen);
- each round: title, brief and columns;
- each row: its cells, its issues (faulty cell, fixed value) and its consequence.
To add your own data, keep one cell per column, put the original of a duplicate within the 6 rows before it, and write a short, concrete consequence. Use fictional data.
Look
- Brand color (
brandPrimary) and highlight color (brandSecondary). - Factory background (16:9 image) and model image.