Data Cleaning Rush

The arcade game where players clean data rows before they reach the model. Rules, skills and fork customization.

Editors4 min read

A data factory at night. Data rows (customer records, orders, sensor readings, job applications) move along a conveyor belt toward "the model". The player picks the right action for each row before it gets there. Dirty data teaches the model the wrong things: every mistake shows, in one line, what the model learned.

Solo session of 5 to 10 minutes, available in English and French.

Play it now, for free and without an account: the Data Cleaning Rush game page.

How it plays

  • 4 rounds, 37 items by default. One item is one data row.
  • Each round opens with a short brief (expected formats, context).
  • For each row, the player:
    • lets it through when it is clean, including a rare but legitimate value;
    • deletes the duplicate of a row already sent to the model;
    • handles a missing value (fill it in or set it aside);
    • fixes a format (date, unit, case, separator);
    • flags an outlier (impossible value).
  • Fixes add up: a row can carry several issues. "Send" goes with the chosen fixes; deleting a duplicate decides at once.
  • A row is right when the chosen actions are exactly the expected ones: nothing missed, nothing overdone.
  • Difficulty grows: one issue per row at first, then formats, outliers and rare legitimate values, and finally several issues per row and near-duplicates.

The "Already in the model" table keeps the last rows sent: that is where duplicates show up.

Feedback

  • The model quality gauge rises with each well-handled row and drops faster with each dirty one.
  • After each decision, the row shows the faulty cells, the fixed value and the right action. A mistake shows its concrete consequence, for example "Blank age read as 0: the model thinks babies are shopping".

Accessibility

  • Keyboard: 2 to 5 to pick an action, 1 or Enter to send and continue.
  • Without a timer (sessionDurationSeconds: null) or with the system's reduced motion, the belt waits for the player, one row at a time.
  • All text is readable HTML, feedback is announced to screen readers, and a refused action plays a sound.

Skills measured

Skill What is measured
data.quality.duplicates Spot a duplicate, even with different casing, without deleting a unique row
data.quality.missing See a missing value and decide to fill it in or set it aside
data.quality.formats Harmonize dates, units, casing and separators
data.quality.outliers Flag an impossible value, and let a rare but legitimate one through

Each row declares its skill. A clean row measures a skill too: letting a freezer at -21 °C through measures data.quality.outliers.

Score

  • A right row earns pointsPerCorrectRow points, a wrong row loses wrongRowPenalty. The score stays between 0 and the maximum.
  • A row that reaches the model undecided goes with the fixes already chosen.
  • The server recomputes the score from the raw answers and the release config.

Customize a fork

Rules

Rule Default Effect
sessionDurationSeconds 600 Session length; null = no timer, belt at the player's pace
baseRowSeconds 16 Seconds for a row to cross the belt in round 1
speedUpPercentPerRound 15 Belt speed-up at each round, in %
rowsPerRound 12 Maximum rows played per round (the first ones)
pointsPerCorrectRow 100 Points for a right row
wrongRowPenalty 30 Points lost for a wrong row
enabledIssues all 4 Issue kinds in play: duplicate, missing, format, outlier

Disabling an issue kind removes the rows that test it, and its button disappears.

Content

Every displayed text is editable, in the fork's source language (the other languages are translated by file):

  • the labels (title, introduction, actions and their hints, feedback, end screen);
  • each round: title, brief and columns;
  • each row: its cells, its issues (faulty cell, fixed value) and its consequence.

To add your own data, keep one cell per column, put the original of a duplicate within the 6 rows before it, and write a short, concrete consequence. Use fictional data.

Look

  • Brand color (brandPrimary) and highlight color (brandSecondary).
  • Factory background (16:9 image) and model image.

Edit this page on GitHub (opens in a new tab)