What is data literacy? Definition, skills and programme

Data literacy: a working definition, the skills that matter day to day, data quality in practice and a sample training programme for your teams.

Everyone4 min read

Data literacy is the ability to read, understand, use and question data in your work. It does not require programming. It does require knowing whether a figure can be trusted, what it really measures, and when to be wary of a table before basing a decision on it.

Why data literacy concerns everyone

Data is no longer one team's business. A salesperson feeds a CRM, an HR manager exports a headcount spreadsheet, a buyer compares suppliers in a workbook, a manager reads a dashboard every Monday. At each step, a simple mistake (a duplicated row, an empty cell, a date in the wrong format) can travel all the way to a decision.

AI raises the stakes: a model learns from the data it is given. Feed it dirty data and it produces wrong results, presented with the same confidence as correct ones. Data literacy is therefore a prerequisite for AI literacy.

Which skills make up data literacy?

They fall into four families:

Family What the person can do
Read understand what an indicator measures, its period, unit and source; read a chart without being misled by its scale
Check spot questionable data: duplicates, missing values, inconsistent formats, impossible values
Reason tell correlation from causation, distrust small samples, compare like with like
Communicate present a figure with its context and limits, protect the personal data you share

Data quality: the most practical foundation

The "check" family is the easiest to train, and the one whose mistakes cost the most, the fastest. Four problems come up again and again.

Duplicates

The same person, order or invoice appears twice, sometimes with different capitalisation or spelling. The typical result: inflated revenue or headcount. The habit: check identifiers before adding things up.

Missing values

An empty cell is not a zero. A missing age read as 0, a blank amount counted as nothing: the average drifts and nobody notices. The habit: decide explicitly whether to fill the value or set the row aside.

Inconsistent formats

Dates as DD/MM/YYYY and MM/DD/YYYY in the same column, amounts in euros and in thousands of euros, different decimal separators: the calculations still run, but they are wrong. The habit: standardise before you analyse.

Outliers

An impossible value (an age of 250, a negative temperature inside an oven) must be flagged. A rare but legitimate value (a freezer at -21 °C) must be kept. The habit: ask whether the value is possible, not just whether it is unusual.

These four habits are exactly what Data Cleaning Rush trains: rows of data move along a conveyor belt towards a model, and the player picks the right action before they arrive. Every mistake shows, in one sentence, what the model learned wrong.

What does a sample programme look like?

Data literacy programmes that work usually follow this pattern:

  1. Start from your own data. Generic examples convince less than the spreadsheets your teams actually handle. A forked game can reuse columns and cases close to yours, with fictitious data.
  2. Begin with data quality, because it is visible, fixable and measurable.
  3. Add critical reading of indicators: period, sample, source.
  4. Connect to AI: once the quality habit is in place, show what dirty data does to a model, then train people to check AI answers.
  5. Measure by skill: duplicates, missing values, formats and outliers, so you know where to focus next.

Data leaders will find guidance on running this kind of programme on the page for data and AI leaders.

Data literacy and data protection

Working with data often means working with personal data. A data literacy programme restates the basics: share only what is needed, prefer pseudonymised or aggregated data, never paste a file of named records into an unapproved external tool. These terms are defined in the glossary.

Get started

Play a game of Data Cleaning Rush from the catalogue, no account needed, then read data and AI literacy at work to build the rest of your programme.

Edit this page on GitHub (opens in a new tab)