Data quality

Spotting duplicates

Recognise records that describe the same thing twice, even when they are not written the same way.

A duplicate is the same customer, order or product recorded more than once. Exact copies are easy to find; the hard cases differ by a capital letter, an extra space, an abbreviation or a typo in an email address.

Duplicates inflate counts, distort averages and send the same letter twice to the same customer. In a dataset used to train or evaluate a model, they give some cases too much weight and can make the model look better than it is.

Spotting them starts with choosing what identifies a record (an email, a customer number, a combination of fields), then normalising values before comparing them. When two rows conflict, decide which one to keep rather than deleting one at random.

In the reports of your games, every answer linked to this skill counts: the skill is shown as acquired from 80% of correct answers, and in progress from 50%.

Games that train this skill

Play them for free, without an account, then adapt them for your teams.

See all the skills