Automation

How to remove duplicate data: a checklist for doing it thoughtfully

Two records for the same customer, a repeated order or a supplier entered with slight variations may seem like a minor issue. Until someone calls the same person twice, consults outdated information or makes a decision based on a distorted report.

This checklist helps you decide what to review before removing duplicate data, how to do so without losing relevant information and when it makes sense to turn the clean-up into a rule or automation. Use it before a migration, when merging spreadsheets, when organising a CRM or when you identify inconsistencies in your records.

A duplicate is not always an exact copy. It may be the same company written in two different ways, a person with two email addresses or the same order entered through different channels. That is why deleting records without reviewing the context can create a new problem.

Review objective

Before changing any data, define what you want to correct. The aim is not to reduce the number of rows or records: it is to keep one reliable record for every customer, contact, product, order or entity you manage.

Tick off this initial check:

  • I have defined the type of data I am going to review: contacts, customers, suppliers, orders, products or another type.
  • I know the impact of the duplicate on my operations: repeated communications, follow-up errors, unclear reports or manual tasks.
  • I have identified where the data is held: one spreadsheet, one application, several connected tools or imported files.
  • I have decided which data point best identifies each record, such as an email address, document number, internal code or a combination of fields.
  • I have established who validates uncertain cases before information is deleted or merged.

If you cannot identify which data point reliably identifies each record, do not start deleting. You will first need to define an identification rule and check whether that data is complete and up to date.

Pre-clean-up checklist

Preparation prevents an apparently simple clean-up from deleting notes, history or useful information that appears in only one of the copies.

  • I have created a copy of the file or exported the records before modifying them.
  • I have recorded the review date and the source of the data I am going to compare.
  • I have separated records that appear to be duplicates from those that are clearly different.
  • I have reviewed whether any fields use inconsistent formats: capitalisation, accents, spaces, telephone prefixes, abbreviations or incomplete addresses.
  • I have checked whether two similar records could belong to different people, for example different contacts at the same company.
  • I have decided which fields must be retained even if the main record is merged or deleted.

A typical case is finding two records with the same first and last name but different email addresses. You should not assume that one is unnecessary: it could be the same person using an old address, two people with the same name or a shared account. The available information and how it is used should guide the decision.

Operational checklist

This section helps you decide which record to keep and how to consolidate data without turning the clean-up into a loss of context.

  • I have selected a primary record for each group of duplicates.
  • I have compared the relevant fields before deleting a copy: name, contact details, address, status, date, notes and history.
  • I have transferred valid information that appears only in the duplicate to the primary record.
  • I have resolved conflicts between different data instead of retaining one automatically.
  • I have checked whether the duplicate record is linked to orders, invoices, tasks, documents or other items.
  • I have left an internal note when the merge requires a decision that another person may need to understand later.
  • I have deleted or archived the copy only after confirming that the primary record contains the necessary information.

In operational data, the date does not always resolve a conflict. A more recent telephone number may be correct, but an older note may explain an issue that is still open. Keep what helps the work continue, not merely the value that appears to be newer.

It is also useful to distinguish between merging, archiving and deleting:

  • Merging brings together valid information from several records in one primary record.
  • Archiving removes a record from day-to-day operations when it may still be needed for reference.
  • Deleting removes a record when you have confirmed that it does not provide necessary information, relationships or history.

Technical checklist

Manual identification may be sufficient for a small, stable list. When you receive data from forms, emails, spreadsheets or multiple applications, you need consistent rules so that the result does not depend on who carries out the review.

  • I have defined the fields that will be used to identify exact matches.
  • I have established which variations will be normalised before comparison, such as extra spaces, capitalisation or telephone formats.
  • I have separated exact matches from likely matches that require human review.
  • I have checked which tool creates each record and which one should serve as the primary source.
  • I have reviewed what happens when two tools update the same data differently.
  • I have tested the rule on a sample before applying it to the full dataset.
  • I have verified that the clean-up does not break links to documents, orders, tasks or subsequent processes.

Normalisation means converting equivalent data into a common format before comparing it. For example, if a telephone number appears with spaces in one record and without them in another, a literal comparison may treat them as different. Normalisation improves detection, but it does not replace the review of ambiguous cases.

Avoid an overly broad rule, such as considering all records that share a name to be duplicates. Instead, combine data points that make sense for your case: name and email address, customer code and company, or order number and date. The right combination depends on what you are managing.

When duplicates are generated as information moves between tools, the solution may require an entry rule, prior validation or more clearly defined synchronisation. AI automations for businesses can help organise that workflow when manual review is no longer sustainable, provided that exceptions and validation criteria are clearly defined.

Maintenance checklist

Removing duplicate data once corrects the historical record; preventing it reduces the work that will reappear. Apply this section once the initial clean-up has been validated.

  • I have defined which fields are mandatory when creating a new record.
  • I have specified a single format for email addresses, telephone numbers, codes and company names where applicable.
  • I have added a match check before creating a new record.
  • I have assigned a primary source when the same data arrives from several tools.
  • I have defined who can create, merge, edit or delete records.
  • I have established a periodic review of identified duplicates.
  • I have documented what to do with uncertain matches rather than deleting them automatically.
  • I have reviewed the forms and imports that may be creating repeated entries.

Automation should not make decisions on its own in cases with commercial, administrative or customer relationship consequences. It can identify matches, propose an action or route the case to a responsible person. The final decision must remain under human control when the information is ambiguous or sensitive.

How to interpret the result

Once you have completed the checklist, classify your situation according to the type of issue you have identified:

You can carry out a controlled clean-up if you have a clear identifier, there are few duplicates and you can compare their data before merging or deleting them. Keep a prior copy and validate the result when finished.

You need to organise your working rules if duplicates result from different formats, incomplete records or differing criteria among people. In that case, first define the mandatory fields, the input format and who resolves uncertainties.

You need to review the workflow between tools if the same contact, order or data point is repeated when information is imported, synchronised or copied between systems. Deleting the copies will relieve the symptom, but it will not prevent them from being generated again.

Do not delete anything yet if you do not know which record is correct, if there is conflicting data or if a record is linked to documents and operations you have not reviewed. Flag these cases for human validation.

If you have identified that duplicates originate from a repetitive process involving forms, emails, spreadsheets or applications, you can tell AVSISTEC about your automation project. The starting point will be to identify which data should take precedence, which exceptions require review and which change can be sustained over time.