Why does a recruitment database fill up with duplicate candidates?
Because the same person arrives through several doors. They apply to one advert, get sourced from LinkedIn by a colleague, are imported in a bulk upload, and are added by hand after a call. Unless each of those routes checks against the same rule, you end up with four records for one person, and the history is split across all four.
Duplicates are usually described as a data quality problem, which understates them. The cost is not untidiness, it is the split history. One record holds the note from the call, another holds the CV, a third holds the fact they were rejected by a client last year. No consultant sees the whole person, and the system cannot rank someone it only half knows.
The worst version of this is commercial. Two consultants submit the same candidate to the same client because each is looking at a different record. That is a conversation with a client no agency wants to have.
Matching cannot be done on name alone. Common names collide constantly, and a rule that merges on name will eventually merge two different people, which is worse than leaving duplicates. Reliable matching uses email, a LinkedIn profile, and phone number, and treats name similarity as a hint rather than a decision.
The three things worth having are prevention at every entry point, detection for the duplicates already there, and a guard on submission that recognises the same person across records even when the records have not been merged yet. The third is the one that stops the client conversation.
Where do duplicate candidate records come from?
Every route into the database is a chance to create a second copy of someone. The fix is a check at each route, and the check that works differs by route.
| Entry point | Why it creates duplicates | The check that helps |
|---|---|---|
| Job board applications | Candidates apply with a different email each time | Match on phone and name before creating a record |
| LinkedIn or browser capture | Profiles often have no email at all | Match on the normalised profile URL |
| Bulk CV uploads | Hundreds of records created at once, unchecked | Run the same matching as single entry, per file |
| Manual entry after a call | Quicker to add than to search | Search-as-you-type on email and phone in the add form |
| Data migrations | Two databases merged, each holding the same people | Deduplicate during the load, not after |
| Referrals and events | Names typed from memory, spelling varies | Treat as a possible match and send to review |
How do you match duplicates reliably?
Rank the signals by how often they are unique to one person. An email address is the strongest, once lower-cased and trimmed. A mobile number is next, once formatted consistently so that 07700 900123 and +44 7700 900123 read as the same number. A LinkedIn profile URL is strong once the tracking parameters and trailing slashes are stripped.
Name, current employer and location are supporting evidence. They can raise confidence that two records with the same phone number are one person, but they should never create a match on their own, and two different people who share a name and an employer are not rare in a large database.
Should duplicates be merged automatically?
Only on the strongest evidence. An exact match on a normalised email is safe to merge without asking. Anything weaker belongs in a review queue where a person confirms it, because an incorrect merge is far worse than a duplicate: two people's CVs, notes and submission histories become one record, and separating them again is slow and error-prone.
A good merge keeps everything from both records: every note, file, application and submission, with their original dates and authors. Where fields conflict, keep the most recent value and retain the other in history rather than discarding it.
Are duplicate records a UK GDPR problem?
When a candidate makes a subject access request, you have one month to give them a copy of their data. If they exist as four records, a search that finds one produces an incomplete response. The same applies to erasure: deleting one record while three others survive means the request was not actually met.
Duplicates also blur the record of when and how you obtained someone's data, and when you told them you held it. Merging keeps that history in one place, which is the easiest way to be able to answer those questions accurately.
How do you stop two consultants submitting the same person?
Check at the point of submission, not only at the point of entry. Before a CV goes to a client, look for any record that shares an email, phone or profile URL with the candidate and has already been submitted to that client in the period your terms of business cover. If there is one, stop and show the consultant who submitted them and when.
Agree the internal rule in advance too: who owns a candidate who has been worked by two consultants, and for how long. Most disputes are not about the software but about the absence of a rule.
Where Vayora fits
Vayora checks for an existing person when candidates arrive through imports, bulk CV uploads and sourcing, matching on email and LinkedIn profile. Weaker matches are suggested for someone to confirm, a name match on its own never merges anyone, and every merge is a deliberate choice of which record survives. A submission is stopped when the same person has already gone to that role, or is still live with that client, even if the earlier submission sits on a different record for them, matched on email, phone or LinkedIn profile. It does not yet warn on the manual add-candidate form.
Common questions
- Can deduplication accidentally merge two different people?
- Yes, if it merges on weak evidence such as a name, or a name and a company. That is why automatic merging should be limited to exact matches on strong identifiers like a normalised email address, with everything else sent to a person to confirm. An incorrect merge mixes two people's histories and is much harder to undo than a duplicate is to fix.
- How often should we check for duplicate candidates?
- Continuously at the point of entry, so new duplicates are caught as they arrive, and with a periodic sweep of the existing database for those that slipped through or were created before the checks existed. Run an extra sweep after any bulk import or migration, since those are the moments most duplicates are created in one go.
- What should happen to consent and source dates when two records merge?
- Keep both. The merged record should show the earliest date you obtained the candidate's data and from where, when you gave them privacy information, and any objections or opt-outs from either record. An opt-out recorded on one duplicate must carry across; losing it in a merge means contacting someone who asked you to stop.