Data quality
How to measure contact database quality
A practical checklist of quality measures — coverage, validity, uniqueness and consistency — and how to compute them yourself.
"High quality" means nothing unless it is measured. These are the measures that matter, how we report them, and how to check them yourself.
1. Coverage (completeness)
Percentage of rows where a field is filled in. Formula: non-empty values ÷ total rows. We compute this for every column of every pack automatically. Try it on any file with our CSV field analyzer.
2. Validity
Filled in is not the same as correct. Check format validity:
- Phone numbers parse to a valid number for the country (see phone normalisation).
- Emails match the basic pattern
name@domain.tld. - Postcodes match the national format.
3. Uniqueness
Exact duplicate rows are counted by our scanner and published per dataset. Near duplicates (same business, slightly different spelling) need normalised keys such as E.164 phone numbers or website domains.
4. Consistency
Look for mixed formats within a column (dates as 01/02/2026 and 2026-02-01, phone numbers with and without country codes). Inconsistency slows every import.
5. Accuracy
Only real-world checks can establish accuracy. Sample rows and verify them against primary sources. No seller can honestly promise 100% accuracy for business contact data, and we do not.
A quick scorecard
| Measure | Where to find it here |
|---|---|
| Coverage per field | Dataset page → Data dictionary |
| Duplicate rows | Dataset page → Quality report |
| Last processed | Dataset page → Quick facts |
| Version history | Dataset page → Changelog |
Updated 2026-09-26
