Skip to content
DataPackHub

Data methodology

How datasets are processed, measured, versioned and delivered in numbered SmartPacks.

This page explains exactly what happens between a data file arriving on our server and a customer downloading it.

1. Files and packs

Each dataset is stored as separate files of approximately 50,000 records ("production packs"). Smaller trial files (about 30,000 records) are kept separately and are only sold as trial products. Every file receives a permanent pack ID, for example UK-0007.

2. Automatic inventory scan

When a file is added, the scanner:

  1. Detects the country from its folder.
  2. Counts data rows (the header row is not counted; blank rows are ignored).
  3. Records the file size and a SHA-256 checksum.
  4. Reads the column headers and identifies common field types (phone, name, email, address, city, region, postcode, and for business data company, website and category).
  5. Measures coverage: for each column, the share of rows with a non-empty value.
  6. Counts exact duplicate rows inside the file (identical in every column, ignoring letter case and surrounding spaces).
  7. Counts malformed rows (rows with more values than the header has columns).
  8. Counts phone values with an impossible digit count, and the share written in international format (with a "+" or "00" country-code prefix).
  9. Flags columns that look like sensitive personal data; such files are withheld from sale until reviewed.

Large files are read as a stream, so the scanner never needs to load a whole file into memory.

3. What the numbers on product pages mean

FigureMeaning
Records availableSum of data rows across packs currently for sale
Field coverageNon-empty values รท records, across all packs for sale
Exact duplicate rowsMeasured inside each pack and summed
Duplicates removed during processingShown only when reported by the operator, labelled "reported by operator"
Last processedMost recent modification date of the files for sale

Anything we have not measured is displayed as "Not currently measured". We never fill gaps with estimates.

4. Versions

Datasets are released as versions (for example AU-CON-2026-09 for contact data or UK-BIZ-2026-09 for business data). New data is added as new packs or a new version. If a file that has already been delivered changes on disk, its checksum no longer matches, downloads of that pack are paused and an administrator is alerted โ€” delivered data is never silently replaced.

5. Allocation

When you buy, the system reserves the lowest-numbered packs you have not received before, holds them while you pay, and allocates them once the payment provider confirms payment. Every allocation stores the pack ID, dataset version, order, date and checksum. See How SmartPacks work.

Last updated 2026-09-30