The core of the SmartCensus is the names and addresses of adults in the UK, collated as annual snapshots from 1997 onwards. It is compiled under more than a dozen Data Licencing Agreements by linking public versions of the Electoral Register and multiple consumer registers provided by GeoDS data partners. GeoDS data scientists document the reliability and provenance of the data, typically in peer reviewed articles in internationally recognised academic journals.
Content
The SmartCensus is held in an ISO27001-accredited Trusted Research Environment (TRE), for access only by professionally accredited GeoDS data scientists. Retaining of individual level observations empowers researchers to undertake individual level research on the adult UK population. Non-disclosive derivative datasets provide detailed estimates and annual updates of the changing characteristics of local populations: for example, ‘research ready data’ created by the GeoDS includes annual neighbourhood modelled ethnicity proportions (using the GeoDS Ethnicity Estimator names software) and estimated origin- destination pairings of residential moves.
The complexity of multiple updated data licencing agreements precludes access to the raw data by third parties, but suggestions for new derivative research ready data products are welcomed..
It is anticipated that data will continue to be updated on an annual basis. GeoDS strives towards continual improvement in the completeness of the SmartCensus in the recent past as well as the present, so each version of the derivative research ready data should be viewed as provisional and subject to some updating as more data become available.
For detailed description of the columns contained within the data, see the Variable Dictionary - and for an overview of the characteristics of the data, see the Data Summary. These files can be downloaded from the bottom of this page.
Quality, Representation and Bias
Individual address records are georeferenced using the Assign algorithm (&reference please). The data are ingested from multiple organisations, and it is possible that some names and addresses are inconsistently formatted between datasets. Cross validation between data sources is used to quantify the veracity of each addressee. Despite applying bespoke rigorous address matching and name matching methodologies, it is still possible that same cases are not matched. Thus the number of unique addresses is slightly overestimated.
Historically, addresses have been recorded as address lines in a separate table. This can render address matching to other data quite difficult as the number and composition of address lines varies by addresses, and between different versions of the data too.
In addition, it is very difficult to determine the completeness of the data. The SmartCensus has near complete coverage of the adult population for some years, while others are more heavily reliant upon carry back of records from years in which our changing portfolio of suppliers provide more complete coverage. Data lags are known to occur in uptake of electoral registration and this is factored into data quality audits. It is acknowledged that adults resident at multiple addresses may have duplicate entries within the data.
Whilst the data providers have endeavoured to compile registers which are as complete and accurate as possible, this does not preclude bias in the data that are collected. Lan and Longley (2022) investigate the potential bias in recording of ethnic minorities by comparing modelled ethnicity proportions estimated from the SmartCensus (previously called Linked Consumer Registers) with those recorded in a UK Census of Population (https://journals.sagepub.com/doi/full/10.1177/01600176221116568). Roughly 84% of individuals were classed as White British, compared to 81% from the 2011 UK Census. We have also identified that areas with higher proportion of adults in rented accommodation had the greatest underrepresentation within cities.
Reliability scores for the dataset have been calculated and are available upon request.