The Voter File: The most important dataset in politics

Article P3-01

Your registration is a public record. What campaigns attach to it, and whether any of it is accurate, is a separate question.

In brief

A voter file holds your registration, address, geography and turnout history, plus the party primary ballot you took. It does not hold how you voted in a general election, because there is nowhere to put that. Anything else on it, such as your race or income, was either bought or inferred from your name and neighbourhood. Inferred race is the one added trait whose error has been measured, and a common commercial labelling step under-counted Black voters in North Carolina by 28.2 percentage points. That test rests on six Southern states, and purchased data has no comparable public accuracy check.

How to use this

Before arguing about campaign data, name the layer you mean. If your complaint is about registration, address or turnout, you are describing a published document you can inspect. Ask your state what it collects and releases. If your complaint is about race, income or consumer details, you are describing something bought or inferred, so ask for accuracy evidence rather than a field list. When someone quotes an inferred trait as though it were recorded, ask how it was produced and whether the prediction was kept continuous or collapsed to a single label. And when a model's errors are described as measured, ask which population the test covered before you extend it to your own.

What the story is about

When people ask what a campaign knows about them, they usually imagine a secret file. The real starting point is plainer. In the United States, almost everything a campaign holds begins with the voter file, a list built from the public record that registering to vote creates. That record exists because you signed up, not because anyone was watching. The file is public, and so is the act that generates it. Which means the useful questions are not about surveillance. They are about what gets attached to that public list afterwards.

The spine of the file is short and dull. North Carolina publishes its own layout, and the registration file carries 74 fields, 37 of them district or geography codes such as precinct, ward, congressional district, school board, water and sewer (North Carolina State Board of Elections, 2026). A parallel history file has 15 fields, and for each election it records the election, the voting method, the precinct, the county and the party of the ballot taken (North Carolina State Board of Elections, 2021). That last field is the party of a primary ballot. It is not a record of anyone's general-election vote, and no field in either layout holds one.

What exists varies by state, so campaigns "perceive voters differently in different areas" even while most of what they know comes from the same public records (Hersh, 2015).

Everything beyond that spine is added. Race shows how the addition works and how well it has been checked. The only place inference can be graded is the six Southern states that record self-reported race at registration, which is where the name dictionaries come from (Rosenman, Olivella and Imai, 2023). Those self-reported race records are Southern and American, and nothing equivalent exists elsewhere. Within it, the errors are measured and patterned: misclassification is "strongly correlated with demographic and socioeconomic factors" (Argyle and Barber, 2024).

The commercial layer has no such test. Catalist describes a unified national voter file whose records go back over 20 years, assembled from election officials and integrated with census and commercial data (Catalist, no date). That tells you what the vendor says it does. It is not a measurement of whether the appended details are right.

So where does a trait nobody recorded come from? Two routes. A campaign can buy it from a commercial file that sells household and consumer details, or it can have it predicted from what the public record does contain.

Prediction works from two things you leave behind without meaning to: your surname and your neighbourhood. The method used for race, Bayesian Improved Surname Geocoding, compares your last name with the surnames common in each racial group, then mixes in the racial makeup of the area where you live. The result is a probability, not a fact.

That difference matters. A high probability that you belong to a group is a statement about people who resemble you. It is not knowledge about you, and it does not turn into knowledge because a database stores it in a field called race.

That kind of error has been measured, and it is large. Dong et al. (2025), in a peer-reviewed study, traced one common commercial practice: taking the single most likely race from a continuous prediction, "as used by a prominent commercial voter file vendor", and writing that one label into each voter's record. In North Carolina, the final step under-counted Black voters by 28.2 percentage points. The vendor's model was not the problem. The labelling step was. The authors also show that keeping the prediction continuous is not enough by itself, because "calibrated continuous models are insufficient to eliminate it".

Put those pieces together and the honest answer to the opening question is narrower than it feels. In the United States, a voter file records your registration, your address, your geography and your turnout history. It does not record how you voted in a general election, because there is nowhere to put that.

The race attached to your record is predicted from your name and your neighbourhood. The income and consumer details attached to it are purchased. The first has been tested, and mostly against data from six Southern states. The second has no comparable public accuracy test.

So what

The point of separating those two things is that they behave differently. The public record is fixed, published and open to inspection. The layer stacked on top of it is a product and a set of estimates, sold or generated by someone else. Which is why an argument about campaign data has to say which layer it means.

A complaint about what sits in the registration file is a complaint about a document you can read. A complaint about the modelled layer is a complaint about accuracy, and accuracy can be measured. Treating the two as one thing produces arguments that can be neither checked nor settled.

For political parties

A party building a targeting programme works from the public backbone: who is registered, where they live, and whether they have turned out before. That is where most of what campaigns know comes from, which is why campaigners treat it as the base.

On top of it sits the modelled layer. Its errors do not spread evenly across the electorate. They cluster, which means a party that treats an inferred race as a fact will be wrong more often about some voters than others.

The practical test is the same as the ethical one. Ask whether a modelling choice serves the decision you are making, or the conclusion you already preferred. A list tuned until it agrees with you is a mirror, not intelligence.

For government

A rule about voter data has to say which layer it governs. The public record is defined and published: what is collected, what is withheld, and who may copy it are all things a legislature can specify.

The purchased and inferred layer is harder to write rules for, because there is nothing comparable to inspect. The one accuracy test the public has is for inferred race, and it rests on the six Southern states that collect self-reported race at registration, so even that test covers a fraction of the country.

That does not make the layer ungovernable. It means a disclosure rule aimed at the file's fields will reach the spine and miss everything stacked on it. If the aim is to know whether the added data is any good, the rule has to ask for accuracy evidence rather than a field list.

References

Argyle, L.P. and Barber, M. (2024) 'Misclassification and bias in predictions of individual ethnicity from administrative records', American Political Science Review, 118(2), pp. 1058–1066. Available at: https://doi.org/10.1017/S0003055423000229 (Accessed: 10 September 2026).

Catalist (no date) Dynamic national database. Available at: https://catalist.us/data/ (Accessed: 10 September 2026).

Dong, E. et al. (2025) 'Addressing discretization-induced bias in demographic prediction', PNAS Nexus, 4(2), pgaf027. Available at: https://doi.org/10.1093/pnasnexus/pgaf027 (Accessed: 10 September 2026).

Hersh, E.D. (2015) Hacking the electorate: how campaigns perceive voters. Cambridge: Cambridge University Press. Available at: https://doi.org/10.1017/CBO9781316212783 (Accessed: 10 September 2026).

North Carolina State Board of Elections (2021) Voter history data file layout (layout_ncvhis.txt). Raleigh, NC: NCSBE. Available at: https://dl.ncsbe.gov/data/layout_ncvhis.txt (Accessed: 10 September 2026). Calculations from this data were made for this article and are available upon request.

North Carolina State Board of Elections (2026) Voter registration data file layout (layout_ncvoter.txt). Raleigh, NC: NCSBE. Available at: https://dl.ncsbe.gov/data/layout_ncvoter.txt (Accessed: 10 September 2026). Calculations from this data were made for this article and are available upon request.

Rosenman, E.T.R., Olivella, S. and Imai, K. (2023) 'Race and ethnicity data for first, middle, and surnames', Scientific Data, 10(1), 299. Available at: https://doi.org/10.1038/s41597-023-02202-2 (Accessed: 10 September 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact