Inside the data: 9,516 polls, two licenses, and an anomaly rate that needs explaining
Over the last few days we completed a new harvest from Italy's official register of political and electoral polls (the database under art. 8 c.3 of Law 28/2000, Presidency of the Council of Ministers): 7,958 polls, collected page by page from the public registry. They sit alongside the 1,558 already in the archive, sourced from onData under an open license. This piece is about what is inside — and above all about what is not.
- 9,516polls in the archive, from June 2012 to August 2026
- 19.4%flagged
suspecton the official source (8.8% on onData) - 90.7%of official polls with no assignable territorial code
- 655days — the largest gap between fieldwork end and filing
Two sources, two licenses, never mixed
The 7,958 polls from the official registry carry no stated reuse license: the collection is public by statutory obligation, but that does not make it freely redistributable. The 1,558 onData polls are released under CC-BY-4.0.
The two sources always stay distinct in our system. We never present them under a single license badge, every row carries its own source, reference and attribution, and exports respect the regime of the source a data point came from.
The registry is not an archive of voting intentions
This is the biggest surprise of the harvest: half the registry is not about elections at all.
Nearly half the filings are thematic surveys — opinions on individual policies, consumption, attitudes — and only just over a quarter are national voting intention. It is an archive of the country's polling activity, not an electoral archive, and that changes how it should be queried.
Why one source has more anomalies (and not for the obvious reason)
The share of polls flagged suspect by our checks is 19.4% on the official source against 8.8% on onData. The intuitive explanation — "the registry holds far more thematic questions, and those are harder to normalise" — is contradicted by the data.
suspect, by question type
Thematic questions have an anomaly rate of 8.4%, below the source average. The flags concentrate instead exactly where the data matters most to us: national voting intention (39.4%), municipal (41.6%) and regional (41.2%).
And comparing the two sources at equal question type does not close the gap — it widens it: on national voting intention alone, 39.4% on the official registry against 7.7% on onData.
The real cause is format. The registry collects documents filed by the polling houses, and the structured fields are few: almost everything has to be extracted from free text and attachments. The most frequent flags say so plainly.
| Quality flag | Polls | What it means |
|---|---|---|
margin_error_confidence_unknown | 4,514 | Margin of error with no stated confidence level |
attachment_unparsed | 2,127 | Attachment present but not machine-extractable |
sample_size_unparsed | 1,679 | Sample size not recognised in the text |
vote_question_unparsed_q1 | 284 | Vote question present but not reconstructible |
geo_unresolved | 142 | Territory named but not resolvable to an entity |
onData, a dataset already normalised upstream by people, does not pay this price. Put differently: the anomaly rate measures distance from the original document, not the reliability of the house that ran the poll.
What the registry does not contain at all
- No ISTAT territorial code. Only a free-text name ("comune di Milano", "Piemonte"). Geolocation therefore runs on name matching: when a name is ambiguous (municipalities sharing a name) or unresolvable, the poll is left without an assigned geography and explicitly flagged — never forced onto the first plausible match. Today 90.7% of official polls have no territory.
- Sample size missing in one case out of five. 1,679 official polls (21.1%) report no readable sample: without it no margin of error can be computed, and the poll cannot enter any weighted calculation.
- No party register. Lists appear as text labels. This is the limit that weighs most on analysis, and the one we are working on now.
That last point is worth making concrete, because it is the kind of trap that produces wrong charts without anyone noticing. In the archive the label "Futuro Nazionale" appears both in three 2013 readings, where it averages 0.6%, and in 184 readings from 2026, where it averages 5.0%: two different political entities under the same name. And the label "Roberto Vannacci", worth over 20% in some readings, is not a list at all — it is a personal approval question that ended up inside a poll classified as national vote.
Until the canonical register is complete, any party-level aggregation on this source has to be done by explicit label matching and declared as such — which is exactly what we do in the other two pieces published today.
When a poll "appears" is not when it was taken
The date a poll shows up in the registry is not the date it was conducted.
| Filing lag after fieldwork end | Official registry | onData |
|---|---|---|
| Median | 3 days | 3 days |
| 90th percentile | 14 days | 5 days |
| Maximum | 655 days | 369 days |
In the typical case the lag is immaterial, but the tail is long and it always runs in the same direction: the data looks more recent than it is. That is why every time series on the platform places a poll on its fieldwork end where available, falling back to the filing date only when the former is missing.
Why this work matters
Polls are not observed electoral data: they are third-party estimates, produced with methods we do not control, on samples we often cannot see. In our architecture they therefore always stay separate from actual election results, and never enter forecasting models as ground truth — only as one contextual feature among many, carrying its own quality flag.
Publishing an archive's inventory of defects alongside the archive is the least eye-catching part of this work, and it is the part that makes everything else usable. We will hold to the same standard with every future addition.
- The archive reaches 9,516 polls from two sources under different licenses, always kept separate.
- The official registry is half thematic surveys: it is not an electoral archive.
- The higher anomaly rate is not driven by thematic questions, which are among the cleanest: it comes from extracting filed documents.
- The real gaps are territory, sample size and party register — we declare them rather than guessing them away.
- A poll should be dated to its fieldwork end, not its filing.
Source: ElectionLab polling archive — official register of political and electoral polls (database under art. 8 c.3 of Law 28/2000, Presidency of the Council of Ministers) and the onData dataset (CC-BY-4.0). Counts as of 29 August 2026 across the whole archive. Explore it in the Polls section, or read the two analyses built on it: the centre-right decline and pollster house effects.
— ElectionLab Studio