How open data is prevented

Twelve patterns used to reject calls for open data.

When open data is debated, the counter-arguments are remarkably stable, whoever makes them, in whatever forum and about whatever specific proposal. The Excuses Compass documents twelve of these recurring patterns, which serve, alone or in combination, to delay open data, to delegitimise it or to prevent it structurally.

The four axes

Some patterns push responsibility onto others, some end in resignation and defeatism, a third group plays down the need to act, and a fourth construes open data itself as the source of harm.

Pushing responsibility away

FIG. 1 · VACUUM OF RESPONSIBILITY The Buck-Passer “That's up to the IT department. Or the statistics unit. Or maybe it really is a federal matter, ask them.”

Diffuse responsibility is not a law of nature. It is the result of missing governance design.

The argument from diffuse responsibility is one of the most effective blocking patterns in the open data debate, because there is barely anything to argue against: if nobody is responsible, nobody can be held to account. The passing back and forth between IT departments, specialist units, municipal data centres and federal authorities is often not the result of genuine uncertainty but of decisions about data governance that were never taken.

Questions of responsibility are solvable. Open data officers, data coordination units and governance frameworks, such as the Berlin open data act or the open government data principle in the German E-Government Act, show how institutional responsibility can be created deliberately. Where there is political will, responsibility is established. Where it is missing, responsibility is described as missing.

Further reading
FIG. 2 · STANDARDS BLOCKADE The Standards Perfectionist “As long as there are no uniform, legally watertight metadata standards, we cannot publish anything at all. Too risky.”

DCAT-AP.de, INSPIRE, Schema.org: the standards exist. Perfection as an excuse is a form of strategic passivity.

The argument that publication must wait for perfect, uniform standards has a structural logic that makes it theoretically extendable for ever: because standards can always be improved further, the argument can in principle never be refuted. In practice it is rarely a technical description and more often a political position.

Robust interoperability standards do exist: DCAT-AP.de for Germany, DCAT 2.0 as a W3C standard for Europe, INSPIRE for geodata, OpenAPI for interfaces. Municipalities and federal authorities publish successfully with these standards every day. Pointing to missing or inadequate standards usually masks an unwillingness to begin at all.

Further reading
FIG. 3 · DATA PROTECTION REFLEX The Data Protection Shield “The moment we start opening data, we are standing in front of the data protection officer. The risk is incalculable.”

Data protection and openness are not opposites. The GDPR contains explicit opening clauses, and most administrative data is not personal data at all.

Data protection arguments against open data are rhetorically effective because they carry moral authority: who argues against protecting personal data? At the same time these arguments are routinely applied to datasets that contain no personal data at all: budget figures, infrastructure information, aggregated statistics, land use plans.

Article 86 of the GDPR contains explicit opening clauses for official documents. The German Act on the Re-Use of Public Sector Information and the EU Open Data Directive 2019/1024 create positive obligations to open up for the public sector. Data protection is a legitimate argument for data that really is personal; for the large majority of administrative datasets it simply does not apply.

Further reading
FIG. 4 · PUBLISH AND PRAY The Laissez-Faire Optimist “We put the data on the portal. What others do with it is up to them, adoption is not our job.”

Open data portals without outreach and feedback loops are digital ghost towns.

The publish-and-pray pattern reflects a technicist understanding of open data: you make the data available, and then, use happens. That understanding ignores the fact that open data is a sociotechnical ecosystem requiring relationships, trust and feedback loops. Portals without documentation, community building and feedback mechanisms have measurably lower usage rates.

Successful open data initiatives, from GovData to the Berlin data portal to the open data portal of the city of Cologne, invest deliberately in communication, named contacts, APIs and user workshops. They understand publication not as a one-way process but as the beginning of a relationship. Publication without care is not open data; it is digital filing.

Further reading

Resignation and defeatism

FIG. 5 · IT'S A FAD The Trend Despiser “Open data is like blockchain, a hype. In three years nobody will mention it. We had the same with e-government.”

Open data is anchored in EU law (Directive 2019/1024) and in national strategies. That is infrastructure, not hype.

Pointing to hype cycles is a rhetorical mechanism that delegitimises the need to act by presenting a phenomenon as short-lived and insubstantial. In the case of open data the empirical premise is false: the open data movement has gained institutional depth continuously since 2009, from Obama's Open Government Directive to the UK National Data Strategy.

Open data is binding law for all member states through EU Directive 2019/1024. It is anchored in the German E-Government Act, in the open government data strategy of the federal interior ministry and in the register modernisation act. International commitments, the G8 Open Data Charter and the Open Government Partnership with 78 member states, underline its strategic permanence. The hype became governance infrastructure long ago.

Further reading
FIG. 6 · DATA ELITISM The Citizen Pessimist “Most citizens can't do anything with raw data. We need data literacy across the board before open data makes sense.”

Universal data literacy is not required. A small group of data-literate intermediaries creates value for many.

The argument that universal data literacy is a precondition combines elitist premises with paternalistic care. It systematically underestimates data literacy in society and structurally overlooks the intermediary effect: a majority does not need to be data-literate for society to benefit from open data.

A small group of data-literate actors, data journalists at outlets like Der Spiegel or the Süddeutsche Zeitung, civil society organisations such as CorrelAid or Code for Germany, start-ups such as Citymapper or contributors to OpenStreetMap, can turn open data into insight and products of broad social use. Beyond that, data literacy grows through use rather than as a precondition for it.

Further reading

Playing down the need to act

FIG. 7 · THE TRANSPARENCY MYTH The Transparency Romantic “We publish the data, so we have transparency! What more do you want? The data is open.”

Open data without context and without a civil society to use it is a heap of files, not accountability.

Equating publication with transparency is a widespread but analytically mistaken idea. Transparency is a social phenomenon; it does not arise from the mere availability of data but from interpretive access, comprehensible presentation and social resonance. Yu and Robinson (2012) described this “new ambiguity of open government” precisely.

Raw data without context can even be counterproductive. It enables selective interpretation, creates information asymmetries between data-literate and data-inexperienced actors, and can lead to false conclusions. Real transparency requires documentation, metadata, contextualisation, explanation, and a civil society that can and wants to make sense of the data.

Further reading
FIG. 8 · QUALITY NIHILISM The Quality Nihilist “Open data means poor quality by definition. The genuinely good, validated data is not something we can release.”

Worries about quality often mask an interest in control. Openness demonstrably improves data quality.

The quality argument has a structural elegance: since flaws can always be found, the argument is theoretically inexhaustible. At the same time it ignores that keeping data closed prevents exactly the improvement that openness would enable. External users identify errors that stay invisible internally.

Open data initiatives regularly show empirically that openness improves quality: error reports from the community lead to corrections; external validation uncovers systematic problems in collection; competition between data producers raises standards. The FAIR principles (findable, accessible, interoperable, reusable) offer a frame in which quality assurance and openness are not opposites but complementary goals.

Further reading

Open data as the source of harm

FIG. 9 · INNOVATION PANIC The Enemy of Business “Open data destroys the business models of geoinformation companies. Who is going to collect quality data then?”

Empirically, open data multiplies value creation. McKinsey, Deloitte and the European Commission are consistently positive.

The argument that open data harms business interests has some empirical plausibility for companies that have so far profited from exclusive access to data collected by the state. The view across the economy as a whole, however, is clear and robustly evidenced: open data generates value rather than destroying it.

The McKinsey Global Institute (2013) estimated the global economic value of open data at three to five trillion dollars a year. The European Commission puts the EU-wide value at 184 billion euros a year and rising. The example of the UK Ordnance Survey, which earned more from value-added services after opening its geodata, shows that open data creates new markets and business models. What it destroys are outdated rent positions, not innovation.

Further reading
FIG. 10 · THE PANORAMA OF MISUSE The Prophet of Misuse “If we open the data it will immediately be misused by trolls, political opponents and foreign actors.”

Keeping data closed is no effective protection. Openness creates external scrutiny, and scrutiny uncovers misuse.

Scenarios of misuse are rhetorically effective because they address real risks and because they are hard to falsify: a plausible case of misuse can always be constructed. At the same time the analogous risks of closed data are routinely ignored: data leaks, corruption inside authorities, uncontrolled commercial trade in administrative data.

The Open Knowledge Foundation formulated precisely in the Open Definition that openness does not mean defencelessness. Licences govern the conditions of use, those conditions are enforceable, and the wider public as a body of scrutiny is often more effective than institutional gatekeepers. When everyone has access, errors, manipulation and misuse can also be identified and corrected faster. Transparency is frequently the better protection.

Further reading
FIG. 11 · COST PANIC The Cost Maximiser “Preparing open data costs a fortune. We have no budget. No new statutory duties without funding to match!”

The cost of open data is systematically overestimated, and the savings from fewer information requests are ignored.

Cost arguments appeal to real budget constraints and are therefore politically hard to counter. At the same time the actual cost of open data is regularly overestimated and the savings structurally ignored: fewer freedom-of-information requests, fewer duplicate surveys, better data quality for internal use, reduced maintenance effort thanks to external validation.

Houghton (2011) showed for Australia that the social returns on open research data exceed the costs many times over. EU programmes such as the Connecting Europe Facility, Horizon Europe and the structural funds co-finance open data infrastructure. And in many case studies the actual cost per published dataset lies far below the original estimates, particularly where data curation is built into the collection process from the start.

Further reading
FIG. 12 · CONTROL PARANOIA The One Who Fears Losing Control “Once we have released the data we lose all influence over how it is used, interpreted and presented.”

Licences, embargo periods and tiered access make controlled openness possible without losing autonomy.

Fears of losing control are understandable. An institution that has controlled its data for decades fears the consequences when others work with it and possibly reach different, more critical conclusions than the institution holding it. This fear is rarely stated explicitly but is often structurally effective.

The open data ecosystem offers differentiated instruments of control. Creative Commons licences (CC BY, CC BY-SA, CC BY-NC) govern the conditions of re-use precisely and are legally enforceable. Embargo periods protect one's own research and competitive interests. Tiered access, fully open, registered, on application, allows openness in stages. Persistent identifiers secure authorship permanently. FAIR rather than merely OPEN is not a step backwards but a differentiated practice.

Further reading

Related

Let us talk about what you have in mind.

A first conversation costs you half an hour.