Tag Archives: news

Data Localization is the Answer. What Was the Question?

Last week I gave a talk at Chatham House for an event on ‘Is Data Localisation the Answer to Data Sovereignty? Location, Control and the Architecture of Trust’,  organised in collaboration with their Digital Ambassadors Forum. I really enjoyed the conversation and questions and had a number of people ask for slides and more details so have put together the following summary.


Four years ago I had the privilege of hosting digital transformation leaders from around the world at Rockefeller’s Bellagio conference centre. In the room we had senior officials from governments that represented at least 2 billion people. Much of our conversation focused on digital public infrastructure (then…a new topic) but at one point the conversation shifted to data and sovereignty. Most leaders in the room insisted that yes, data mattered and, obviously, data localization was the answer. Everyone, that is, except the Ukrainian in the room. We’ll come back to that.

If data localization is the solution… what problem is it solving? What is the nature of the “data sovereignty” we think is a threat? My experience is that data’s location says very little about sovereignty… and the reduction of local vs abroad on where data is stored both presupposes some fairly pre-digital notions of sovereignty and ignores a lot of options and tradeoffs that could befuddle a policy maker. If you’re too busy to read any of this post, this first slide sums up much of what I’m getting at. On one axis is a range of models for storing data. Some are “local,” others are not. On the other axis is a range of threats one should be thinking about when storing data (this list isn’t even exhaustive). The resulting matrix (which would be unique for any country or company) tries to lay out the real tradeoffs between storage models and risk types. But more than anything else it suggests localization is not a magic cure-all for one’s policy or economic woes.

I think sovereignty (or as I prefer, agency) is best achieved via adopting multiple models to provide coverage against a range of threats. This also has tradeoffs, and, to be most effective suggests governments should be trying to shape the cloud market into being more interoperable. This argument – which I refer to as a “commoditized stack” – offers a more promising path to agency than anything currently on the table.

There is Both New Infrastructure and a Genuine Problem

First, some context. Any conversation about data, data localization and/or digital sovereignty has to contend with how and where data is stored. This quickly gets you to a topic I’m deeply interested in: the cloud. I recently presented to Europe’s finance ministers (the memo I wrote for them is published on the Eurogroup’s website) and I first sought to impress on them that “the cloud” is 21st-century infrastructure in exactly the way water systems, railways, electricity and telecoms were essential infrastructures in previous centuries.

The cloud – which at a minimum is the provisioning of storage and compute at scale – is now critical to any modern economy. It is a core input into services ranging from food delivery to government benefits to banking. Happily, most people get this, and it is a reason why concern over who controls data has become so important.

Whose Data are we Talking About?

There’s a second question hiding inside the “data localization” question. When a government says “our data must be localized,” which data does it mean? Its own — the tax records, health files and registries it holds in trust for citizens? Or the data generated by the entire economy — every bank, hospital, startup and grocery chain?

These are radically different problems. The first is, in many ways, a government procurement question. This is in part because states carry a special custodial duty. Citizens cannot opt out of giving the state their data. But it is also a function of the fact that states are expected to be actors of last resort, capable of functioning under even the most brutal circumstances – such as a state of war.

The second is industrial policy. Requiring the whole economy to localize means repricing compute and storage for every firm in the country. Most localization debates slide between “the government’s data” and “the nation’s data” without noticing they have changed the subject. For this piece, let’s assume we are just talking about government data, but you can easily expand this to the whole economy and see how the implications and tradeoffs become more daunting.

What is the Threat Model?

To return to our title, if data localization is the solution, what is the problem? Here policy makers would benefit from a little threat modeling, to understand what the threats (and potentially opportunities) data localization is seeking to answer.

At present there is a threat that many policy makers have in mind. In most European and OECD countries (and beyond), many people, when pressed, land on the same answer: lawful access (or unlawful access, depending on your point of view). Specifically, they are worried about the CLOUD Act, the 2018 US law that lets American authorities compel US providers (like AWS, Azure or Google Cloud) to hand over data stored on servers they manage, regardless of where they are located around the world. Somewhat related, people also worry about a denial of access, whereby, via legal means, a government might be denied access to its own data.

I too find the CLOUD Act deeply, deeply problematic, and an excellent reason why one should think carefully about becoming reliant on US (or Chinese) cloud providers. But it isn’t the only threat to a government’s data. Here, for example, is a fuller list:

Lawful (or unlawful) access remains a threat. And there are others:

  • Access denial: not someone reading your data, but someone simply turning the servers off, or severing connectivity to them, because they can.
  • Cyber attack: an actor stealing or ransomwaring your data.
  • Data colonialism: an actor exporting data outside your jurisdiction to be exploited by foreign firm(s).
  • Lack of capacity: an inability to use or protect your own data.
  • Competitiveness: when storing, managing and accessing your data is simply more expensive than in other jurisdictions, leaving you at a competitive disadvantage.
  • Act of God: Your ability to access your data (particularly if stored in a single location or region) is at risk from a natural disaster.

Any one of these can compromise data integrity or security – which would suggest it is core to discussions about “data sovereignty” – but almost none of these figure in the political discourse. Everyone focuses on the lawful access and access denial problem. (To be fair, I find most engineers and business people are thinking about all these threats; it’s mostly policy people zeroing in on the first two.) But if you start to think about various possible threats, solutions become trickier as each threat points in a different direction.

Where does data actually live?

Once you’ve wrapped your head around the numerous ways the universe might conspire to destroy your data, you have to reconcile that with the options you have to store your data. There are several and one can somewhat imperfectly align them along an axis of “foreign controlled” to “locally controlled.”

At one end you have your classic hyperscaler model, a US firm offering almost limitless storage, with high redundancy (as well as a range of other platform services) all subject to the CLOUD Act, with your data primarily stored in a data centre in Virginia. At the other end, a government-owned data centre you nominally control entirely (we could debate the provenance of the technology inside it, but set that aside). In between sit a range of actors we rarely bother to distinguish: a firm headquartered in your country that hosts your data in a datacentre located abroad; a hyperscaler’s local data centre on your soil (but still subject to the CLOUD Act!); a domestic joint venture running on licensed foreign technology; a domestic firm hosting domestically. Very different rules apply to each.

In addition, what is foreign controlled and locally controlled has little alignment with data localization. Indeed the localization debate flattens them into “local good, foreign bad” which, depending on your threat model, may have profound, and not necessarily positive implications for your data’s security or sovereignty.

To make this more obvious you can run the aforementioned threats against this spectrum of options. This is where it gets uncomfortable.

Lawful access: localization doesn’t help

Start with the threat everyone cares about. The CLOUD Act applies to data sitting in Ireland, or the UK, or Montreal on an American hyperscaler exactly as much as it applies to data sitting in Virginia. The Act follows the company, not the geography. Moving your data into a hyperscaler’s local data centre — the single most popular “sovereignty” measure on offer (and one whole procurement frameworks are built around!) — does not address this threat.

This is one reason why governments are being persuaded to invest in their own “local” or “national” clouds. Sovereignty is derived from ownership (which, as we’ve seen, it may not be) and comes with its own tradeoffs…

What actually helps against lawful access is holding your own encryption keys, or using a provider genuinely outside the requesting state’s jurisdiction — and each of those brings its own trade-offs from elsewhere on the list.

Cyber attack: you might be safer on the hyperscaler

Against a serious state-backed attacker, where do you want your data? The honest answer is awkward: probably on the platform with the largest, best-resourced defensive security team on earth. If you’re a middle power government, or a large, successful private company, do you trust your state agencies to out-defend firms whose security teams and budgets dwarf your own? If your “A” team is protecting national secrets, who’s left to protect more mundane things like social benefits or the DMV? If you’re concerned about attacks from North Korea, Russia or China — states with sophisticated capabilities – this analysis becomes even more awkward. Localizing your data onto weaker domestic infrastructure can make this threat worse.

Access denial: ownership beats location — sometimes

Access denial is the threat I think deserves far more attention than it gets. Forget reading your data; someone with control over your infrastructure can simply deny you access to it. Maybe via legal process, or… not. They could just turn the power off or sever a key network cable.

Here the spectrum behaves differently. Physical denial is possible for anything hosted outside your borders — including by a domestically-owned firm operating internationally. Legal denial can reach even a foreign-owned entity that happens to sit on your soil: if the parent company is ordered to stop serving you, the local data centre goes dark just the same. Only the bottom of the spectrum — genuinely domestic operations — offers much protection.

Localization can help here… with caveats. A domestic operation still depends on hardware, software and services that are not indigenous to your country. You may have traded a low risk of someone flipping the off-switch for an elevated cyber risk and the slow grind of a degraded service. But of all the threats on this list, access denial is the one localization most clearly addresses.

Attack and act of god: the safest place may not be your country

Remember the Ukrainian in the room at Bellagio, the one person least sold on localization? They were the ones presenting. It was a talk I still think about: photo after photo of bombed-out Ukrainian data centres. When your threat model includes missiles, “keep the data at home” reads very differently. Data localization isn’t necessarily the best defence against hostile acts. Ukraine has become a significant adopter of US hyperscalers. They are not alone. Estonia drew this lesson a decade earlier and implemented its data embassies, placing core data abroad on purpose to ensure continuity of service in the event of a hostile kinetic attack.

Data localization can also carry risks and tradeoffs. Last September, South Korea’s National Information Resources Service — the government’s own data centre in Daejeon — experienced a fire. Six hundred and forty-seven government services went down. Ninety-six systems were destroyed outright, including G-Drive, the government’s internal file store: roughly 858 terabytes of working documents, gone. There was… no backup. Last time I searched, it was still unrecoverable. This wasn’t an attack. It was infrastructure, in one place, and that one place burned.

The hard reality is, depending on your threat model, the safest place for your data may not be in your country. My home of Canada has the luxury of being enormous — an act of God is unlikely to take out data centres in Vancouver and Montreal simultaneously (if it does, I suspect we collectively, as a planet, have bigger problems). But many countries don’t have that luxury. If you’re Belgium or an island nation, I’m not sure you do. As mentioned above, Estonians have decided, explicitly, that they don’t. An Estonian “data embassy” in Luxembourg (and possibly elsewhere) stores sovereign Estonian data abroad, precisely because it’s safest there.

Data colonization: it’s about who collects, not where it sits

One more, because it exposes the frame’s deepest flaw. If your worry is the extraction of your data’s economic value — data colonialism, industrial loss — then where the data is stored is only part of the issue. Equally important is who collects it. A foreign application harvesting your citizens’ data will do so regardless of which data centre it rents. Does a domestic company that collects the data, then chooses to store it abroad, raise bigger questions than a foreign company storing your data locally?

Every Choice is a Trade-Off

To really drive this home I created an illustrative matrix of these trade-offs. The ratings would likely differ from country to country, but it gives one a sense of the terrain. There is no row that is all green. There never will be. And that is my point. Data localization is a solution to a very specific threat model… an important one, but a specific one. And the notion that you should spend billions or more to create a national champion to solve for that problem is… a trade-off in its own right that, even executed well, may carry significant risks and downsides.

So the first conclusion of the talk: data localization is not a strategy. It is a tactic that addresses some threats, may sometimes be effective for the headline threat, and actively worsens others. If someone is leading with localization, the serious response is: against which threat? Until that question has an answer, it’s hard to balance risk and reward, costs and options.

Moving from Sovereignty to Optionality

So what would a more successful strategy for achieving “sovereignty” look like? First and foremost, it means not starting with a solution but with a threat model — and being honest about the tradeoffs in addressing each threat. As we’ve just seen, that exercise rarely lands on a single answer.

I think it also lies in creating optionality, both domestically and internationally. This is what I’ve argued for in The Path to a Sovereign Tech Stack is Via a Commodified Tech Stack. The core thesis being that standardizing storage at the cloud level, and then, ideally, platform-as-a-service functions, would bring portability, competition and federative options to the cloud layer – including, critically, around data storage.

Focusing on choice and interoperability creates a range of more interesting options than focusing on “sovereignty” since, as I previously mentioned, that conjures up unhelpful notions that conflate territoriality and ownership with control and security. Mike Bracken and I first suggested agency was a more helpful term in our Lisbon Council piece on Europe’s Digital Strategy. Mike’s gone on to write some additional good thoughts, but here is one of my favourite parts of that original piece:

Which brings us to a deeper concern: the framing of digital sovereignty itself.

We understand the instinct. In a world where digital infrastructure is often foreign-owned and opaque, wanting more control is natural. But sovereignty, if defined as owning every layer of a technological stack, can quickly become counterproductive. It leads to balkanized systems, limited interoperability and a retrenchment from global cooperation. The incentives for creative development disappear. Competition and innovation subside.

Instead, we should be aiming for digital resilience and interoperability.

Summary Advice

If you’ve read this far, here is the version to carry into your next meeting.

Data localization is an answer waiting for a question. Before adopting it, name your threat model. If the threat is the CLOUD Act, localization does not help you — the law follows the company, not the geography. If the threat is a state-backed cyber attack, localization may hurt you. If the threat is access denial, it helps — read the fine print on ownership. If the threat is fire, flood or missiles, the safest place for your data may be another country entirely; ask Estonia, or ask South Korea, which kept its data sovereign, domestic, in one building, and lost 858 terabytes of it in an afternoon.

If you carry only three questions out of this post, make them these.

Whose data? A rule that is prudent stewardship for a tax agency becomes an economy-wide tax when applied to every firm in the country.

Against which threat? Localization genuinely helps with some threats, does little for the headline one, and makes others worse. Name the threat, and the right storage model usually names itself.

Can you leave? Sovereignty that depends on any single provider’s goodwill — foreign or domestic — isn’t sovereignty. The durable kind comes from being able to move. That’s the commoditized stack argument, and it’s the subject of my next post.


The Chatham House conversation was held under a mix of rules; nothing here draws on anything said in the room — this is the argument I brought in. Related: Parting Clouds (with Curtis McCord, for the Canadian Anti-Monopoly Project), The Path to a Sovereign Tech Stack is via a Commoditized Tech Stack (Tech Policy Press), and NATO’s Digital Back End Could Fall Apart Without Change (Foreign Policy). If you’re working on any of this elsewhere in the world, I’d like to compare notes.

Connectedness, Volleyball and Online Communities

I’m currently knee deep into Connected: The Surprising Power of Our Social Networks and How They Shape Our Lives by Christakis & Fowler and am thoroughly enjoying it.

One fascinating phenomenon the book explores is how emotions can spread from person to person. In other words, when you are happy you increase the odds your friends will be happy and, interestingly, that your friends’ friends will be happy. Indeed Christakis & Fowler’s research suggests that emotional states are, to a limited degree, contagious.

All this reminded me of playing competitive volleyball. I’ve always felt volleyball is one of the most psychologically difficult games to play. I’ve regularly seen fantastic teams collapse in front of vastly inferior opponents. I used to believe that the unique structure of volleyball accounted for this. The challenge is that the game pauses at the end of every point, allowing players to reflect on what happened and, more importantly, assign blame (which can often be allocated to a single individual on the team). This makes it easy for teams to start blaming a player, over-think a point, or get frustrated with one another.

As a result, even prior to reading Connected, I knew team cohesion was essential in volleyball (and, admittedly, any sport) . This is often why, between points, you’ll see volleyball teams come together and high-five or shake hands even if they lost the point. If emotions and confidence are contagious, I can now see why it is a team starts to lose a little confidence and consequently then plays a little worse causing them to lose still more confidence and then, suddenly they are in a negative rut and can’t escape.(Indeed, this peer reviewed paper showed that tactile touch among NBA players was a predictor of individual and team success)

Of course, I’ve also long believed the same is true of online (and, in particular, open source) communities. That poisonous and difficult people don’t just negatively impact the people they are in direct contact with, but also a broader group of people in the community. Moreover, because communication often takes place in comment threads the negative impact of poisonous people could potentially linger, dragging down the attitude and confidence of everyone else in the community. I’ve often thought that the consequence of negative behaviours in the online communities has been underestimated – Christakis and Fowler’s research suggests there are some more concrete ways to measure this negative impact, and to track it. Negative behaviour fosters (and possibly even attracts) still more negative behaviour, creating a downward loop and likely causing positive or constructive people to opt out, or even never join the community in the first place.

In the end, finding ways to identify, track and mitigating negative behaviour early on – before it becomes contagious – is probably highly important. This is just an issue of having people be positive, it is about creating a productive and effective space, be it in pursuit of an open source software product, or a vibrant and interesting discussion at the end of an online newspaper article.

Why Old Media and Social Media Don't Get Along

Earlier today I did a brief drop in phone interview on CPAC’s Goldhawk Live. The topic was “Have social media and technology changed the way Canadians get news?” and Christoper Waddell, the Director of Carleton University’s School of Journalism and Chris Dornan, Director of Carleton University’s Arthur Kroeger School of Public Affairs were Goldhawk’s panel of experts.

Watching the program prior to being brought in I couldn’t help but feel I live on a different planet from many who talk about the media. Ultimately, the debate was characterized by a reactive, negative view on the part of the mainstream media supporters. To them, threats are everywhere. The future is bleak, and everything, especially democratic institutions and civilization itself teeter on the edge. Meanwhile social media advocates such as myself are characterized as delusional techno-utopians. Nothing, of course, could be further from the truth. Indeed, both sides share a lot in common. What distinguishes though, is that while traditionalists are doom and gloom, we are almost defined by the sense of the possible. New things, new ideas, new approaches are becoming available every day. Yes, there will be new problems, but there will also be new possibilities and, at least, we can invent and innovate.

I’m just soooooo tired of the doom and gloom. It really makes one want to give up on the main stream media (like many, many, many people under 30 have). But, we can’t. We’ve got to save these guys from themselves – the institutions and the brands matter (I think). So, in that pursuit, let’s tackle the beast head on, again.

Last, night the worse offender was Goldhawk, who tapped into every myth that surrounds this debate. Let’s review them one by one.

Myth 1: The average blog is not very good – so how can we rely on blogs for media?

For this myth, I’m going to first pull a little from Missing the Link, now about to be published as a chapter in a journalism textbook called “The New Journalist”:

The qualitative error made by print journalists is to assume that they are competing against the average quality of online content. There may be 1.5 million posts a day, but as anyone whose read a friend’s blog knows, even the average quality of this content is poor. But this has lulled the industry into a false sense of confidence. As Paul Graham describes: “In the old world of ‘channels’ (e.g. newspapers) it meant something to talk about average quality, because that’s what everyone was getting whether they liked it or not. But now you can read any writer you want. Consequently, print media isn’t competing against the average quality of online writing, they’re competing against the best writing online…Those in the print media who dismiss online writing because of its low average quality are missing an important point. No one reads the average blog.”

You know what though, I’m going to build on that. Goldhawk keeps talking about the average blog or average twitterer (which of course, no one follows, we all follow big names, like Clay Shirky and Tim O’Reilly). But you know what? They keep comparing the average blog to the best newspapers. The fact is, even the average newspaper sucks. The Globe represents the apex of the newspaper industry in Canada, not the average, so stop using it as an example. To get the average, go into any mid-sized town and grab a newspaper. It won’t be interesting. Especially to you – an outsider. It will have stories that will appeal to a narrow audience, and even then, many of these will not be particularly well written. More importantly still, there will little, and likely no, investigative journalism – that thing that allegedly separates blogs from newspapers. Indeed, even here in Vancouver, a large city, it is frightening how many times press releases get marginally touched up and then released as “a story.” This is the system that we are afraid of losing?

Myth 2: How will people sort good from low quality news?

I always love this myth. In short, it presumes that the one thing the internet has been fantastic at developing – filters – simple won’t evolve in a part of the media ecosystem (news) where people desperately want them. At best, this is naive. At worse, it is insulting. Filters will develop. They already have. Twitter is my favourite news filter – I probably get more news via it than any other source. Google is another. Nothing gets you to a post or article about a subject you are interested in like a good (old-fashioned?) google search. And yes, there is also going to be a market for branded content – people will look for that as short cut for figuring out what to read. But please people are smarter than you think at finding news sources.

Myth 3: People lack media savvy to know good from low quality news.

I love the elitist contempt the media industry sometimes has towards its readers. But, okay, let’s say this is true. Then the newspapers and mainstream media have only themselves to blame. If people don’t know what good news is, it is because they’ve never seen it (and by and large, they haven’t). The most devastating critique on this myth is actually delivered by one of my favourite newspaper men: Kenneth Whyte is his must listen-to Dalton Camp Lecture on journalism. In it Whyte talks about how, in the late 19th and early 20th century NYC had dozens and dozens of newspapers that fought for readership and people were media savvy, shifting from paper to paper depending on quality and perspective. That all changed with consolidation and a shift from paying for content to advertising for content. Advertisers want staid, plain, boring newspapers with big audiences. This means newspapers play to the lowest common denominator and are market oriented to be boring. It also leaves them beholden to corporate interests (when was the last time the Vancouver Sun really did a critical analysis of the housing industry – it’s biggest advertisement source?). If people are not media savvy it is, in part, because the media ecosystem demands so little of them. I suspect that social media can and will change this. Big newspapers may be what we know, but they may not be good for citizenship or democracy.

Myth 4: There will be no good (and certainly no investigative) journalism with mainstream media.

Possible. I think the investigative journalism concern is legitimate. That said, I’m also not convinced there is a ton of investigative journalism going on. There may also be more going on in the blogs than we might know. It could be that these stories a) don’t get prominence and b) even when they do, often newspapers don’t cite blogs, and so a story first broken by a blog may not be attributed. But investigative journalism comes in different shapes and sizes. As I wrote in one of my more viewed posts, The Death of Journalism:

I suspect the ideal of good journalism will shift from being what Gladwell calls puzzle solving to mystery solving. In the former you must find a critical piece of the puzzle – one that is hidden to you – in order to explain an event. This is the Woodward and Bernstein model of journalism – the current ideal. But in a transparent landscape where huge amounts of information about most organizations is being generated and shared the critical role of the journalist will be that of mystery solving – figuring out how to analyze, synthesize and discover the mystery within the vast quantity of information. As Gladwell recounts this was ironically the very type of journalism that brought down Enron (an organization that was open, albeit deeply  flawed). All of the pieces of that lead to the story that “exposed” Enron were freely, voluntarily and happily given to reports by Enron. It’s just a pity it didn’t happen much, much sooner.

I for one would celebrate the rise of this mystery focused style of “journalism.” It has been sorely needed over the past few years. Indeed, the housing crises that lead to the current financial crises is a perfect example of case where we needed mystery solving not puzzle solving, journalism. The fact that sub-prime mortgages were being sold and re-packaged was not a secret, what was lacking was enough people willing to analyze and write about this complex mystery and its dangerous implications.

And finally, Myth 5: People only read stories that confirm their biases.

Rather than Goldhawk it was Christopher Waddell who kept bringing this point up. This problem, sometimes referred to as “the echo chamber” effect is often cited as a reason why online media is “bad.” I’d love to know Waddell’s sources (I’m confident he has some – he is very sharp). I’ve just not seen any myself. Indeed, Andrew Potter recently sent me a link to “Ideological Segregation Online and Offline.” What is it? A peer reviewed study that found no evidence the Internet is becoming more ideologically segregated. And the comparison is itself deeply flawed. How many conservatives read the Globe? How many liberals read the National Post? I love the idea that somehow main stream media doesn’t ideologically segregate an audience. Hasn’t any looked at Fox or MSNBC recently?

Ultimately, it is hard to watch (or participate) in these shows without attributing all sorts of motivations to those involved. I keep feeling like people are defending the status quo and trying to justify their role in the news ecosystem. To be fair, it is a frightening time to be in media.

When someone demands to know how we are going to replace newspapers, they are really demanding to be told that we are not living through a revolution. They are demanding to be told that old systems won’t break before new systems are in place. They are demanding to be told that ancient social bargains aren’t in peril, that core institutions will be spared, that new methods of spreading information will improve previous practice rather than upending it. They are demanding to be lied to.

And I refuse to lie. It sucks to be a newscaster or a journalist or a columnist. Especially if you are older. Forget about the institutions (they’ve already been changing) but the culture of newsmedia, which many employed in the field cling strongly to, is evolving and changing. That is a painful process, especially to those who have dedicated their life to it. But that old world was far from perfect. Yes, the new world will have problems, but they will be new problems, and there may yet be solutions to them, what I do know is that there aren’t solutions to the old problems in the old system and frankly, I’m tired of those old problems. So let’s get on with it. Be critical, but please, stop spreading the myths and the fear mongering.