LowEndBox - Cheap VPS, Hosting and Dedicated Server Deals

When Data Sovereignty Rules Bite You Hard: Some AWS Data in UAE, Bahrain is Gone for Good

AWS on FireThe great promise of cloud computing has always been that hardware failures become someone else’s problem.  Disks die, power fails, network links get cut, and entire buildings can disappear from service, but applications keep running because the infrastructure behind them is redundant.

Amazon has spent years and a great deal of sales capital teaching customers to think in exactly those terms.  Spread a workload across multiple Availability Zones, keep backups, design for failure, and the death of an individual machine or even an individual data center should be a non-event.

That worked until the failure domain became a war zone.

On September 15, AWS said it could not restore access to resources and data hosted exclusively in its Middle East (Bahrain) Region, me-south-1.  The company said damage there spanned multiple Availability Zones and exceeded what its regional and multi-AZ services were designed to withstand.  AWS reached a similar conclusion for one Availability Zone in the UAE Region, mec1-az2, while recovery work continues for other affected UAE resources.

The damage dates back to March, when drone strikes physically hit AWS infrastructure in the Gulf.  AWS said two facilities in the UAE were directly struck, while a strike near a Bahrain facility caused physical damage.  The attacks caused structural damage and power disruption, and in some cases fire suppression created additional water damage.  A second Bahrain Availability Zone was later disrupted in April, taking the Bahrain Region offline.

Most customers, AWS says, were eventually able to re-establish operations in other Regions using backups or data that remained accessible.  But, of course, if “most customers” = “all customers,” you wouldn’t be reading this story.  For data that existed only inside the damaged Bahrain Region, AWS now says the recovery options have been exhausted.

Remember “You Wouldn’t Notice”?

There is an old AWS quote that has aged in a particularly interesting way.

In 2017, CBS asked AWS executive Matt Wood a rather pointed question:

CBS: “I don’t mean to give anyone ideas, but let’s say I figured out that one of these unmarked buildings was an AWS data center, and I blew it up.  Are you saying that it’s so backed up and redundant that you probably wouldn’t notice?”

AWS: “Yeah, you wouldn’t notice.  I mean, we might be a bit upset, but you wouldn’t notice!”

Nine years later, somebody effectively ran the experiment, although on a much larger scale than the loss of a single building, and people did, in fact, notice.

Data Sovereignty Handcuffs

Forcing businesses and users to keep data in places their governments tell them to keep them is all the rage these days.  The theory is that by forcing citizens to keep data in regions under that government’s control, they can more effectively protect it against spying, snooping, and hackers.

The problem with this idea is that it concentrates data, in sometimes very unhealthy ways from a disaster protection perspective.  And that’s what happened here.  Bahrain and UAE have data sovereignty laws.  Unfortunately for them, AWS doesn’t have multiple regions in their countries.  You can distribute the workload among Availability Zones inside the Region, but all of those zones still exist inside the same geopolitical failure domain.

A multi-AZ architecture protects you from a lot of things.  It can protect against a failed power system, a fire, a network outage, a hardware failure, or the loss of an individual facility.  It cannot guarantee survival when the event is large enough to damage several facilities across the Region.  At that point, “geographically redundant” needs to mean something larger than “another building nearby.”

From a purely engineering standpoint, the obvious technical answer is cross-Region replication.  Keep a copy in Europe, the United States, another Middle Eastern country, or anywhere else sufficiently distant that the same physical event is unlikely to destroy both copies.  But that doesn’t always fly with the people doing the compliance audits.

I Guess Maybe Now They’ll Do 3-2-1 Backups

The gold standard for data protection is “3 copies, 2 different kinds of media, 1 of which is offsite”.

I’m all for the cloud, but if it’s important data and you’re limited to keeping all of your cloud data in one set of buildings…then for pity’s sake, have off-site backups.  You don’t need to have another failover site ready to light up if Amazon goes down for you, but you shouldn’t lose the data.  Replicate it off-site.  I mean, forget even acts of war – what if hackers get into your systems?  Into Amazon’s systems?  What if an employee goes rogue?  What if, what if…

A replicated off-site solution with ransomware protection and disk snapshots (or heaven forbid, tape!) doesn’t mean you can restore service quickly, but it means you won’t lose the data.  In other words, go ahead and say “if we have to go to that dire situation, we’re not promising a recovery time objective (RTO).  But we are promising a Recovery Point Objective (RPO).”

Losing data sucks.

No Comments

    Leave a Reply

    Some notes on commenting on LowEndBox:

    • Do not use LowEndBox for support issues. Go to your hosting provider and issue a ticket there. Coming here saying "my VPS is down, what do I do?!" will only have your comments removed.
    • Akismet is used for spam detection. Some comments may be held temporarily for manual approval.
    • Use <pre>...</pre> to quote the output from your terminal/console, or consider using a pastebin service.

    Your email address will not be published. Required fields are marked *