Infrastructure diversity and disaster avoidance

Server racks in a data centre, still being populated with machinery - by Photo by

I’ve always had a simple belief – you can’t depend on any one service, supplier or machine. You always need a level of redundancy. Which is fine… except it can get rather expensive.

But how expensive is not managing it?

I’d like to start with a few links out to important stories in this space:

Fastly failure brings down Amazon, Reddit, The Guardian and Gov.uk
Jaguar Land Rover cyber attack cost the UK economy £1.9bn
Marks and Spencer cyber attack wipes out profits

The first failure just made some websites unavailable for a day and was down to an error. The next two, however, both had factors in common and cost a huge amount of money. There are many more instances, some of which are unreported.

Every one of those organisations had professional approaches to cyber security and resilience. Like a bank, they checked for all the known vulnerabilities. But they made what I believe are some key strategic failures. Some very human in nature, and some very technical in nature.

#1 The human element

Humans are corruptible. You need three things for corruption to happen: Opportunity (e.g. access to systems, bank accounts), Motivation (financial stress, hunger, addictions, workplace pressure), and Rationalisation (I need it, the system is unjust, the company owes me, everyone does it, the directors take too much money).

In the UK, for instance, a lot of young people are finding life a little hopeless – housing is super expensive and job opportunities have thinned out. At JLR this, coupled with Tata marking its own cybersecurity homework using Tata Consultancy Services caused a real opportunity for the attackers.

M&S, guess what… used Tata Consultancy Services also. So this created the opportunity.

The rationalisation was likely the need that these young men who were arrested in the UK felt – that in a world where the rich just seem to get richer whilst most others have barely felt a change to their wealth in two decades is unfair. An increasingly unjust socio-economic background will feed this.

And then there’s the motivational pressure itself. People want to have the nice things they see others having. Many are following influencers on social media with carefully groomed Instagram accounts. This is a perfect storm for cybersecurity to be harder than ever.

I’ve always had a few rules in business to help stave off fraud, and it’s worked so far:

  1. Pay people enough that they can afford decent housing, healthcare and transport.
  2. Minimise the difference in incomes between the top and bottom of the organisation.
  3. Have a reasonable set of checks and counterbalances.
  4. Invest in training against fraud.
  5. Ensure a good cultural fit between team members so they understand each other.

#2 The technical element

This, in a way, is harder than the human element. Humans don’t have much predictability, but the majority are honest and if you look after them they tend to look after you. Machines, however, are cold, hard things made of metal, plastic and stone. And computers and the software they run will always follow the rules given, not the rules intended.

That makes them vulnerable. Whether they are deterministic machines like classical computing approaches, or stochastic like the current LLM type AIs. It’s very easy to trick an AI with prompt injection, for example, with some dire consequences. Coercing an LLM doesn’t need a fraud triangle in place – it just needs an understanding of how the LLM works, and that guardrails for them are just like guardrails in real life – they’re easy to climb over.

Consequently we should always always assume that the systems in place are inherently vulnerable. And that leads neatly to why we need redundancy and diversity.

As any ecologist will tell you, monocultures are highly vulnerable to infection. This is a common problem with bananas – most bananas we eat are essentially all clones of each other – they are incredibly similar to one another. A previous banana variety, the Gros Michel, is no longer widely available because a fungus found a way to ruin yields. So currently most of us in Europe eat Cavendish bananas – and they too are starting to get an infection. Well, your fleet of Windows or Mac computers are, similarly, clones of each other.

Consequently, if someone gets into your systems there’s a very real risk that an unpatched vulnerability could spread to everything. This could come through where you store your codebase through to your file servers. Even printers can be used as an attack vector.

I’ve always felt that the two pathways to avoiding these problems are multiple:

  1. Keep critical systems available across multiple platforms.
  2. Redundancy, redundancy, redundancy.
  3. Critical data should be immutable.
  4. Have to identity providers available to your teams.

Very rarely are any of these approaches in place. For example, a publication may have a huge archive of content that is important to their business revenues – but it’s almost always kept in forms that are mutable. At least consider taking a weekly backup burned onto an old school DVD ROM, for example, that cannot easily be changed. Database engines and software can be written or configured with immutable properties too, so that only a handful of people with the very highest privileges can change the data.

Conclusion

It’s impossible to give a simple guide to a complex problem, but I believe these rather basic principles can be very powerful. Imagine if your Windows machines were hacked but everything can still be run from Macs? Or if your website goes down because of a major vulnerability in a popular WordPress plugin compromises it, could you fail over to an isolated version of the site that is very cut down but still maintains your business and income stream?

Of course, like all insurance, most people don’t value it that much until they need it… and many people feel that insurance that’s never claimed upon was clearly a waste of money. I’ve always struggled to persuade customers to take these matters seriously and to spend enough – enough on developers, enough on security tools, and enough on redundancy. In the end I spent a large amount of money developing an off-cloud backups platform which was useful insurance for me when selling to customers but it was hard to recoup the costs as nobody would pay for it!

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

pentulo.com
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.