Arnie, Danny and the Problem with Our OSS Digital Twins

We invest heavily in production environments (aka Arnie) and in reactive network fault-fix measures. Most OSS teams think of non-prod (aka Danny) as somewhere to test releases and patches.

But if our non-prod environments were true digital twins of production, their value could be far greater.

They could become environments where we simulate proposed changes, stress-test operational scenarios, validate interactions across domains and deliberately break things without breaking anything in production (ie anything that customers directly depend on).

Most importantly, they could drastically increase the number of incidents we prevent rather than fix.

But let’s take a closer look at other benefits of building non-prod environments that are a closer replica of prod environments typically are today.

.

A Twin in Name Only

Production environments (equivalent to Arnold Schwarzenegger in the modified film photo above) are impressive – large, complex, highly interconnected and constantly changing. They need to be because they carry and manage live customer services on the global telecommunications network.

Our non-production environments (equivalent to Danny DeVito) are often far less impressive / sophisticated. They tend to be smaller. They contain subsets of data. Integrations are stubbed out or missing completely. Network domains are represented differently. Data is incomplete, often stale and not even reflective of the production data it’s supposed to mirror. Interfaces behave differently. Some downstream systems simply aren’t there.

We imply that non-prod is a representation of production, but sometimes the resemblance is about as convincing as Arnie and Danny as Twins in the 1988 movie of the same name that some of you might remember.

A “Danny Environment” might be perfectly adequate if we’re only asking limited questions like, “Does this software release / patch / upgrade install correctly and basically remain working?

But I hope you’d agree with me that that’s an incredibly limited use of something that could be much more valuable.

What if we started treating our Danny environment as an offline simulation for our entire OSS and network ecosystem? What could we ask it to do in a safer way than running these experiments on production (which the change board would never allow)?

.

What if we Could Simulate Before we Operate?

A sufficiently representative digital twin gives us somewhere to experiment before touching anything in production.

That opens up possibilities well beyond traditional release testing:

  1. Prevent incidents before they happen – simulate proposed changes and failure scenarios, identify unexpected consequences and address them in change planning before customers are exposed
  2. Test across domains rather than inside silos – validate complete workflows spanning domains and systems (as this previous article indicates, most Sev 1 outages these days arise from combinational / cascading events that are never adequately tested. Intra-domain testing tends to be far more thorough than interoperability / cross-domain testing)
  3. Rehearse major network changes – practice complex migrations, upgrades, topology changes and transformation activities in the Danny environment (non-prod), then before / during / after carrying them out for real in the Arnie environment (prod)
  4. Replay production incidents – recreate the conditions surrounding a major incident and forensically explore alternative responses, fixes and preventative controls
  5. Stress-test automation safely – expose closed-loop automation to edge cases, conflicting inputs, cascading effects, load testing and unusual failure combinations before granting it approval or greater authority in production
  6. Run a multitude of “what if?” scenarios – there are generally so many “what if?” questions running through my inquisitive mind, such as:- what happens if a resource fails, a dependency disappears, capacity changes or a migration sequence is altered, then observe the outcome without touching a customer-facing Arnie environment
  7. Shift investment left – It’s not uncommon for ~50% of network outages to arise during (and as a result of) change windows in production. Instead, a holistic digital twin. This leaves the potential of using realistic simulation, validation and cross-domain / interoperability testing to move more operational investment towards prevention, rather than relying predominantly on responding to incidents after they’ve occurred (ie detection, incident response and fault-fix)
  8. Find hidden dependencies – expose relationships and single points of failure (SPoF) between systems, services and resources that are difficult to discover when testing each application independently
  9. Train people against realistic conditions – allow trainee operators and engineers to experience life-like major failures, unusual alarm patterns and complex recovery situations in Danny environments where mistakes have no real consequences or customer impact
  10. Turn non-prod into an operational asset – create an environment that supports engineering, operations, transformation, training, automation and planning rather than existing primarily as a software test platform.

The important shift is from testing whether a standalone application works to simulating whether the entire operational ecosystem works.

.

The Biggest Problems Often Live Between Systems

This becomes particularly important in OSS because some of our hardest problems don’t live inside individual applications.

As mentioned in this article earlier, an inventory platform can pass its tests. An assurance platform can pass its tests. An orchestrator can pass its tests. The network domain controllers can all pass theirs. And the end-to-end service can still fail, in part because doing interoperability testing fell into the “too hard” category. It’s deemed that there’s too much coordination and setup required between stakeholders.

Many production issues emerge from the interactions between vendors, domains, APIs, data models, timing assumptions, orchestration flows and operational processes.

Unfortunately, those boundaries are often exactly what simplified non-production environments reproduce least accurately.

Passing ten isolated application tests is not the same as passing one realistic end-to-end test.

A true digital twin gives us the ability to validate those interactions across domains. It gives us somewhere to test interoperability at a scale and level of complexity much closer to reality.

That is potentially far more valuable than simply proving that each component works independently.

.

Shift Left – From Fault-Fix to Fault-Prevention

There is another, more philosophical, question to ask to determine the value of a true digital twin, “where do we want to spend our operational effort?”

The ITIL diagram below provides a useful lens on answering that question. Most people will focus on the status quo, which is to spend almost all of the budget on a Service Desk, Incident Management, Problem Management and Event Management as well as all the NOC / operations processes and systems that are built around that answer. Unfortunately, all of those are invoked after the network / service has failed or degraded (ie too late).

Moving from right to left (from red to green), we expend greater effort on disciplines that help us design, validate, test and protect services, which prevent problems from occurring (green boxes below), as opposed to disciplines that detect, diagnose and resolve problems after they’re already affecting production (red boxes).

Figure: A simplified view of the opportunity to shift operational effort left – from reactive fault-fix (red boxes) towards greater prevention, simulation and testing / validation (green boxes)

A sophisticated digital twin will probably never eliminate the need for reactive operations (red boxes). However, it should give us the opportunity to need them less often and therefore invest in them less, shifting funding to green boxes.

But it totally makes sense why we do invest so heavily in the right-hand side of this picture. Service Desks, Event Management, Incident Management and Problem Management are essential because, when something fails in production, we need to detect it quickly, understand what happened and restore service.

But by the time these “red” processes are involved, the problem has already had an impact on production (and possibly customers).

.

The Business Case for More Complete Twins

A highly production-like “Danny” environment gives us our best opportunity to move more of that effort to the left (from red to green). If Danny genuinely mirrors Arnie, we can invest more heavily in Service Validation & Testing, Change Management, Availability Management, Capacity Management, continuity and engineering activities, thus better validating a change before it ever encounters a customer.

Note that that shift becomes even more powerful if/when the digital twin represents the whole OSS and network ecosystem rather than individual applications. That would allow us to perform realistic cross-domain and interoperability testing, simulate unusual combinations of events and deliberately search for the dependencies that would otherwise only reveal themselves during an incident.

In other words, the goal isn’t simply to become faster at fixing faults. It’s to increase the percentage of faults that never make it into production in the first place, obviously.

That flipped perspective changes the entire business case for non-production environments to look more like Arnies and less like Dannys.

 

PS. Categories of digital twin/triplet and reality twin/triplet can be found here.

If this article was helpful, subscribe to the Passionate About OSS Blog to get each new post sent directly to your inbox. 100% free of charge and free of spam.

Our Solutions

Share:

Most Recent Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.