disaster recovery testing?

Overview

“Built-in DR capabilities of services and automated deployment are strong enablers for disaster recovery testing”

Disaster recovery testing has traditionally been one of the hardest tasks for IT Ops teams because of high costs, complexity and a low level of automation. In reality, DR testing hardly ever took place. For many crucial (stateful) services, recovery to another data centre or region is out-of-the-box available. Combined with deployment automation of solutions, disaster recovery is now actually achievable at very cost-effective rates.?

Since the platform is crucial for all solutions run by the enterprise, DR testing must be high on the agenda of the team running the platform.?

The platform design should take the DR scenario pertaining to the solution with the highest recovery ambitions into account. At the same time, the standard recovery options can already be set up for a very high level, resistant to data centre outages.?

There are scenarios available for all requirements, up to near zero data loss and zero recovery time. ?

DR testing of solutions involves, among others, connectivity testing, data-integrity testing and application consistency testing. ?

For strict cloud native solutions chaos engineering techniques and tools can be considered.?

The platform should be resilient, but nevertheless, making sure the entire platform can automatically be restored or rebuilt from scratch, is generally a good idea. Since DR testing is usually one of the last activities and easily forgotten (or ignored rather), make sure to add it to the ‘definition of done’.

Activities checklist

Initial:

  • Retrieving the RTO/RPO requirements and CIA classifications ?
  • Drafting a DR test plan for the platform and solutions?
  • Creating DR patterns for a variety of application architectures?
  • Automating the deployment of platform services, solutions and their infrastructure services ?
  • Automating the deployment of shared services?
  • Creating DR test-plans for shared services?
  • Adding DR testing to back log?
  • Determining the recovery region and creating the basic infrastructure

Recurring:

  • Making DR a chapter in all solution designs ?
  • Retrieving the RTO/RPO requirements and CIA classifications ?
  • Assessing the DR requirements of new solutions and impact on design of the platform if needed.?
  • Providing synchronisation and recovery of data?
  • Testing DR platform services and solutions

RASCI

cloud consultantinformedtransformation consultant
cloud architectresponsiblecloud partners
cloud security specialistinformedDevOps teaminformed
cloud developerinformedbusiness stakeholderaccountable
cloud engineersupportingarchitecture
cloud analystinformedsecurity
product owner CCoEinformedfinance
managementprocurement
Scroll to Top