Ensuring Business Continuity with Warm Standby Disaster Recovery
How a cross-region DR solution brought a secondary region online in 24 minutes.
Oct 21, 2024
case Page image

Summary

An asset management company needed a reliable disaster recovery solution to safeguard operations for millions of customers and meet stringent audit requirements. Bajaj Tech.AI designed and implemented an AWS Warm Standby disaster recovery solution going beyond a standard multi-availability-zone setup to a full cross-region deployment, replicating infrastructure from the primary region (Mumbai) into a secondary DR region (Hyderabad). The entire solution was built and fully validated in under three months, with the final DR drill bringing the secondary region's infrastructure online in just 24 minutes.

Business Challenge

The asset management company faced three interconnected priorities in building its disaster recovery capability:

  • Compliance: The client needed to establish a disaster recovery setup to comply with industry regulations and audits not an optional capability, but a regulatory requirement.
  • Uninterrupted operations: Ensuring seamless service delivery to millions of customers was paramount, even in the event of unforeseen disruptions.
  • Risk mitigation: Protecting against risks posed by natural disasters, cyberattacks, or system failures was a top priority for a company managing customer assets at scale.

For an asset management company, a disaster recovery gap isn't just an operational risk, it's a direct threat to the audit and regulatory standing that the business depends on to operate at all.

Solution Approach

Partnering closely with the client, Bajaj Tech.AI designed and implemented an AWS Warm Standby disaster recovery solution. While Disaster Recovery is commonly available across two availability zones, this solution went further by following an automation-first approach and extending DR across two full AWS regions.

The warm standby approach is an Active-Passive DR model: all infrastructure and necessary AWS services from the primary region (Mumbai) were replicated at a lower scale into the secondary DR region (Hyderabad). This meant Production Environment infrastructure or services normally hosted in Mumbai could be quickly recovered in the event of a disaster by scaling the DR infrastructure (EC2, RDS) in Hyderabad to match production size, performing database failover from the primary into the DR region, and redirecting production traffic to the DR environment through updated Public DNS configuration in the CDN.

In less than 3 months' time, the DR solution was in place and fully validated. The sample architecture diagram can be seen in the figure below.


Image

The project moved through several distinct phases:

  1. Objective alignment: Defined Recovery Time Objective (RTO) and Recovery Point Objective (RPO), ensuring the disaster recovery plan met precise business needs rather than a generic industry benchmark.
  2. Deployment strategy: Orchestrated the deployment of 10+ critical AWS services into the DR region, spanning ECS, ECR, EC2, S3, AWS RDS, Lambda, ACM, Route 53 Hosted Zones, VPC, and Subnets.
  3. Database migration: Collaborated on setting up DB instances and executed migration strategies for various databases, following the optimal approach for each individual database.
  4. Automated infrastructure setup: Employed AWS CloudFormation templates, AWS Backup Service, and Lambda functions to automate infrastructure setup in the DR region, tailoring configurations to regional requirements.
  5. Application services deployment: Ensured smooth deployment by addressing infrastructure configuration challenges, updating AWS SDK versions, and migrating critical application services to the DR region.
  6. Testing excellence: Conducted a meticulous Mock DR drill, resolving identified issues before the actual DR drill to ensure the disaster recovery plan's effectiveness.
  7. Final disaster recovery drill: Conducted the final DR drill, in which the secondary region's infrastructure was brought online in 24 minutes, followed by the migrations.

Running a Mock DR drill before the real one is what caught issues while they were still low-stakes to fix finding a configuration gap during a rehearsal is very different from finding it during an actual disaster.

Business Impact & Results

  • Compliance: Adherence to audit requirements was achieved through the establishment of a robust DR setup and regular DR drills, giving the company a defensible position with regulators.
  • Risk mitigation: Protection against natural and technical disasters affecting entire AWS regions, not just individual availability zones or servers.
  • Data loss prevention: Continuous database replication ensured minimal data loss in the event of a failover.
  • Improved efficiency: Automation reduced the risk of errors and improved recovery time, turning what could have been a manual, error-prone process into a repeatable, tested procedure.

The 24-minute recovery time for the final DR drill is the number that matters most, it's the practical difference between a brief, managed interruption and an extended outage affecting millions of customers.

Key Takeaways

  • Cross-region DR provides protection that multi-availability-zone DR alone cannot, an entire region failing is a real, if rare, risk category
  • An automation-first approach, using CloudFormation and Lambda, is what makes cross-region DR operationally practical rather than a manual, error-prone burden
  • A Mock DR drill before the real drill is what turns disaster recovery from a theoretical plan into a tested, trustworthy capability
  • A 24-minute recovery time for a full secondary region is achievable when infrastructure setup and database failover are both automated end-to-end
  • This kind of resilience engineering pairs directly with the high availability PostgreSQL and SRE practices Bajaj Tech.AI has built elsewhere

Conclusion

For this asset management company, the adoption of a warm standby disaster recovery solution demonstrated a genuine commitment to operational excellence, resilience, and compliance delivered in under three months, with a final DR drill bringing the secondary region online in 24 minutes. This strategic investment enabled the client to safeguard business continuity and mitigate risks in today's dynamic business landscape. Financial institutions facing similar audit and continuity requirements can draw a direct lesson from this engagement: going beyond availability-zone-level DR to full cross-region protection, built on automation rather than manual runbooks, is what makes disaster recovery both compliant and genuinely reliable under pressure.

This kind of cloud resilience work complements Bajaj Tech.AI's infrastructure and API monitoring and modernization engagements, all part of building infrastructure that's resilient, observable, and cost-efficient at once.

Looking to build a disaster recovery solution that meets both compliance and continuity requirements? Connect with our experts to explore the right approach for your organization.

Written by
Vikram Shivtare
Principal
Ensuring Business Continuity with Warm Standby Disaster Recovery | Bajaj Tech.AI