Optimising Cloud Spend While Scaling Up Business
How a BFSI company cut cloud costs 25-50% across 50+ AWS services.
Dec 24, 2024
case Page image

Summary

A leading BFSI company in India, managing a complex hybrid technology landscape across over 50 AWS services, struggled to control cloud costs while scaling operations. Manual processes, reliance on licensed tools, and a lack of automation compounded the problem, making it increasingly difficult to control cloud spend without a structured, dedicated approach.

Bajaj Tech.AI formed a dedicated Site Reliability Engineering (SRE) team spanning AWS, DevSecOps, security, and infrastructure as code, applying a structured, checklist-driven approach across pricing models, auto-scaling, automation, and governance. The result: a 25-50% reduction in average monthly cloud costs, alongside far greater visibility and predictability in cost forecasting.

Business Challenge

The company's hybrid technology landscape comprised multiple applications with different technologies, operating systems, and availability requirements. With multiple lower environments shared across teams and manual operations governing much of the infrastructure, cost optimization became increasingly complex. The lack of automation made it difficult to control cloud costs effectively, and heavy reliance on licensed tools added further to operational expenses together making it genuinely challenging to improve margins and support business growth.

At 50+ AWS services and multiple shared environments, the challenge wasn't identifying individual cost-saving opportunities, it was building a structured enough approach to find and act on all of them systematically, rather than chasing the most visible ones.

Solution Approach

A Site Reliability Engineering (SRE) team with expertise in AWS, DevSecOps, security, and infrastructure as code was formed. The team adopted a structured approach, categorizing the landscape and analyzing each category using standard checklists. The cost optimization project was divided into several categories:

  • Pricing models: Workloads were grouped into lower and higher environment categories based on usage patterns and requirements, enabling the implementation of appropriate pricing models on-demand, reserved instances, and spot instances matched to each workload's actual behavior.
  • Auto-scaling policies: Policies based on metrics like CPU, memory, RPM, and peak traffic patterns were defined to ensure efficient resource utilization and cost savings, rather than static over-provisioning.
  • Automations: Categories were created for repetitive manual tasks, DevOps tasks, cost optimization tasks, and governance and reporting tasks, with various automation stacks built around them, including a self-service portal for starting or stopping lower environments, a DIY portal for monitoring ECS service or server health and starting/stopping application servers with a given count, governance reports like EC2 inventory, housekeeping tasks like auto-deletion of snapshots or AMIs, utilities to generate heap dumps, and various Jenkins-based automations.
  • Tools development: Tools for patching frameworks, AWS Spot frameworks, auto-healing of incidents, and process frameworks for monitoring cost and infrastructure capacity were developed to automate tasks and improve efficiency across the board.

The structure here matters as much as the individual tactics categorizing workloads before applying pricing models, and building reusable automation stacks rather than one-off scripts, is what made the approach scale across 50+ services instead of being limited to a handful of easy wins.

Business Impact & Results

The implementation of this comprehensive solution resulted in a significant reduction in average monthly cloud costs, with savings ranging from 25% to 50%. The company gained visibility and predictability in cost forecasting, tracking, and optimization, allowing them to control cloud spend effectively while scaling up business operations.

By adopting a holistic approach to cost optimization and automation, the company was able to enhance margins and achieve greater operational efficiency in managing its hybrid technology landscape turning what had been a source of unpredictable expense into a controlled, forecastable line item.

Key Takeaways

  • A checklist-driven, categorized approach scales cost optimization across dozens of services far better than ad hoc tactical fixes
  • Matching pricing models (on-demand, reserved, spot) to actual workload behavior is more effective than applying one pricing strategy universally
  • Self-service and DIY portals reduce the ongoing manual burden that would otherwise erode cost-optimization gains over time
  • A 25-50% cost reduction range reflects how much waste typically accumulates in a hybrid landscape without dedicated SRE ownership
  • This kind of structured optimization complements the cloud database cost optimization and festive traffic surge cost work Bajaj Tech.AI has done elsewhere

Conclusion

For this BFSI company, forming a dedicated SRE team and applying a structured, checklist-driven approach to cost optimization delivered a 25-50% reduction in average monthly cloud costs while simultaneously improving visibility and predictability for future planning. Organizations managing a large, hybrid AWS footprint can draw a direct lesson from this engagement: sustainable cost optimization requires ongoing ownership and automation, not a one-time audit, since manual processes and licensed-tool sprawl tend to creep back in without dedicated attention.

This kind of cloud cost discipline pairs naturally with Bajaj Tech.AI's infrastructure and API monitoring and disaster recovery work, all part of the same SRE discipline of making infrastructure both cost-efficient and resilient.

Looking to bring predictability and savings to your cloud spend? Connect with our experts to explore the right cost optimization approach for your organization.

Written by
Vikram Shivtare
Principal
Optimising Cloud Spend While Scaling Up Business | Bajaj Tech.AI