
AWS MSP · Managed cloud operations
Reelo: high availability and auto scaling on AWS
A multi-AZ, auto-scaling platform that sustains 1,000+ concurrent requests per minute, monitored and managed 24x7.
- AWS Advanced Tier Services Partner
- Public case study
About the customer
Reelo is a customer engagement and marketing automation platform designed for restaurants and retail businesses. It helps increase repeat customers through loyalty programs, automated campaigns, and customer insights.
The business challenge
Reelo was experiencing rapid growth in website traffic and needed infrastructure capable of handling more than 1,000 concurrent requests per minute without impacting application performance or availability. The existing environment was not designed to scale automatically during peak traffic, leading to performance bottlenecks and potential service disruptions.
- Handling over 1,000 concurrent requests per minute during peak business hours.
- Maintaining low response times and high application performance under heavy traffic.
- Eliminating single points of failure to ensure high availability.
- Automatically scaling application servers based on traffic demand.
- Supporting business growth without requiring manual infrastructure intervention.
Goals and objectives
Design and implement a highly available, secure and auto-scalable cloud infrastructure capable of processing 1,000+ concurrent requests per minute while maintaining optimal performance, minimizing downtime, and reducing operational overhead through automation.
- Build a highly available infrastructure with no single point of failure.
- Support 1,000+ concurrent requests per minute with consistent performance.
- Enable automatic horizontal scaling based on CPU, memory, or request load.
- Maintain application uptime of 99.9% or higher.
How we solved it on AWS
A cloud-native architecture on AWS in the Mumbai region (ap-south-1), designed so that capacity follows demand and no single component can take the platform down.
- 1
Auto Scaling
- Deploy the application on an Auto Scaling Group or Kubernetes cluster.
- Automatically add or remove application instances based on CPU, memory, or request count.
- Define minimum, desired, and maximum capacity to optimize resource usage.
- 2
Load balancing
- Use an Application Load Balancer (ALB) or NGINX to distribute incoming traffic across multiple application instances.
- Perform health checks and route traffic only to healthy instances.
- 3
Monitoring and logging
- Implement CloudWatch and Datadog for infrastructure and application metrics.
- Use centralised logging with CloudWatch log groups.
- Configure alerts for CPU, memory, latency, error rates, and application health.
The AWS toolkit
| AWS service | Purpose |
|---|---|
| Amazon Route 53 | DNS management and public traffic routing. |
| Amazon CloudFront | Edge caching and global content delivery in front of the origin. |
| Amazon VPC | Isolated, multi-AZ network in ap-south-1 with segregated public and private subnets, an Internet Gateway and NAT Gateways for controlled egress. |
| Elastic Load Balancing | Application Load Balancer fronting the production API, performing health checks and distributing traffic across healthy targets in multiple Availability Zones. |
| Amazon EC2 | Production API application fleet and the self-managed data layer (ClickHouse cluster, MongoDB replica set, Redis), all deployed without public IP addresses. |
| EC2 supporting services | EBS, NAT, Elastic IPs and Auto Scaling Group capacity management for storage, networking and application scaling. |
| Amazon ECS | ECS on Fargate for running containerized service tasks in private subnets without managing instances. |
| AWS Lambda | Serverless, event-driven processing, VPC-integrated for private resource access. |
| Amazon RDS | Managed MySQL in a Multi-AZ deployment, with the primary writer in one Availability Zone and standby readers in the other two. |
| Amazon S3 | Static assets, application data, backups and log retention across 65+ buckets. |
| Amazon SQS | Decoupled, asynchronous message queuing between services. |
| Amazon SNS | Notification fan-out, including CloudWatch alarm delivery. |
| Amazon EventBridge | Scheduled jobs and event-driven rules that trigger downstream processing. |
| Amazon Kinesis | Real-time data streaming and ingestion. |
| AWS Secrets Manager | Centralized storage and rotation of database credentials and application secrets. |
| AWS IAM | Least-privilege identity, role and permission management across the account. |
| AWS KMS | Managed encryption keys covering EBS volumes, S3 objects and RDS storage at rest. |
| Amazon GuardDuty | Continuous threat detection across account, network and data-plane activity. |
| AWS Security Hub | Aggregated security posture and compliance findings across the account. |
| AWS Config | Resource configuration recording and drift detection against desired state. |
| AWS CloudTrail | Account-wide API audit logging. |
| VPC Flow Logs | Network traffic visibility and forensic investigation. |
| Amazon CloudWatch | Logs, metrics, dashboards and alarms underpinning monitoring and alerting. |
Key outcomes
- Successfully sustained 1,000+ concurrent requests per minute during peak business hours with no degradation in response times.
- Eliminated single points of failure by distributing traffic across multiple Availability Zones via the Application Load Balancer, removing the manual failover dependency that existed in the legacy setup.
- Reduced manual infrastructure intervention during traffic surges. Auto Scaling Groups now handle scale-out and scale-in automatically based on CPU and memory thresholds, versus previous manual provisioning, cutting manual ops effort by 70%.
- Improved incident detection and response time through centralized CloudWatch and Datadog monitoring and alerting, reducing mean-time-to-detect (MTTD) for performance issues.
- Zero customer-facing outages since go-live during peak marketing campaign periods.
Challenges and lessons learned
| Challenge | Resolution |
|---|---|
| Running a mixed footprint (EC2, ECS, Lambda) increased architectural complexity and made cost and ownership boundaries less clear across services. | Standardized workload placement guidelines (steady-state services on ECS, event-driven components on Lambda) and used consistent tagging for cost allocation and visibility. |
| Alert volume from CloudWatch and Datadog initially caused alert fatigue for the engineering team, with low-priority notifications drowning out critical ones. | Re-tiered alerts by severity (P0–P3), routing high-severity alerts to the communication channel and email for immediate visibility, while informational metrics moved to dashboards reviewed periodically instead of triggering real-time notifications. |
| Initial Auto Scaling policies were tuned too conservatively, causing brief latency spikes before new instances came online during sudden traffic bursts. | Adjusted scaling thresholds and cooldown periods, and introduced predictive and step scaling based on request-count metrics rather than CPU alone, reducing scale-out lag. |
Monitoring and anomaly detection
Monitoring and alerting are part of the managed service: 24x7 infrastructure monitoring, alerting and cloud incident management through CloudWatch and Datadog. Detection works in two layers, chosen by whether a fixed limit is meaningful for the signal.
Deterministic threshold monitors
Signals with an agreed service commitment, such as latency, error rate and availability, are monitored against explicit thresholds in Datadog, evaluated on short windows so a breach is detected within a minute. Where a defined remediation exists, the alert routes to an automation webhook that applies it before human involvement is needed, and also to the team Slack channel so the event is visible whether or not remediation succeeds.
Statistical and machine-learning anomaly detection
Where a static threshold is unreliable or too slow, signals are judged against seasonality-aware baselines, so deviation is measured against the expected pattern for that service and period. Daily cloud spend is monitored with AWS Cost Anomaly Detection, which sets a per-service expected range and reports deviation as a percentage of expected. Datadog Watchdog detects deviations in application traces and infrastructure metrics without a monitor being defined in advance.
Triage and review
Detections route to the team Slack channel and to email. Events that do not self-clear are triaged by the on-call engineer. Monitor sensitivity is reviewed monthly and tuned against the false-positive rate.
How it is set up for Reelo
Datadog covers the production service with APM traces, infrastructure metrics and logs, with service naming standardised so traces, metrics and logs resolve to a single service identity. The application runs on EC2 under pm2 process management.
| Monitor type | Signal | Method | Action on alert |
|---|---|---|---|
| Threshold | High p95 latency (above 10 seconds) | p95 of request traces on the production service, alert above 10s | Slack alert and an automated process-restart webhook |
| Statistical, seasonality-aware | AWS Cost Anomaly Detection | Daily per-service spend against an expected range, deviation reported as a percentage of expected | Alert subscription |
| Machine learning | Datadog Watchdog | Automatic detection across traces and infrastructure metrics without a predefined monitor | Surfaced on the monitor timeline |
| Threshold | CloudWatch alarms | Queue depth, CPU and memory utilisation, and account security events, across 142 alarms | Actions enabled on every alarm, routed to SNS |
On 2 August 2026 at 04:30 IST, outside business hours, the p95 latency monitor entered ALERT after request latency exceeded 10 seconds over a one-minute window. A notification went to the team Slack channel and the process-restart webhook applied the defined remediation. The monitor returned to OK 60 seconds after alert, with no engineer intervention and no customer impact. The same pattern recurred and recovered the same way on 4 August 2026.
On 4 August 2026, AWS Cost Anomaly Detection reported a deviation of 1,207.89% above the expected range for Amazon Comprehend, a cost impact of only $4.59 over one day. A static budget threshold set at any practical level against monthly spend would not have surfaced a movement that small. The seasonality-aware baseline flagged it as a 1,207% deviation from expected for that service, which is the signal that matters for early detection of misconfiguration or unintended usage.
A slow-query condition on the ClickHouse query log, exceeding 30 seconds, was raised on 14 June 2026 at 01:22 IST, outside business hours, and monitoring was configured in response the same day.
Service desk and escalation
P1 and P2 incidents are supported 24x7x365. P3 and P4 incidents and service requests are handled during business hours, 09:00–18:00 IST, Monday to Friday, excluding Indian public holidays. Response and resolution targets are set per severity.
| Severity | Definition | Response | Resolution | Coverage |
|---|---|---|---|---|
| P1 | Production down or unusable for all users; data loss or active security incident | 15 minutes | 4 hours | 24x7x365 |
| P2 | Major degradation; core function impaired; no workaround available | 30 minutes | 8 hours | 24x7x365 |
| P3 | Partial or minor impact; workaround available | 4 business hours | 3 business days | Business hours |
| P4 | Service request, cosmetic issue, how-to query | 1 business day | 5 business days | Business hours |
- Response targets are met for at least 95% of incidents in each calendar month, and a written root cause analysis is delivered for every P1 within 5 business days.
- Work is tracked in Asana with a named assignee, comment trail and completion record. Each party designates a primary and an escalation contact. Channels in active use are Asana, email and a shared Slack channel.
- Out-of-hours engagement in June and July 2026 included a slow-query condition raised and closed on a Sunday, ClickHouse production EBS volume encryption completed overnight ahead of a production release, and a MongoDB failover test run between 06:00 and 07:00 IST to avoid production impact.
Project timeline
- Project start
- 1 May 2026
- Go-live in production
- 24 May 2026
- Project end
- 31 May 2026
- Current status
- In production
Want your AWS platform run like this?
A 30-minute call with the engineers who run it. We look at your traffic, your uptime targets and what we would put in place first.