Skip to content
VA2PT.COM
Back to all case studies

AWS MSP · Managed cloud operations

Reelo: high availability and auto scaling on AWS

A multi-AZ, auto-scaling platform that sustains 1,000+ concurrent requests per minute, monitored and managed 24x7.

  • AWS Advanced Tier Services Partner
  • Public case study

About the customer

Reelo is a customer engagement and marketing automation platform designed for restaurants and retail businesses. It helps increase repeat customers through loyalty programs, automated campaigns, and customer insights.

The business challenge

Reelo was experiencing rapid growth in website traffic and needed infrastructure capable of handling more than 1,000 concurrent requests per minute without impacting application performance or availability. The existing environment was not designed to scale automatically during peak traffic, leading to performance bottlenecks and potential service disruptions.

  • Handling over 1,000 concurrent requests per minute during peak business hours.
  • Maintaining low response times and high application performance under heavy traffic.
  • Eliminating single points of failure to ensure high availability.
  • Automatically scaling application servers based on traffic demand.
  • Supporting business growth without requiring manual infrastructure intervention.

Goals and objectives

Design and implement a highly available, secure and auto-scalable cloud infrastructure capable of processing 1,000+ concurrent requests per minute while maintaining optimal performance, minimizing downtime, and reducing operational overhead through automation.

  • Build a highly available infrastructure with no single point of failure.
  • Support 1,000+ concurrent requests per minute with consistent performance.
  • Enable automatic horizontal scaling based on CPU, memory, or request load.
  • Maintain application uptime of 99.9% or higher.

How we solved it on AWS

A cloud-native architecture on AWS in the Mumbai region (ap-south-1), designed so that capacity follows demand and no single component can take the platform down.

  1. 1

    Auto Scaling

    • Deploy the application on an Auto Scaling Group or Kubernetes cluster.
    • Automatically add or remove application instances based on CPU, memory, or request count.
    • Define minimum, desired, and maximum capacity to optimize resource usage.
  2. 2

    Load balancing

    • Use an Application Load Balancer (ALB) or NGINX to distribute incoming traffic across multiple application instances.
    • Perform health checks and route traffic only to healthy instances.
  3. 3

    Monitoring and logging

    • Implement CloudWatch and Datadog for infrastructure and application metrics.
    • Use centralised logging with CloudWatch log groups.
    • Configure alerts for CPU, memory, latency, error rates, and application health.

The AWS toolkit

AWS servicePurpose
Amazon Route 53DNS management and public traffic routing.
Amazon CloudFrontEdge caching and global content delivery in front of the origin.
Amazon VPCIsolated, multi-AZ network in ap-south-1 with segregated public and private subnets, an Internet Gateway and NAT Gateways for controlled egress.
Elastic Load BalancingApplication Load Balancer fronting the production API, performing health checks and distributing traffic across healthy targets in multiple Availability Zones.
Amazon EC2Production API application fleet and the self-managed data layer (ClickHouse cluster, MongoDB replica set, Redis), all deployed without public IP addresses.
EC2 supporting servicesEBS, NAT, Elastic IPs and Auto Scaling Group capacity management for storage, networking and application scaling.
Amazon ECSECS on Fargate for running containerized service tasks in private subnets without managing instances.
AWS LambdaServerless, event-driven processing, VPC-integrated for private resource access.
Amazon RDSManaged MySQL in a Multi-AZ deployment, with the primary writer in one Availability Zone and standby readers in the other two.
Amazon S3Static assets, application data, backups and log retention across 65+ buckets.
Amazon SQSDecoupled, asynchronous message queuing between services.
Amazon SNSNotification fan-out, including CloudWatch alarm delivery.
Amazon EventBridgeScheduled jobs and event-driven rules that trigger downstream processing.
Amazon KinesisReal-time data streaming and ingestion.
AWS Secrets ManagerCentralized storage and rotation of database credentials and application secrets.
AWS IAMLeast-privilege identity, role and permission management across the account.
AWS KMSManaged encryption keys covering EBS volumes, S3 objects and RDS storage at rest.
Amazon GuardDutyContinuous threat detection across account, network and data-plane activity.
AWS Security HubAggregated security posture and compliance findings across the account.
AWS ConfigResource configuration recording and drift detection against desired state.
AWS CloudTrailAccount-wide API audit logging.
VPC Flow LogsNetwork traffic visibility and forensic investigation.
Amazon CloudWatchLogs, metrics, dashboards and alarms underpinning monitoring and alerting.

Key outcomes

  • Successfully sustained 1,000+ concurrent requests per minute during peak business hours with no degradation in response times.
  • Eliminated single points of failure by distributing traffic across multiple Availability Zones via the Application Load Balancer, removing the manual failover dependency that existed in the legacy setup.
  • Reduced manual infrastructure intervention during traffic surges. Auto Scaling Groups now handle scale-out and scale-in automatically based on CPU and memory thresholds, versus previous manual provisioning, cutting manual ops effort by 70%.
  • Improved incident detection and response time through centralized CloudWatch and Datadog monitoring and alerting, reducing mean-time-to-detect (MTTD) for performance issues.
  • Zero customer-facing outages since go-live during peak marketing campaign periods.

Challenges and lessons learned

ChallengeResolution
Running a mixed footprint (EC2, ECS, Lambda) increased architectural complexity and made cost and ownership boundaries less clear across services.Standardized workload placement guidelines (steady-state services on ECS, event-driven components on Lambda) and used consistent tagging for cost allocation and visibility.
Alert volume from CloudWatch and Datadog initially caused alert fatigue for the engineering team, with low-priority notifications drowning out critical ones.Re-tiered alerts by severity (P0–P3), routing high-severity alerts to the communication channel and email for immediate visibility, while informational metrics moved to dashboards reviewed periodically instead of triggering real-time notifications.
Initial Auto Scaling policies were tuned too conservatively, causing brief latency spikes before new instances came online during sudden traffic bursts.Adjusted scaling thresholds and cooldown periods, and introduced predictive and step scaling based on request-count metrics rather than CPU alone, reducing scale-out lag.

Monitoring and anomaly detection

Monitoring and alerting are part of the managed service: 24x7 infrastructure monitoring, alerting and cloud incident management through CloudWatch and Datadog. Detection works in two layers, chosen by whether a fixed limit is meaningful for the signal.

  • Deterministic threshold monitors

    Signals with an agreed service commitment, such as latency, error rate and availability, are monitored against explicit thresholds in Datadog, evaluated on short windows so a breach is detected within a minute. Where a defined remediation exists, the alert routes to an automation webhook that applies it before human involvement is needed, and also to the team Slack channel so the event is visible whether or not remediation succeeds.

  • Statistical and machine-learning anomaly detection

    Where a static threshold is unreliable or too slow, signals are judged against seasonality-aware baselines, so deviation is measured against the expected pattern for that service and period. Daily cloud spend is monitored with AWS Cost Anomaly Detection, which sets a per-service expected range and reports deviation as a percentage of expected. Datadog Watchdog detects deviations in application traces and infrastructure metrics without a monitor being defined in advance.

  • Triage and review

    Detections route to the team Slack channel and to email. Events that do not self-clear are triaged by the on-call engineer. Monitor sensitivity is reviewed monthly and tuned against the false-positive rate.

  • How it is set up for Reelo

    Datadog covers the production service with APM traces, infrastructure metrics and logs, with service naming standardised so traces, metrics and logs resolve to a single service identity. The application runs on EC2 under pm2 process management.

Monitor typeSignalMethodAction on alert
ThresholdHigh p95 latency (above 10 seconds)p95 of request traces on the production service, alert above 10sSlack alert and an automated process-restart webhook
Statistical, seasonality-awareAWS Cost Anomaly DetectionDaily per-service spend against an expected range, deviation reported as a percentage of expectedAlert subscription
Machine learningDatadog WatchdogAutomatic detection across traces and infrastructure metrics without a predefined monitorSurfaced on the monitor timeline
ThresholdCloudWatch alarmsQueue depth, CPU and memory utilisation, and account security events, across 142 alarmsActions enabled on every alarm, routed to SNS
Detection with automated remediation

On 2 August 2026 at 04:30 IST, outside business hours, the p95 latency monitor entered ALERT after request latency exceeded 10 seconds over a one-minute window. A notification went to the team Slack channel and the process-restart webhook applied the defined remediation. The monitor returned to OK 60 seconds after alert, with no engineer intervention and no customer impact. The same pattern recurred and recovered the same way on 4 August 2026.

Statistical anomaly detection on spend

On 4 August 2026, AWS Cost Anomaly Detection reported a deviation of 1,207.89% above the expected range for Amazon Comprehend, a cost impact of only $4.59 over one day. A static budget threshold set at any practical level against monthly spend would not have surfaced a movement that small. The seasonality-aware baseline flagged it as a 1,207% deviation from expected for that service, which is the signal that matters for early detection of misconfiguration or unintended usage.

Related detection

A slow-query condition on the ClickHouse query log, exceeding 30 seconds, was raised on 14 June 2026 at 01:22 IST, outside business hours, and monitoring was configured in response the same day.

Service desk and escalation

P1 and P2 incidents are supported 24x7x365. P3 and P4 incidents and service requests are handled during business hours, 09:00–18:00 IST, Monday to Friday, excluding Indian public holidays. Response and resolution targets are set per severity.

SeverityDefinitionResponseResolutionCoverage
P1Production down or unusable for all users; data loss or active security incident15 minutes4 hours24x7x365
P2Major degradation; core function impaired; no workaround available30 minutes8 hours24x7x365
P3Partial or minor impact; workaround available4 business hours3 business daysBusiness hours
P4Service request, cosmetic issue, how-to query1 business day5 business daysBusiness hours
  • Response targets are met for at least 95% of incidents in each calendar month, and a written root cause analysis is delivered for every P1 within 5 business days.
  • Work is tracked in Asana with a named assignee, comment trail and completion record. Each party designates a primary and an escalation contact. Channels in active use are Asana, email and a shared Slack channel.
  • Out-of-hours engagement in June and July 2026 included a slow-query condition raised and closed on a Sunday, ClickHouse production EBS volume encryption completed overnight ahead of a production release, and a MongoDB failover test run between 06:00 and 07:00 IST to avoid production impact.

Project timeline

Project start
1 May 2026
Go-live in production
24 May 2026
Project end
31 May 2026
Current status
In production

Want your AWS platform run like this?

A 30-minute call with the engineers who run it. We look at your traffic, your uptime targets and what we would put in place first.