GenAI infrastructure on AWS, Azure and GCP: a buyer's guide for Indian startups
By VA2PT Team, . 9 min read

Choosing where to run your GenAI workloads is less about which model is smartest this month and more about four boring questions: where does the data go, how do you connect privately, what will it cost at scale, and what happens when you need to switch. This guide answers those for the three managed platforms most Indian startups end up choosing between: Amazon Bedrock, Azure OpenAI Service and Google Vertex AI, and then covers the point where self-hosting on GPUs starts to make sense.
We use one scenario throughout: a Chennai-based logistics SaaS building a customer support assistant with retrieval over shipment documents and a nightly batch that classifies 500,000 delivery notes.
The regional question comes first
For Indian companies the first filter is not price. It is whether the model you want is served from an Indian region, because that determines whether prompts leave the country (relevant under the DPDPA and sector regulation) and what latency your users get.
The honest state of play:
- AWS: the Mumbai region (
ap-south-1) is the primary Indian region, with Hyderabad (ap-south-2) as a second. Bedrock is available in Mumbai, but model availability differs by region; the newest models often land in US regions first. Bedrock also supports cross-region inference profiles, which can route requests to other regions; check that you have not enabled one that sends data outside India. - Azure: Azure has Central India (Pune), South India (Chennai) and West India (Mumbai). Azure OpenAI model availability by region changes frequently; some deployment types (for example, "Global" deployments) explicitly process in any region worldwide, while "Data Zone" and regional deployments constrain where processing happens. If residency matters, choose a regional deployment and read the deployment type carefully.
- GCP: Google Cloud has Mumbai (
asia-south1) and Delhi (asia-south2). Vertex AI's Gemini models have a documented list of supported regions, and Google also offers a "global" endpoint that does not guarantee a location. As with Azure, pick a regional endpoint if you need one.
Because availability shifts every quarter, the rule is: decide the required region, then pick from the models served there, not the other way round. Put the region in the model inventory and alert on any call that leaves it.
Private connectivity: do not call the model over the internet
All three platforms let you keep traffic on the provider's backbone:
| Platform | Private access mechanism | What to check |
|---|---|---|
| Bedrock | VPC interface endpoint (com.amazonaws.<region>.bedrock-runtime) |
Endpoint policy limits which model ARNs can be invoked |
| Azure OpenAI | Private Endpoint on the Azure OpenAI resource; disable public network access | Private DNS zone for privatelink.openai.azure.com |
| Vertex AI | Private Service Connect or VPC Service Controls perimeter | Perimeter includes the storage bucket used for RAG documents |
Terraform for the Bedrock endpoint, since it is the one most teams forget:
resource "aws_vpc_endpoint" "bedrock_runtime" {
vpc_id = var.vpc_id
service_name = "com.amazonaws.ap-south-1.bedrock-runtime"
vpc_endpoint_type = "Interface"
subnet_ids = var.private_subnet_ids
security_group_ids = [aws_security_group.bedrock_endpoint.id]
private_dns_enabled = true
policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Effect = "Allow"
Principal = "*"
Action = ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"]
Resource = "arn:aws:bedrock:ap-south-1::foundation-model/*"
}]
})
}
Pair the endpoint with an IAM policy on the calling role that lists the exact model IDs allowed. That is your allowlist, and it is what stops a developer from quietly switching to a model in another region.
Token cost control: the maths you must do before launch
The support assistant will be used roughly 40,000 times a day. Each request carries a system prompt (800 tokens), retrieved context (2,500 tokens), the user's question (100 tokens) and produces a 300-token answer. That is about 3,400 input and 300 output tokens per request, 136 million input tokens and 12 million output tokens a day.
At list prices in the range of a few dollars per million input tokens for mid-tier models, that is a monthly bill in the low thousands of dollars for input alone, and it scales linearly with retrieved context. Three controls cut it:
- Prompt caching. Bedrock, Azure OpenAI and Vertex AI all offer some form of cached input pricing for repeated prefixes. Put the system prompt and static instructions first so the cache hits. On a stable prefix of 800 tokens this alone removes a quarter of input cost.
- Retrieval budget. Cap retrieved context at what the answer needs: five chunks of 300 tokens beats twelve chunks of 500. Measure answer quality against context size with your evaluation suite; the curve flattens fast.
- Model routing. Classify requests first with a small, cheap model (or a rule) and send only the hard 20% to the large model. For the nightly classification batch, use batch inference APIs, which all three providers offer at a discount for non-urgent work.
Track cost per resolved conversation, not cost per token. If the assistant deflects a ticket that would have cost an agent 6 minutes, the token spend is a rounding error; if it doesn't, no discount will save it.
RAG reference architectures on each platform
The pattern is the same everywhere: ingest documents → chunk → embed → store vectors → at query time retrieve → build prompt → call model → filter output. The managed pieces differ.
AWS
S3 (documents) → Bedrock Knowledge Bases (chunking + embeddings with Titan or Cohere)
→ OpenSearch Serverless or Aurora pgvector (vector store)
API Gateway → Lambda → Bedrock Agents / RetrieveAndGenerate → Guardrails for Bedrock → response
Use Guardrails to enforce topic denial and PII masking on both input and output; it is the cheapest place to put that control.
Azure
Blob Storage → Azure AI Search (integrated vectorisation, hybrid search)
App Service / Container Apps → Azure OpenAI (chat + embeddings) → Content Safety → response
Azure AI Search's hybrid (keyword + vector) retrieval with semantic ranker is strong for document-heavy support; it is the reason many teams pick this stack even when their app runs elsewhere.
GCP
Cloud Storage → Vertex AI Search (or Vector Search) → Cloud Run → Gemini on Vertex AI → response
Vertex AI Search does chunking, embedding and retrieval as a managed service with grounding metadata in the response, which shortens the path to "show the source".
Whichever you pick, own three things yourself: the chunking strategy (test it, do not accept defaults), the evaluation suite, and the metadata on each chunk (tenant ID, document ID, principal ID) so you can filter by tenant and delete on request.
When self-hosting on GPUs starts to make sense
Self-hosting an open-weights model on your own GPUs is not cheaper by default. It becomes cheaper when three conditions hold:
- High, steady volume. A single NVIDIA L4 or A10G class GPU serving a 7B to 8B parameter model with vLLM handles hundreds of thousands of short requests a day. At that volume, per-token API pricing starts to exceed the hourly cost of the instance.
- Latency or residency you cannot get from the API. If the model you need is not served from India and the data cannot leave, self-hosting in Mumbai is the only option.
- You have someone to run it. GPU drivers, model updates, autoscaling and observability are real work. If your platform team is two people, the managed API is the right answer.
Break-even sketch for the nightly classification batch: 500,000 notes at 400 tokens each is 200 million input tokens per night. On a per-token API at a low-tier price, that is a few hundred dollars per month at batch discounts. A single on-demand GPU instance running for the two hours the batch needs costs a few dollars per night. The self-hosted path wins here because the workload is bursty, predictable and tolerant of a small model. The support assistant, with its unpredictable traffic and need for a stronger model, stays on the managed API.
When you do self-host, use the managed Kubernetes services (EKS, AKS, GKE) with GPU node pools, the NVIDIA device plugin, and vLLM or TGI behind an autoscaler. We cover that setup in detail in a separate article on Kubernetes and AI workloads.
FinOps tagging for AI spend
AI spend is invisible in most cloud bills because it hides inside "Bedrock", "Cognitive Services" or "Vertex AI" line items with no team attribution. Fix that on day one:
- Tag every AI resource (endpoints, knowledge bases, vector stores, GPU node pools) with
team,featureandenv. Enforce with a tag policy or an Azure Policy or GCP Organization Policy constraint. - Use application inference profiles on Bedrock, which let you tag invocations by application so that cost allocation reports split model usage by feature.
- On Azure, create one Azure OpenAI resource per feature or per team, not one shared resource, so cost analysis can split by resource.
- On GCP, use separate projects per feature for AI services; labels on the project flow through to billing exports.
- Export token counts from your application logs to your metrics store and plot cost per feature next to usage. Cloud billing lags by a day; your logs do not.
Set a budget alert per feature at 120% of the forecast. AI features fail expensively and quietly: a retry loop or a runaway agent can double the bill overnight.
A decision table
| You need | Choose | Because |
|---|---|---|
| Widest managed model choice with strong guardrails and India region | Bedrock in ap-south-1 |
Multiple model families behind one API, Guardrails, Knowledge Bases |
| GPT-family models and Microsoft 365 or Entra integration | Azure OpenAI, regional deployment | Deployment types allow explicit residency; AI Search is excellent for RAG |
| Gemini models, strong managed search and BigQuery-adjacent data | Vertex AI in asia-south1 |
Vertex AI Search, grounding, tight BigQuery integration |
| Predictable high-volume batch on a small model | Self-hosted on GPU nodes in Mumbai | Cost per token collapses at volume; residency guaranteed |
Common mistakes
- Enabling global or cross-region routing for latency without noticing it moves data out of India.
- Calling the model over the public internet from a private subnet through a NAT gateway, and paying NAT data processing charges on every token.
- No evaluation suite, so nobody can safely switch models when prices or availability change.
- One shared API key used by every feature, so cost and abuse cannot be attributed.
- Choosing the largest model for every request. Most support questions do not need it.
What to do this week
- Write down the region requirement and confirm which models are actually served there today, on your chosen platform.
- Put a private endpoint in front of the model API and restrict the callable model list in IAM or policy.
- Reorder your prompts so the static prefix comes first and turn on prompt caching.
- Tag every AI resource and set a budget alert per feature.
- Estimate your monthly tokens with the arithmetic above and compare the managed price to a single GPU instance. Keep the spreadsheet; you will revisit it every quarter.
The platforms will keep changing. The controls above do not.
- genai
- aws
- azure
- gcp
- bedrock
- vertex-ai
- finops
- india