DPDPA 2023: an engineering checklist for AI and cloud teams
By VA2PT Team, . 11 min read

Most DPDPA articles are written for lawyers. This one is written for the people who will actually get the ticket: the backend engineer who has to make "delete my account" mean something, the platform team that owns the S3 buckets, and the ML engineer who just shipped a support chatbot that reads customer tickets.
India's Digital Personal Data Protection Act, 2023 (DPDPA) is short by the standards of privacy law, and that is the problem. It states outcomes ("erase personal data when the purpose is served") without telling you how. Below, each obligation is translated into a system requirement you can put in a backlog, with the cloud-specific traps we keep seeing in audits.
A note on status: the Act received assent in August 2023, and the government has been phasing in the Act and its rules since. Timelines for specific obligations depend on the notified rules, so treat the deadlines your legal team gives you as the source of truth and use this article for the how.
The scenario we will use
A Bengaluru fintech with 4 million users runs on AWS in ap-south-1. They store KYC documents in S3, customer records in RDS PostgreSQL, events in Kinesis and ClickHouse, and they have just launched an LLM support assistant on Amazon Bedrock that answers questions using past tickets. Their marketing team also exports segments to a US-based email tool.
Every point below is checked against that stack.
1. Who you are under the Act decides your obligations
The Act names three roles:
- Data Principal: the individual the data is about.
- Data Fiduciary: the entity that decides why and how personal data is processed. That is you, if you own the product.
- Data Processor: an entity that processes on behalf of a fiduciary. AWS, your email vendor and your LLM API provider are processors for you.
The Act also allows the government to designate Significant Data Fiduciaries (SDFs) based on volume and sensitivity of data, with extra duties such as a Data Protection Officer, independent audits and data protection impact assessments. A fintech with millions of users should plan as if it will be designated.
System requirement: maintain a processor register. It is a table, not a document: vendor, what personal data flows to them, region, contract reference, and the technical control that limits the flow (IAM policy, allowlisted fields, VPC endpoint). If you cannot name the control, you do not have one.
2. Consent has to be a record, not a checkbox
The Act requires consent to be free, specific, informed, unconditional and unambiguous, given through a clear affirmative action, and as easy to withdraw as it was to give. Consent requests must be accompanied by a notice that states what data is collected and for what purpose.
The engineering consequence is that consent is an event with a lifecycle, and you must be able to prove it.
-- One row per consent decision; never update in place, always insert.
CREATE TABLE consent_events (
id bigserial PRIMARY KEY,
principal_id uuid NOT NULL,
purpose text NOT NULL, -- 'kyc', 'marketing_email', 'support_llm'
action text NOT NULL CHECK (action IN ('granted','withdrawn')),
notice_version text NOT NULL, -- hash or semver of the notice shown
channel text NOT NULL, -- 'web', 'android', 'ivr'
occurred_at timestamptz NOT NULL DEFAULT now(),
evidence jsonb NOT NULL -- ip, user agent, screen id, locale
);
CREATE INDEX ON consent_events (principal_id, purpose, occurred_at DESC);
Current state is a query, not a column:
SELECT DISTINCT ON (purpose) purpose, action, occurred_at
FROM consent_events
WHERE principal_id = $1
ORDER BY purpose, occurred_at DESC;
Three rules that make this work in practice:
- Purpose is a first-class key. "Marketing" is not a purpose; "marketing email about loan offers" is. Every downstream job that touches personal data declares the purpose it runs under, and the job refuses to run for principals without an active grant.
- Withdrawal propagates. A withdrawal event must trigger the same pipeline that handles erasure for that purpose (section 4 below), because the Act says processing must stop within a reasonable time after withdrawal.
- Keep the notice. Store the exact notice text (or its hash and a versioned copy in S3 with Object Lock) so you can show what the person agreed to.
Common mistake: storing marketing_opt_in boolean on the users table and overwriting it. You lose the history, the notice version and the channel, which is exactly what a regulator or an auditor will ask for.
3. Data residency versus cross-border transfer
The DPDPA does not mandate that all personal data stays in India. It allows transfer to any country except those the central government restricts by notification. Sector regulators, however, do impose localisation: the RBI's 2018 payment data directive requires payment system data to be stored only in India, and SEBI and IRDAI have their own rules. So for our fintech, the answer is: payment data stays in ap-south-1, full stop; other personal data may leave, subject to the notified list and your own risk appetite.
Turn that into controls:
{
"Version": "2012-10-17",
"Statement": [{
"Sid": "DenyOutsideIndia",
"Effect": "Deny",
"Action": "*",
"Resource": "*",
"Condition": {
"StringNotEquals": { "aws:RequestedRegion": ["ap-south-1", "ap-south-2"] },
"ArnNotLike": { "aws:PrincipalArn": "arn:aws:iam::*:role/GlobalServicesBreakGlass" }
}
}]
}
Attach that as a Service Control Policy at the OU that holds regulated workloads. Then check the leaks that an SCP does not catch:
- S3 Cross-Region Replication rules pointing outside India.
- CloudFront with origin shield or logging buckets in
us-east-1. - Global services (IAM, Route 53, CloudFront) that legitimately need
us-east-1; hence the break-glass role above. - SaaS vendors: your email tool in the US receives name, email and segment membership. That is a cross-border transfer. Record it in the processor register, confirm the country is not on the restricted list, and minimise the fields (send a hashed ID and a segment code, not the phone number).
- LLM APIs: check the region of the model endpoint you call. Bedrock has model availability in Mumbai for some models, not all; Azure OpenAI and Vertex AI have their own regional lists. If the endpoint is in Singapore or the US, that is a transfer too, and the prompt is the payload.
4. Erasure has to reach every copy
The Act requires erasure when consent is withdrawn or when the specified purpose is no longer served, unless retention is required by law. Retention laws are the tricky part: the Prevention of Money Laundering Act rules require regulated entities to keep KYC records for five years after the business relationship ends. So "delete my account" for our fintech means: delete everything you can, put the KYC records into a legal-hold state where they are inaccessible for any purpose except the legal one, and set a timer.
Here is a retention matrix that fits in a Confluence page and survives an audit:
| Data set | Store | Purpose | Retention after purpose ends | Erasure mechanism |
|---|---|---|---|---|
| Profile, preferences | RDS | service delivery | 0 days | hard delete row, cascade |
| KYC documents | S3 (Object Lock, Compliance) | legal (PMLA) | 5 years | lifecycle expiration after hold |
| Transaction history | RDS + ClickHouse | legal/accounting | per statute | partition drop by month |
| Product events | Kinesis → S3 | analytics | 90 days | S3 lifecycle, key-partitioned by principal_id |
| Support tickets | RDS + OpenSearch | support | 2 years | delete + reindex |
| Ticket embeddings | vector DB | LLM assistant | same as tickets | delete by metadata filter |
| Backups | RDS snapshots | DR | 35 days | rolling; document the lag |
Points that trip teams up:
- Vector databases hold personal data. An embedding of a ticket that says "my card ending 4421 was charged twice, I am Priya from Pune" is derived personal data. Store
principal_idas metadata on every vector and delete by filter. If your vector store cannot delete by metadata efficiently, that is a reason to pick another one. - Backups. You cannot surgically delete a person from a 35-day-old RDS snapshot. The accepted approach is to document the backup retention window, ensure restores re-run pending erasures (keep an
erasure_queuetable that a restore hook replays), and keep the window short. - Logs. Application logs with emails or phone numbers in them are personal data. Mask at the log library level, not at the SIEM.
- Object Lock in Compliance mode cannot be shortened, even by the root account. Set it to the legal minimum, not "10 years to be safe".
5. Breach notification and the CERT-In clock
Under the DPDPA, a data fiduciary must notify the Data Protection Board and each affected data principal in the event of a personal data breach, in the form and manner prescribed by the rules. Separately, CERT-In's April 2022 directions require reporting of specified cyber incidents to CERT-In within six hours of noticing them. The two regimes overlap but are not the same: CERT-In cares about the incident class; the DPDPA cares about personal data.
Engineering requirements:
- A single incident timeline with a first-detected timestamp that everyone agrees on. Six hours starts from "noticing", so log when the alert fired, when a human acknowledged it, and when it was classified as a breach.
- Scope query in minutes, not days. You need to answer "which principals, which fields" fast. That means personal data is tagged at the column level (a
piicomment or a data catalog entry) and access is logged (RDS audit logs, S3 server access logs, CloudTrail data events on sensitive buckets). - Templates ready. The CERT-In format and the DPDPA notice content are known in advance. Keep them in the incident runbook with the fields pre-mapped to your logs.
- Log retention of 180 days is a CERT-In requirement for ICT logs, kept within India. Check your CloudWatch and S3 log lifecycle rules against that number.
6. What an LLM feature changes
Our fintech's support assistant reads past tickets and answers new ones. Four new obligations appear:
Purpose creep. Tickets were collected for support. Using them to train or fine-tune a model, or to build a retrieval index, is a new purpose. Either it is covered by the original notice, or you need fresh consent, or you rely on a legitimate use ground the Act recognises. Do not assume; write down which one.
Prompts are data flows. Every prompt that includes a customer's ticket text is a disclosure to your model provider. If the provider is outside India, section 3 applies. Bedrock and Azure OpenAI both state that customer prompts are not used to train foundation models; verify the current terms and record them in the processor register.
Erasure reaches the index. When a principal is erased, the retrieval index and any cached responses that quote their data must be purged. Cache keys should include the principal ID so purges are targeted.
Minimise before you send. Redact PAN, Aadhaar, card numbers and phone numbers from ticket text before it goes into the prompt. A regex pass catches most; a small NER model catches the rest.
import re
PATTERNS = {
"pan": re.compile(r"\b[A-Z]{5}[0-9]{4}[A-Z]\b"),
"aadhaar": re.compile(r"\b\d{4}\s?\d{4}\s?\d{4}\b"),
"card": re.compile(r"\b(?:\d[ -]?){13,19}\b"),
"phone": re.compile(r"(?:\+91[\-\s]?)?[6-9]\d{9}\b"),
}
def redact(text: str) -> str:
for name, rx in PATTERNS.items():
text = rx.sub(f"[{name.upper()}_REDACTED]", text)
return text
Run this on the way into the prompt, and again on the model's output before it is shown or logged. It is not perfect. It is the difference between a controlled disclosure and an uncontrolled one.
7. Rights requests need an API, not an inbox
Principals have the right to access a summary of their data and the processing activities, to correct and erase, to grievance redressal, and to nominate someone to exercise rights on their behalf. The rules set response timelines; your legal team will tell you the number. The engineering question is whether you can fulfil a request at all without a two-week manual scramble.
Build a principal_data_export job that walks the retention matrix and produces a JSON archive, and a principal_erasure job that walks it in reverse. Both are idempotent, both write an audit record, and both are exposed to support staff through an internal tool with approval. If you have done section 4 properly, this is a week of work. If you have not, this is the thing that forces you to.
Common mistakes we see in audits
- Encrypting everything and calling it compliance. Encryption at rest with the default KMS key does nothing for consent, purpose limitation or erasure.
- Treating analytics events as anonymous because they carry a UUID. A UUID you can join back to a user is personal data.
- Vendors added by product managers through a credit card, never entered in the processor register.
- Data in developers' laptops, notebooks and S3 "scratch" buckets, copied from production for debugging.
- LLM prompts logged in full to a third-party observability tool in the US.
What to do this week
- Build the processor register as a table and populate it from your AWS bill and your SaaS spend. Ask every vendor two questions: where is the data stored, and is it used for training.
- Write the retention matrix for your top five data stores. Include the vector database and the backups.
- Add the consent events table and stop overwriting opt-in booleans.
- Put the SCP above in audit mode (CloudTrail query for
RequestedRegionoutside India) before enforcing it. - Add PII redaction in front of any LLM call, and confirm the model endpoint region.
- Put the CERT-In six-hour clock and the DPDPA notice template into your incident runbook, and run one tabletop exercise.
None of this requires a compliance platform. It requires that the people who own the data stores own the obligations, and that the controls are in code where they can be reviewed.
- dpdpa
- ai-compliance
- ai-grc
- data-privacy
- india
- cloud-security