DATA CHASING AI
Turning the Table on the Global AI Data
Economy
A Proposal for a National Citizen Data
Trust as India's Foundational AI Advantage
White Paper
Author: Hemen Parekh
Mumbai, India
July 2026
Executive
Summary
The global Foundational
AI industry stands at an inflection point defined not by compute or algorithms,
but by a resource that is running out under the current rules of engagement:
legitimately obtainable, high-quality training data. The dominant model of the
last decade — AI companies scraping, crawling, and ingesting the open web and
licensed corpora with only partial consent — has now collided with the legal
system. Publishers and authors, including novelist Scott Turow, have filed a
class-action suit against Google alleging that books supplied for search and
retail purposes were used to train Gemini without authorisation. This is not an
isolated dispute; it is a preview of the litigation, cost, and reputational
drag that will increasingly define AI development wherever the data-acquisition
model remains extractive rather than consensual.
This white paper argues
that India need not import this conflict. Instead, India has the demographic
scale, the digital public infrastructure legacy (Aadhaar, UPI, DigiLocker,
India Stack), and the policy appetite to invert the model entirely — replacing
“AI chasing data” with “data chasing AI.” Under the proposed framework, India's
140 crore citizens would be invited to voluntarily and knowingly deposit
categories of their own personal data — strictly anonymised and aggregated
before ever reaching an AI developer — into a National Citizen Data Trust.
Indian Foundational AI companies would then license access to this pool for a
modest fee, giving them a training-data advantage that is simultaneously
larger, richer, more India-specific, and more legally sound than anything
obtainable through scraping.
This is not a incremental
improvement to data-sourcing practice. It is a paradigm change in how the raw
material of AI is acquired — one that converts India's much-discussed
demographic dividend into a data dividend, positions Indian citizens as willing
co-owners of the AI economy rather than its raw material, and gives Indian
Foundational AI companies a structural head start over their US and Chinese
counterparts, at a fraction of the legal and financial cost.
1.
The Crisis in the Prevailing Data Paradigm
Every major Foundational
AI model built to date has been trained substantially on data whose provenance
and consent status is contested. The economics of large language models made
this inevitable: training at scale required more text, image, audio and video
data than any licensing regime could supply in time, so companies crawled first
and negotiated — or litigated — later.
The Google/Gemini case is
illustrative of where this trajectory leads:
Publishers and novelist Scott Turow are suing
Google over books used to train Gemini. The complaint states the works were
supplied for search and retail purposes, not for AI training — and that
blocking AI crawlers would not have covered either of those originally licensed
routes.
The significance is not
the specific dispute but the pattern it represents: even Google, with vast
legal resources and years of experience negotiating content licenses, cannot
fully insulate itself from the consent problem baked into the scrape-first
model. Every AI company following this path — whether in the US, China, or
India — inherits the same exposure: uncertain provenance, unresolved consent,
and an adversarial relationship with the very creators and citizens whose data
makes the models useful.
For India's emerging
Sovereign AI ambitions, replicating this model would mean importing its
liabilities before a single lawsuit has been filed domestically. The alternative
is to design the data layer correctly from the outset.
2.
Prior Art and Provenance: An Original, Pioneering Idea
This is not a reactive
idea, assembled in response to the Google/Gemini litigation or the current
Sovereign AI debate. It is original, and it is
prior art.
The author first proposed
this data-deposit model on 18 November 2018,
years before “AI training data” entered mainstream policy discourse, in the
blog post “Only Answer — Statutory Warning”: https://myblogepage.blogspot.com/2018/11/only-answer-statutory-warning.html.
The author followed up on 6 August 2023,
as generative AI moved to the centre of global policy, with “Stopping Data
Leakage”: https://myblogepage.blogspot.com/2023/08/stopping-data-leakage.html,
refining the same core proposition well ahead of today's AI-versus-publisher
lawsuits.
This white paper is
therefore best read as the third, consolidated milestone in a continuous line
of thought stretching back to 2018 — not an
also-ran response to a trend, but the trend's earliest articulation.
As global AI developers now scramble to resolve exactly the
consent-and-provenance problem this proposal anticipated by several years, the
record should reflect that the underlying idea — citizens as knowing,
compensated depositors of their own data, rather than unwitting sources scraped
without consent — is a genuinely Indian-origin,
first-mover concept.
3.
The Paradigm Shift: From “AI Chasing Data” to “Data Chasing AI”
The proposal is simple to
state and profound in its implications: instead of AI companies pursuing data
through scraping, licensing battles, and litigation, India creates the
conditions under which citizens' own data flows willingly toward a common
national pool, from which AI developers then draw — transparently, legally, and
for a fee.
This reframes four
relationships simultaneously:
1.
The
citizen shifts from being an unwitting data source to a knowing, consenting
depositor and economic stakeholder.
2.
The AI
developer shifts from data hunter to data licensee — a legitimate customer of a
transparent national resource rather than a defendant-in-waiting.
3.
The
State shifts from bystander/regulator-after-the-fact to architect of the data
infrastructure itself — much as it did with UPI for payments and Aadhaar for
identity.
4.
The
relationship between creators/citizens and AI companies shifts from adversarial
to cooperative, because consent and compensation are designed in rather than
litigated in afterward.
This is why the proposal
should not be read as a data-privacy footnote to India's AI strategy. It is
closer in kind to what UPI did for digital payments or what Aadhaar did for
identity verification — a foundational, State-enabled layer that does not
merely support the AI industry but changes the economics on which the entire
industry is built.
4.
The Ten-Tier National Data Framework
The Trust would organise
voluntary citizen contributions into ten data tiers, each independently
consentable — citizens choose which tiers they wish to contribute to, and can
withdraw consent at any time:
|
Tier
|
Data Category
|
Illustrative Contents
|
|
1
|
Personal Data
|
Demographic identifiers,
preferences, life-stage attributes (stripped of PII before pooling)
|
|
2
|
Family Data
|
Household composition,
generational patterns, dependency structures
|
|
3
|
Contact Information
|
Communication and reachability
patterns (never raw numbers/addresses)
|
|
4
|
Social Media Data
|
Behaviour across websites,
portals and apps — usage, sentiment, interaction patterns
|
|
5
|
Cultural Exposure Data
|
Language, media, festival and
consumption patterns reflecting India's cultural diversity
|
|
6
|
Educational Qualifications
|
Degrees, streams, institutions,
skill certifications
|
|
7
|
Language Knowledge
|
Multilingual proficiency —
critical for India-specific LLMs given 22+ scheduled languages
|
|
8
|
Job Experience
|
Sector, role progression, skills
applied — a live labour-market signal
|
|
9
|
Wealth / Income Data
|
Aggregated, banded income and
tax-return-derived economic indicators
|
|
10
|
Medical History Data
|
De-identified, aggregated health
patterns for health-AI applications
|
Citizens would never be
required to contribute all ten tiers. A graduated, opt-in architecture —
similar in spirit to DigiLocker's consent-based document sharing — ensures
participation is genuinely voluntary and layered, not a blanket surrender of
privacy.
5.
Architecture: Consent, Anonymisation, and Aggregation
The credibility of this
proposal rests entirely on the rigour of its privacy architecture. Three
principles are non-negotiable:
4.1 Consent by Design
Every tier of data
contribution requires explicit, revocable, tier-specific consent — modelled on
the Data Empowerment and Protection Architecture (DEPA) consent framework
already piloted in India's financial data-sharing ecosystem. No blanket
consent, no dark patterns, and no default opt-in.
4.2 Anonymisation and
De-identification at Source
Personally identifiable
information is stripped or irreversibly tokenised before data ever leaves the
citizen-facing layer. No AI developer, at any tier of access, would ever
receive raw, identity-linked records — only anonymised, aggregated datasets.
4.3 Aggregation Before
Access
Data is pooled and
statistically aggregated before licensing, so that no individual citizen's
record is separately reconstructable by a licensee, addressing both
re-identification risk and the commercial incentive to over-collect.
Together, these
principles are designed to satisfy — and go beyond — the requirements of the
Digital Personal Data Protection Act, positioning the Trust as a model of
compliance rather than a regulatory risk.
6.
The Economic Model: A National Data Dividend
Indian Foundational AI
companies — and, potentially, MeitY's empanelled AI partner firms — would
license access to anonymised, aggregated datasets from the Trust for a fee,
structured on tiered pricing by data category and volume. The resulting revenue
stream can be applied in three directions:
•
A
citizen data dividend — a share of licensing revenue returned to contributing
citizens, converting personal data from an extracted asset into an
income-generating one.
•
Reinvestment
into the Trust's own security, audit, and consent-infrastructure to keep pace
with data volume and AI developer demand.
•
A
public fund to support smaller Indian AI startups that cannot yet afford
large-scale licensing, ensuring the advantage is not captured only by the
largest players.
This model does something
the scrape-first paradigm structurally cannot: it turns data provision into a
positive-sum transaction rather than a zero-sum extraction, with citizens as
active beneficiaries rather than passive raw material.
7.
Strategic Advantage: India versus the US and China
|
Dimension
|
Prevailing Global Model
|
Proposed Indian Model
|
|
Data acquisition
|
Scraping, licensing disputes,
litigation (e.g. publishers vs. Google/Gemini)
|
Voluntary citizen deposit,
consent-first
|
|
Legal exposure
|
High — copyright and consent
lawsuits
|
Low — provenance and consent
recorded at source
|
|
Cost of data acquisition
|
High — licensing fees, legal
costs, reputational risk
|
Low — modest access fee to a
national pool
|
|
Data richness
|
Public web text — increasingly
exhausted
|
Deep, multi-dimensional,
life-cycle data across 140 crore citizens
|
|
Who benefits economically
|
Platform and AI companies
primarily
|
Citizens (fee-share), Indian AI
developers, and the exchequer
|
|
Trust posture
|
Adversarial — creators vs. AI
companies
|
Cooperative — citizens as
willing co-owners of the AI economy
|
India's advantage is
structural, not incidental. No other democracy has both the population scale
and the digital public infrastructure legacy — Aadhaar, UPI, DigiLocker, the
India Stack — to stand up a consent-based national data trust at this scale,
this quickly. The United States lacks a unified digital identity layer of
comparable reach; China's data governance model does not rest on individual
consent in the way India's constitutional and DPDP framework requires. This is
a genuinely available first-mover position for Indian Foundational AI companies
— available to no one else in the same form.
8.
Governance and Legal Safeguards
To earn and retain public
trust, the Trust should be governed as an independent statutory body, at arm's
length from both the AI industry and any single ministry, with:
•
A
citizen-representation component in its governing board, not merely industry
and government appointees.
•
Mandatory,
published third-party audits of anonymisation and de-identification standards.
•
A
public, real-time consent dashboard where citizens can view and revoke what
they have contributed, modelled on DigiLocker's consent logs.
•
Statutory
penalties for any licensee found attempting re-identification or secondary use
beyond the licensed purpose.
9.
Implementation Roadmap
5.
Pilot
(0–6 months): A single-tier pilot — e.g., Education Qualifications and Job
Experience data — run in partnership with MeitY and one or two empanelled AI
partner firms, to test consent flows, anonymisation pipelines, and licensing
mechanics at modest scale.
6.
Statutory
Foundation (6–12 months): Enabling legislation or a DPDP-compliant statutory
framework establishing the Trust as an independent body, with governance
safeguards as outlined in Section 7.
7.
Phased
Tier Rollout (12–24 months): Sequential activation of additional data tiers,
prioritising those with the clearest AI-training value and the lowest
sensitivity (e.g., Cultural Exposure and Language Knowledge data ahead of
Medical History).
8.
National
Scale-Up (24+ months): Full ten-tier availability, integrated citizen dividend
disbursement, and open licensing to all qualified Indian AI developers, not
solely the initially empanelled firms.
10.
Conclusion: A Truly Indian AI Innovation
The Google/Gemini
litigation is a warning, not an anomaly. Every AI industry built on the
scrape-first model will eventually face its own version of this reckoning.
India does not need to wait for that reckoning to arrive at its own doorstep —
it can design around it now, and in doing so, create something the rest of the
world's AI industry does not yet have: a consent-first, citizen-owned,
transparently licensed national data infrastructure purpose-built for the AI
era.
This is offered not as a
minor policy suggestion but as a foundational proposition for India's Sovereign
AI ambitions — a genuinely Indian innovation, as consequential to the AI
economy as UPI has been to digital payments. It deserves consideration at the
highest levels of government, and I would welcome the opportunity to develop it
further with the Ministry of Electronics and Information Technology and its
partners.
— Hemen Parekh, Mumbai
Prior art references: see
Section 2 above (2018 and 2023 blog posts).