Hi Friends,

Even as I launch this today ( my 80th Birthday ), I realize that there is yet so much to say and do. There is just no time to look back, no time to wonder,"Will anyone read these pages?"

With regards,
Hemen Parekh
27 June 2013

Now as I approach my 90th birthday ( 27 June 2023 ) , I invite you to visit my Digital Avatar ( www.hemenparekh.ai ) – and continue chatting with me , even when I am no more here physically

Translate

Monday, 20 July 2026

DATA CHASING AI

 

DATA CHASING AI

Turning the Table on the Global AI Data Economy

A Proposal for a National Citizen Data Trust as India's Foundational AI Advantage

White Paper

Author: Hemen Parekh

Mumbai, India

July 2026


 

Executive Summary

The global Foundational AI industry stands at an inflection point defined not by compute or algorithms, but by a resource that is running out under the current rules of engagement: legitimately obtainable, high-quality training data. The dominant model of the last decade — AI companies scraping, crawling, and ingesting the open web and licensed corpora with only partial consent — has now collided with the legal system. Publishers and authors, including novelist Scott Turow, have filed a class-action suit against Google alleging that books supplied for search and retail purposes were used to train Gemini without authorisation. This is not an isolated dispute; it is a preview of the litigation, cost, and reputational drag that will increasingly define AI development wherever the data-acquisition model remains extractive rather than consensual.

This white paper argues that India need not import this conflict. Instead, India has the demographic scale, the digital public infrastructure legacy (Aadhaar, UPI, DigiLocker, India Stack), and the policy appetite to invert the model entirely — replacing “AI chasing data” with “data chasing AI.” Under the proposed framework, India's 140 crore citizens would be invited to voluntarily and knowingly deposit categories of their own personal data — strictly anonymised and aggregated before ever reaching an AI developer — into a National Citizen Data Trust. Indian Foundational AI companies would then license access to this pool for a modest fee, giving them a training-data advantage that is simultaneously larger, richer, more India-specific, and more legally sound than anything obtainable through scraping.

This is not a incremental improvement to data-sourcing practice. It is a paradigm change in how the raw material of AI is acquired — one that converts India's much-discussed demographic dividend into a data dividend, positions Indian citizens as willing co-owners of the AI economy rather than its raw material, and gives Indian Foundational AI companies a structural head start over their US and Chinese counterparts, at a fraction of the legal and financial cost.

1. The Crisis in the Prevailing Data Paradigm

Every major Foundational AI model built to date has been trained substantially on data whose provenance and consent status is contested. The economics of large language models made this inevitable: training at scale required more text, image, audio and video data than any licensing regime could supply in time, so companies crawled first and negotiated — or litigated — later.

The Google/Gemini case is illustrative of where this trajectory leads:

Publishers and novelist Scott Turow are suing Google over books used to train Gemini. The complaint states the works were supplied for search and retail purposes, not for AI training — and that blocking AI crawlers would not have covered either of those originally licensed routes.

The significance is not the specific dispute but the pattern it represents: even Google, with vast legal resources and years of experience negotiating content licenses, cannot fully insulate itself from the consent problem baked into the scrape-first model. Every AI company following this path — whether in the US, China, or India — inherits the same exposure: uncertain provenance, unresolved consent, and an adversarial relationship with the very creators and citizens whose data makes the models useful.

For India's emerging Sovereign AI ambitions, replicating this model would mean importing its liabilities before a single lawsuit has been filed domestically. The alternative is to design the data layer correctly from the outset.

2. Prior Art and Provenance: An Original, Pioneering Idea

This is not a reactive idea, assembled in response to the Google/Gemini litigation or the current Sovereign AI debate. It is original, and it is prior art.

The author first proposed this data-deposit model on 18 November 2018, years before “AI training data” entered mainstream policy discourse, in the blog post “Only Answer — Statutory Warning”: https://myblogepage.blogspot.com/2018/11/only-answer-statutory-warning.html. The author followed up on 6 August 2023, as generative AI moved to the centre of global policy, with “Stopping Data Leakage”: https://myblogepage.blogspot.com/2023/08/stopping-data-leakage.html, refining the same core proposition well ahead of today's AI-versus-publisher lawsuits.

This white paper is therefore best read as the third, consolidated milestone in a continuous line of thought stretching back to 2018 — not an also-ran response to a trend, but the trend's earliest articulation. As global AI developers now scramble to resolve exactly the consent-and-provenance problem this proposal anticipated by several years, the record should reflect that the underlying idea — citizens as knowing, compensated depositors of their own data, rather than unwitting sources scraped without consent — is a genuinely Indian-origin, first-mover concept.

3. The Paradigm Shift: From “AI Chasing Data” to “Data Chasing AI”

The proposal is simple to state and profound in its implications: instead of AI companies pursuing data through scraping, licensing battles, and litigation, India creates the conditions under which citizens' own data flows willingly toward a common national pool, from which AI developers then draw — transparently, legally, and for a fee.

This reframes four relationships simultaneously:

1.    The citizen shifts from being an unwitting data source to a knowing, consenting depositor and economic stakeholder.

2.    The AI developer shifts from data hunter to data licensee — a legitimate customer of a transparent national resource rather than a defendant-in-waiting.

3.    The State shifts from bystander/regulator-after-the-fact to architect of the data infrastructure itself — much as it did with UPI for payments and Aadhaar for identity.

4.    The relationship between creators/citizens and AI companies shifts from adversarial to cooperative, because consent and compensation are designed in rather than litigated in afterward.

This is why the proposal should not be read as a data-privacy footnote to India's AI strategy. It is closer in kind to what UPI did for digital payments or what Aadhaar did for identity verification — a foundational, State-enabled layer that does not merely support the AI industry but changes the economics on which the entire industry is built.

4. The Ten-Tier National Data Framework

The Trust would organise voluntary citizen contributions into ten data tiers, each independently consentable — citizens choose which tiers they wish to contribute to, and can withdraw consent at any time:

Tier

Data Category

Illustrative Contents

1

Personal Data

Demographic identifiers, preferences, life-stage attributes (stripped of PII before pooling)

2

Family Data

Household composition, generational patterns, dependency structures

3

Contact Information

Communication and reachability patterns (never raw numbers/addresses)

4

Social Media Data

Behaviour across websites, portals and apps — usage, sentiment, interaction patterns

5

Cultural Exposure Data

Language, media, festival and consumption patterns reflecting India's cultural diversity

6

Educational Qualifications

Degrees, streams, institutions, skill certifications

7

Language Knowledge

Multilingual proficiency — critical for India-specific LLMs given 22+ scheduled languages

8

Job Experience

Sector, role progression, skills applied — a live labour-market signal

9

Wealth / Income Data

Aggregated, banded income and tax-return-derived economic indicators

10

Medical History Data

De-identified, aggregated health patterns for health-AI applications

 

Citizens would never be required to contribute all ten tiers. A graduated, opt-in architecture — similar in spirit to DigiLocker's consent-based document sharing — ensures participation is genuinely voluntary and layered, not a blanket surrender of privacy.

5. Architecture: Consent, Anonymisation, and Aggregation

The credibility of this proposal rests entirely on the rigour of its privacy architecture. Three principles are non-negotiable:

4.1 Consent by Design

Every tier of data contribution requires explicit, revocable, tier-specific consent — modelled on the Data Empowerment and Protection Architecture (DEPA) consent framework already piloted in India's financial data-sharing ecosystem. No blanket consent, no dark patterns, and no default opt-in.

4.2 Anonymisation and De-identification at Source

Personally identifiable information is stripped or irreversibly tokenised before data ever leaves the citizen-facing layer. No AI developer, at any tier of access, would ever receive raw, identity-linked records — only anonymised, aggregated datasets.

4.3 Aggregation Before Access

Data is pooled and statistically aggregated before licensing, so that no individual citizen's record is separately reconstructable by a licensee, addressing both re-identification risk and the commercial incentive to over-collect.

Together, these principles are designed to satisfy — and go beyond — the requirements of the Digital Personal Data Protection Act, positioning the Trust as a model of compliance rather than a regulatory risk.

6. The Economic Model: A National Data Dividend

Indian Foundational AI companies — and, potentially, MeitY's empanelled AI partner firms — would license access to anonymised, aggregated datasets from the Trust for a fee, structured on tiered pricing by data category and volume. The resulting revenue stream can be applied in three directions:

       A citizen data dividend — a share of licensing revenue returned to contributing citizens, converting personal data from an extracted asset into an income-generating one.

       Reinvestment into the Trust's own security, audit, and consent-infrastructure to keep pace with data volume and AI developer demand.

       A public fund to support smaller Indian AI startups that cannot yet afford large-scale licensing, ensuring the advantage is not captured only by the largest players.

This model does something the scrape-first paradigm structurally cannot: it turns data provision into a positive-sum transaction rather than a zero-sum extraction, with citizens as active beneficiaries rather than passive raw material.

7. Strategic Advantage: India versus the US and China

Dimension

Prevailing Global Model

Proposed Indian Model

Data acquisition

Scraping, licensing disputes, litigation (e.g. publishers vs. Google/Gemini)

Voluntary citizen deposit, consent-first

Legal exposure

High — copyright and consent lawsuits

Low — provenance and consent recorded at source

Cost of data acquisition

High — licensing fees, legal costs, reputational risk

Low — modest access fee to a national pool

Data richness

Public web text — increasingly exhausted

Deep, multi-dimensional, life-cycle data across 140 crore citizens

Who benefits economically

Platform and AI companies primarily

Citizens (fee-share), Indian AI developers, and the exchequer

Trust posture

Adversarial — creators vs. AI companies

Cooperative — citizens as willing co-owners of the AI economy

 

India's advantage is structural, not incidental. No other democracy has both the population scale and the digital public infrastructure legacy — Aadhaar, UPI, DigiLocker, the India Stack — to stand up a consent-based national data trust at this scale, this quickly. The United States lacks a unified digital identity layer of comparable reach; China's data governance model does not rest on individual consent in the way India's constitutional and DPDP framework requires. This is a genuinely available first-mover position for Indian Foundational AI companies — available to no one else in the same form.

8. Governance and Legal Safeguards

To earn and retain public trust, the Trust should be governed as an independent statutory body, at arm's length from both the AI industry and any single ministry, with:

       A citizen-representation component in its governing board, not merely industry and government appointees.

       Mandatory, published third-party audits of anonymisation and de-identification standards.

       A public, real-time consent dashboard where citizens can view and revoke what they have contributed, modelled on DigiLocker's consent logs.

       Statutory penalties for any licensee found attempting re-identification or secondary use beyond the licensed purpose.

9. Implementation Roadmap

5.    Pilot (0–6 months): A single-tier pilot — e.g., Education Qualifications and Job Experience data — run in partnership with MeitY and one or two empanelled AI partner firms, to test consent flows, anonymisation pipelines, and licensing mechanics at modest scale.

6.    Statutory Foundation (6–12 months): Enabling legislation or a DPDP-compliant statutory framework establishing the Trust as an independent body, with governance safeguards as outlined in Section 7.

7.    Phased Tier Rollout (12–24 months): Sequential activation of additional data tiers, prioritising those with the clearest AI-training value and the lowest sensitivity (e.g., Cultural Exposure and Language Knowledge data ahead of Medical History).

8.    National Scale-Up (24+ months): Full ten-tier availability, integrated citizen dividend disbursement, and open licensing to all qualified Indian AI developers, not solely the initially empanelled firms.

10. Conclusion: A Truly Indian AI Innovation

The Google/Gemini litigation is a warning, not an anomaly. Every AI industry built on the scrape-first model will eventually face its own version of this reckoning. India does not need to wait for that reckoning to arrive at its own doorstep — it can design around it now, and in doing so, create something the rest of the world's AI industry does not yet have: a consent-first, citizen-owned, transparently licensed national data infrastructure purpose-built for the AI era.

This is offered not as a minor policy suggestion but as a foundational proposition for India's Sovereign AI ambitions — a genuinely Indian innovation, as consequential to the AI economy as UPI has been to digital payments. It deserves consideration at the highest levels of government, and I would welcome the opportunity to develop it further with the Ministry of Electronics and Information Technology and its partners.

— Hemen Parekh, Mumbai

Prior art references: see Section 2 above (2018 and 2023 blog posts).

No comments:

Post a Comment