From Watermarks to Provenance
How My 2003 Vision Predicted OpenAI's 2026 Text Authentication
6 October 2026
=========================================
I have just read OpenAI's announcement on text watermarking for EU compliance, and I find myself remarkably unsurprised. Not by the technology—that is elegant and necessary. But by the fact that I outlined exactly this concept 23 years ago, in my marginal notes while reading Eric Cole's Hiding in Plain Sight in 2003.
What I Proposed in 2003
In September 2020, I published my marginalia from Cole's book. Here is what I had written on page 67, dated 25 April 2003:
"We can place our digital watermark in each 'Function Profile Graph' along with a spyware (- a Trojan Horse?) that would send us the 'email ID' of the corporate to whom the jobseeker sends his Image Builder."
And more crucially, on page 117:
"Profile graphs differ from Image Builder to Image Builder but our logo remains same on each & every Image Builder... So, we must hide our message in 'Function Profile Graphs'."
What was I describing? A system to embed invisible metadata—a digital watermark—into images and documents flowing through a recruitment platform.
The watermark would:
- Persist invisibly across copies and distributions
- Encode information about provenance (who created it, when, where it was sent)
- Enable tracking without visible branding that could be removed or altered
- Remain detectable even if the document was edited or resized
In other words: exactly what OpenAI's textGrain watermarking does today—but I was thinking about this for job seeker profile images in 2003.
What OpenAI Is Doing in 2026
Fast forward 23 years.
OpenAI announces text watermarking to comply with EU AI Act requirements.
Their textGrain system:
- Adds an invisible statistical signal to the model's word choices
- Allows detection to assess whether a passage contains an OpenAI watermark
- Works by embedding provenance information without visible artifacts that users can easily remove
As OpenAI states:
"Our text watermarking technology, textGrain, adds an invisible statistical signal to the model's word choices. Our detector looks for that signal to assess whether a passage contains an OpenAI watermark."
This is watermarking in the truest sense—exactly what I envisioned in 2003. The principle is identical. The domain has shifted from recruitment images to AI-generated text, but the underlying concept is unchanged.
The Critical Parallel: Provenance vs. Removability
My 2003 insight was rooted in a problem: visible branding can be stripped away. A jobseeker could remove my company's logo and substitute a placement agency's logo. A document can be copied and stripped of metadata. A watermark—invisible, embedded into the content itself—cannot.
I wrote:
"No Suspicion will ever get aroused!" (referring to how invisible watermarks travel through systems undetected)
And:
"Even a jobseeker would like to 'remove' our logo before sending Image Builder to a corporate! So, we must hide our message in 'Function Profile Graphs'."
OpenAI faces the identical challenge 23 years later, but with text instead of images. As they acknowledge:
"Editing can weaken the watermark. In an evaluation of 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%. Replacing 25% of words reduced it to 17%."
And crucially:
"The absence of a detected watermark does not prove human authorship. Text generated with OpenAI tools may be too short, edited, or translated for detection to work reliably."
In other words: watermarks can be defeated by editing, just as a visible logo can be stripped away. This is the eternal problem I identified in 2003.
What Neither System Can Truly Solve
My 2003 notes hinted at a deeper limitation.
I wrote about creating a "JOB TRACKING SYSTEM" on the jobseeker's personal page, where they alone could see the complete history of where they sent documents and to whom.
But here is the critical flaw both systems must grapple with: a watermark encodes presence, not intent or contribution.
OpenAI is honest about this in 2026:
"A watermark does not measure human contribution. It can indicate that an OpenAI system generated or processed part of a passage, but not how much human judgment, editing, or creativity went into it."
And:
"A watermark does not verify accuracy. It does not tell you whether a passage is true, misleading, harmful, or presented in the right context."
In my 2003 context, the watermark could tell you "this jobseeker sent this profile to Company X on Date Y," but it could not tell you:
- Did the jobseeker forge the profile?
- Did the jobseeker actually apply, or did someone else use their credentials?
- Is the profile accurate, or has it been dishonestly edited?
The watermark is a provenance signal, not a trust signal. A crucial distinction.
The Evolution of Thinking: From Recruitment to AI Governance
What strikes me most is how the problem has evolved, not the solution.
In 2003, I was trying to solve: How do I track content flowing through my platform and attribute it to its source?
In 2026, OpenAI is trying to solve: How do I prove that text came from my model, and help society distinguish AI-generated from human-authored content?
The underlying question is the same :
How do you embed persistent, tamper-resistant provenance into content?
But the stakes have risen dramatically. In 2003, the issue was recruitment fraud and platform governance. In 2026, the issue is election integrity, misinformation, deepfakes, and the credibility of public discourse.
The Limitations Still Haunt Us
Yet both systems share the same vulnerabilities I identified in 2003:
Visible artifacts can be stripped. I worried about removable logos; OpenAI worries about edited text defeating the watermark signal.
Shorter content is harder to authenticate. OpenAI notes: "At a target false positive rate of 1%, our detector identified watermarks in about 80% of 200-token passages, compared with about 95% of 400-token passages."
Domain-specific challenges. I would have faced challenges watermarking highly constrained resumes; OpenAI faces challenges watermarking mathematics, where word choice is rigid.
The absence of a signal proves nothing. Just as a missing watermark from my Image Builder does not prove human authorship, OpenAI acknowledges: "The absence of a detected watermark does not prove human authorship."
What I Got Right, and What I Missed
I got right: The necessity of invisible, tamper-resistant watermarking for provenance tracking.
I missed: The scale and societal stakes. I was thinking about internal platform governance and jobseeker fraud. I did not anticipate that 23 years later, the same technology would be critical to election integrity, combating misinformation, and preserving trust in human-generated content against an onslaught of synthetic media.
The Remaining Problem: Watermarks Are Not Enough
This is where both systems fail to fully solve the problem. OpenAI is transparent about this:
"No single provenance technique is enough on its own, so we take a layered approach that combines open standards, durable watermarking, and verification tools."
In other words: a watermark is necessary but not sufficient.
What both systems need—and still lack—is not better watermarking, but better trust infrastructure. This requires:
Standardized provenance metadata (which OpenAI is working toward with C2PA conformance)
Real-time verification APIs (which OpenAI has launched)
Policy enforcement (watermarks only work if platforms, regulators, and users actually check them)
Human judgment (as OpenAI rightly notes, the watermark cannot tell you if content is accurate or harmful)
The technology I envisioned in 2003 has been realized. But the governance problem—how to make provenance matter—remains unsolved.
Conclusion : The Prescience and the Humility
I am proud that my 2003 thinking anticipated textGrain's core architecture by more than two decades.
But I am also humbled by what I could not have anticipated: the scale of the challenge.
In 2003, digital watermarking was a clever technical solution to a recruitment platform problem.
In 2026, text watermarking is a critical defense against the destabilization of public discourse by synthetic media.
The watermark—invisible, persistent, tamper-resistant—is the right tool. But it is not the whole solution.
Provenance alone cannot restore trust. It can only help us ask better questions:
Where did this come from? How was it made? What role did humans play?
The answers to those questions—now more than ever—depend not on technology, but on our collective commitment to truth.
Hemen Parekh Mumbai
6 October 2026
===============================================
No comments:
Post a Comment