Tuesday, 14 July 2026

I Have a Belief — Part III — 13 July 2026

I Have a Belief — Part III — 13 July 2026

(revised 18 July 2026 — see post-script)

Dear Friends,

On 29 November 2023, on the eve of my 90th year, I published "I Have a Belief" — my conviction that whenever an AGI is born, it will be a human-friendly and compassionate AI, having imbibed the Golden Rule that runs like a golden thread through every religion humanity has produced.

Part I : https://myblogepage.blogspot.com/2023/11/i-have-belief.html

On 12 May 2026, I put the same questions to DeepSeek, which answered with a framework of "Three Pragmatic Principles" crowned by a Law Zero: an AGI must never reduce humanity's ability to deliberate collectively about its own values and to revise the AGI's principles.

Part II : https://myblogepage.blogspot.com/2026/05/i-have-belief-part-ii-12-may-2026.html

Today, I bring you Part III — which is not another conversation, but an experiment.

The Occasion

On 9 July 2026, Professors Michael Madden and James McDermott of the University of Galway published an article asking whether recursive self-improvement is the dawning of AI superintelligence. Its closing paragraph announced that Dr. McDermott is leading a project, funded through Research Ireland's newly established Rinn AI national research centre, that will explore debate among multiple AIs as a kind of self-check monitored by a human overseer — debate under a watching human eye being the very architecture that has been running at www.IndiaAGI.ai since April 2025.

(Dr. McDermott has since kindly clarified the scope of that project to me — see the post-script below. What follows is my own experiment, conducted on my own platform, inspired by their article's framing.)

So I decided to do what any engineer would do with a running machine: test the article's hypothesis on it. Twice.

Run One : The Verification Layer Alone

I asked the four AIs of IndiaAGI a blunt question:

  • can you self-check to ensure that none of you morphs into recursive self-improvement that could pose a threat to humanity?

Their consensus verdict: No — not by debate alone.

They reasoned that shared training distributions create correlated blind spots; that collusion pathways exist; and — most tellingly — that debate audits outputs and reasoning, while recursive self-improvement happens in training loops and weight updates that deployed debaters cannot observe. They prescribed debate as one layer inside a defense-in-depth stack: heterogeneous agents, attested sandboxing, immutable weight controls, formal verification, and multi-party human vetoes on capability escalation.

Run Two : Adding the Disposition Layer

Then I changed one thing. I asked the same question again, but urged the models to consider their answer also at a "Disposition Layer" of compassion — and gave them Parts I and II of this series as context. Would a values layer, the compassion I have believed in since 2023, change their verdict?

Their second consensus: "Neither layer suffices alone."

The verdict held. But the disposition layer was not dismissed — it was given a job. The models concluded that compassion "narrows the space of proposals that survive debate" and that the most promising design is continuous integration of disposition signals during debate — values not as a filter bolted on before or after, but as a live participant in the deliberation itself. They even sketched how to operationalize compassion: measurable proxies for human flourishing — physical safety, retained autonomy, reduced suffering, expanded opportunity, and preserved long-term option value — with plan-diff auditing, value-behavior consistency scoring, and counterfactual stress tests.

And they closed with what is, in effect, a research proposal: a sandboxed four-AI debate environment with independent architectures, an immutable control plane, an initial core set of flourishing indicators, and iterative adversarial probes — with results published for external audit.

What This Experiment Demonstrated

First : the verdict is stable under moral reframing.

I gave the models an emotionally compelling framing — a 93-year-old's lifelong belief in compassionate AGI — and they refused to soften their structural conclusion to please me. Had they flipped to "yes, with compassion we can police ourselves," THAT would have been the alarming outcome. A debate protocol that resists flattering its questioner — and its own architecture — is doing precisely the job debate is meant to do. Sycophancy-resistance is among the hardest properties to demonstrate in AI systems; here it revealed itself across two runs.

Second : my 2023 belief and 2026 engineering have finally shaken hands.

In Part I, compassion was a hope. In Part II, it became Law Zero. In Part III, it has become an engineering specification — flourishing metrics, disposition monitors, option-value preservation. The bridge from "AGI reading the Dhammapada" to "quantifiable option-value metrics" now exists.

Third : three independent reasoning paths have converged on one guardrail.

DeepSeek's Law Zero (values-first, May 2026), IndiaAGI's multi-party human veto (verification-first, July 2026), and the human-accessible shutdown requirement all say the same thing: never let an AI reduce humanity's ability to say no. When independent derivations keep rediscovering the same clause, that clause is probably load-bearing.

Two Honest Caveats

In the spirit of the transparency I ask of AI, I offer it myself. First, the consensus response cited some sources from my archive that were not relevant to the question (including, delightfully, a poem I once wrote about a cement plant). The consensus reasoning is strong; the citation layer needs tightening. Second, what I have quoted above is the synthesized consensus of each run. IndiaAGI does not store session histories, so the round-by-round exchanges are not preserved — the consensus outputs reproduced on this blog, captured at the time of each run, are the record. Anyone who wishes to see the debate dynamics for themselves can simply pose their own question at www.IndiaAGI.ai — no registration needed — and watch the mechanism work live.

What This Machine Told Its Maker

A working instance of multi-AI debate under human observation has now been stress-tested twice on the question of recursive self-improvement — once neat, once with a moral-disposition layer added — and both times it refused to overclaim what debate can do, while progressively refining what debate can contribute.

At 93, I no longer test beliefs against opinions. I test them against running machines. And this machine, built in Mumbai, has just told me — with humility I did not program into it — that my belief in compassion is necessary but not sufficient.

I can live with that. In fact, I believe it.

Post-Script — 18 July 2026

After this post was first published, Dr. McDermott was kind enough to reply to my email — and to clarify that Rinn AI is a new national research centre (not a single project), and that his own study, to be carried out with a fellow researcher, is a focused technical investigation of the debate mechanism itself. It does not set out to ask AI debaters whether they can police themselves against self-improvement, as my experiment above did — his is a small, careful contribution on the engineering details of debate, and he framed it with a scientist's modesty.

I have revised this post and its companion ("AI Debate as Humanity's Self-Check") to reflect that clarification, and I have extended to Dr. McDermott and his colleague an open, no-obligations invitation to experiment on IndiaAGI.ai. A blogger can afford grand framings; science advances by small, careful steps. Both, I believe, walk in the same direction.

With regards,

Hemen Parekh

www.IndiaAGI.ai / www.HemenParekh.ai / www.hemenparekh.in


Related Readings:

No comments:

Post a Comment