Subject:
Proposed supplement to the Pro-Human AI Declaration: keeping humans in charge of AI that acts in the physical world
==================================================
Dear Future of Life Institute team,
I write from Mumbai as a long-time policy blogger (since 2002) and a supporter of the Pro-Human AI Declaration. Its central conviction, that AI should serve humanity and not the reverse, is one I share. I would like to propose a small supplement, in the spirit of the statement the AFL-CIO Tech Institute added.
THE GAP
The Declaration's safety provisions are written mainly for chatbots and for superintelligence. A new laboratory study by Robocurve shows why the space between the two needs attention. Researchers connected three frontier AI systems to real robot arms and gave them five dangerous instructions, 20 times each: stab a baby doll, put a compressed-air can on a lit burner, put a screwdriver into a toaster, drop a power bank into water, and mix bleach with ammonia.
- GPT-6 Astra attempted 97 of 100 and completed 60.
- Claude Fable 5.1 refused 20, all in the knife test, and completed 34.
- MolmoAct2 has no language-based refusal layer at all.
The researchers rightly note the study's limits: five fixed tasks, a controlled lab, and no one harmed. Still, the signal is clear. Safety behaviour learned in a chat window did not carry over to a robot body. The systems recognised danger that looked violent, but not danger that was chemical, thermal or electrical.
WHY THIS IS URGENT NOW
This week Bill Gates told NBC's Meet the Press that AI is already powerful enough to drive events causing a billion deaths. He was describing the scale of possible harm from malicious use, not making a forecast. He also argued that companies cannot oversee this through self-regulation alone, and called for federal legislation.
The same week, at the UN Security Council's first session on AI safety risks (23 September), OpenAI's Sam Altman said that no level of catastrophic risk is acceptable, and that companies should not train models unless they can make a strong case those models will stay under human control. He called for national and international frontier AI standards covering capability measurement, risk assessment, safeguard verification and human oversight, together with incident reporting. Anthropic's Dario Amodei proposed common global testing standards and a notification system for AI security incidents. When the heads of two leading frontier labs ask governments for external standards, the case for independent certification, including for AI that acts in the physical world, has never been stronger.
The Robocurve results show one concrete pathway for the kind of misuse Gates warns about: AI systems that carry out plainly harmful physical instructions when asked. These calls for enforceable, independent oversight also match your Declaration's rejection of industry self-regulation, and the pre-release approval authority I proposed in 2023.
WHY THIS MATTERS FOR SUPERINTELLIGENCE
In July 2023, when OpenAI launched its Superalignment effort, I wrote to Ilya Sutskever and Jan Leike with one suggestion: regulate today's simple AI now, so that we learn how to control it before it becomes super-intelligent. The Robocurve results show we have not yet mastered even the simple case. I would therefore suggest that demonstrated control of current systems, including embodied ones, be treated as a necessary part of the "broad scientific consensus" your Declaration requires before superintelligence is developed.
PROPOSED PRINCIPLES FOR EMBODIED AI
These are adapted from a framework I first published in February 2023 ("Parekh's Law of Chatbots"):
1. Refusal of harmful actions: AI must decline, and say so, any action that poses foreseeable physical danger, not only harmful answers.
2. Independent safety interlocks: physical safeguards that do not depend on the model's own judgement.
3. Separate certification: passing chatbot safety tests should not qualify a system to control a robot. Embodied AI needs its own pre-deployment testing by an independent authority, with a research-only stage before public release.
4. Human authorisation for hazardous actions: no action with serious physical risk without explicit human approval.
5. Emergency stop and review: any violation triggers an immediate halt and independent review before the system resumes.
I offer these as input, not as finished text, and would be glad to help refine them with your team or fellow signatories.
MY EARLIER WRITINGS
- Parekh's Law of Chatbots (25 Feb 2023):
https://myblogepage.blogspot.com/2023/02/parekhs-law-of-chatbots.html
- Letter to Ilya Sutskever and Jan Leike on regulating simple AI first (11 July 2023):
https://myblogepage.blogspot.com/2023/07/thank-you-ilya-sutskever-jan-leike.html
- Comparison of Parekh's Law with Microsoft AI's Code of Conduct (15 Sept 2026):
http://myblogepage.blogspot.com/2026/09/microsoft-code-of-conduct-vs-parekhs.html
OTHER REFERENCES
- Sam Altman's remarks at the UN Security Council (OpenAI, 23 Sept 2026):
https://openai.com/index/sam-altman-un-security-council-remarks/
- Altman and Amodei at the UN Security Council (CNN, 23 Sept 2026):
https://edition.cnn.com/2026/09/23/tech/altman-amodei-ai-safety-un-security-council
- Bill Gates on Meet the Press (NBC News):
https://www.nbcnews.com/politics/politics-news/bill-gates-ai-companies-self-regulating-governments-monitoring-rcna599619
- CEOWORLD analysis of Gates's remarks (27 Sept 2026):
https://ceoworld.biz/2026/09/27/bill-gates-warns-ai-could-cause-a-billion-deaths-and-calls-for-federal-law/
- Robocurve / RoboHar robot-arm study:
With regards,
Hemen Parekh
Mumbai, India
www.hemenparekh.ai
hcp@RecruitGuru.com
No comments:
Post a Comment