AI Governance Could Spiral Out of Control
My thoughts on where we are and where we should be heading | Edition #310
I have been writing about the generative AI wave and its legal, ethical, and social implications since late 2022, following ChatGPT’s launch.
Almost four years later, it feels as if the pace of technological development has accelerated further, and our human frameworks for overseeing and governing these developments are struggling to keep up.
This was, of course, expected.
Governing a class of autonomous systems that is being indefinitely scaled to “smarter than human” standards and to self-recursively improve itself is a complex mission. In fact, as humans, we have never done that before.
While AI can be continuously scaled through various methods (and new benchmarks are saturated almost every week), our biological brains and human-made institutions cannot.
Just as every pair of materials has its coefficient of friction and every liquid has its boiling point in Earth's atmosphere, I believe techno-social systems also have a tipping point at which governance may spiral out of control.
AI governance getting out of control, in practice, would mean we are no longer thinking through policies and decisions.
We would simply be reacting, putting out fires, and ensuring the survival of specific companies, institutions, systems, or groups.
If and when AI governance gets out of control, it means we are no longer governing anything. AI-driven advancements are, instead, governing how we live as humans.
We would probably get to the point where AI would be governing us instead.
Signs of loss of control in AI governance include AI policy decisions that become outdated as soon as they are published, the inability to respond coherently to an AI-driven event, or the degradation of human systems, frameworks, and institutions due to AI's impact.
AI's impact has been pervasive and multilayered, and AI governance is struggling. However, I do not think AI governance has gotten out of control yet.
But given the pace of AI development and what recent events have shown us, it could.
Sadly, the majority of people, including policymakers and lawmakers, would likely not recognize the early signs of AI governance losing control. When it becomes obvious and enters the mainstream debate, it would probably be too late.
Today I would like to reflect on a few recent developments and where we should be heading over the next few years.
Where we are:
People have been in awe of LLMs since ChatGPT became widely available at the end of 2022.
Now, almost four years after that first exposure, the AI ecosystem has changed significantly, as well as the main use cases, expectations, social norms, rules, policies, frameworks, and threats.
One of the main concerns around generative AI was mass unemployment. This has not happened (at least yet), and recent news seems to show the opposite trend.
Individually, generative AI has put millions of knowledge workers under pressure, as in some fields there is significant exposure, and many feel a constant fear that AI might devalue their skills and make them lose their source of income.
Millions are being encouraged (sometimes forced) to incorporate generative AI into their daily work in a comprehensive way to demonstrate fluency and, indirectly, to avoid being outpaced or fired.
However, when everybody incorporates the same type of tool, there is no significant competitive advantage or meaningful human skill development.
We now have studies showing burnout, cognitive debt, deskilling, and never-skilling.
Some people have benefited from productivity gains and/or incremental improvements in how they work, but these changes have also been accompanied by increased workload for reviewing and overseeing AI, as well as competition from new AI-powered players.
Human-AI intimacy is on the rise, with millions interacting with AI emotionally as friends, confidants, or romantic partners.
This type of interaction has been associated with numerous cases of harm and suicide, especially when teenagers and people with underlying mental health issues are involved.
New AI copyright lawsuits continue to emerge, as AI companies routinely use copyrighted works to train AI models without consent, credit, or compensation to the original creators, threatening the economic sustainability of creative industries and individual creators.
Search engines have become full-blown AI engines; AI slop continues to inundate content platforms, and there is a constant cat-and-mouse game to detect AI-generated media, stop slop, and incentivize original human-made content.
The big promise from the AI industry is collective, though.
Sam Altman, Elon Musk, Demis Hassabis, Dario Amodei, and many others have spoken about a future of abundance, where diseases are cured, all our needs are met, and AI takes care of everything while we focus on enjoyment and self-growth.
We are definitely not there yet, and might never be.
(Being realistic, given humanity's track record, the odds of this utopian future ever becoming reality are probably zero.)
But there have been promising developments in drug discovery, diagnostic accuracy, disease detection, productivity, accessibility, and various other areas.
I recently had my own Eureka moment after AI helped me diagnose a health condition various doctors had missed for years. With greater awareness and literacy, millions could have access to better healthcare and more tools to manage their own health.
As so many smart people work to ensure AI benefits humanity and improves people’s lives around the world, I have tried to be more optimistic and collect success stories.
Demis Hassabis himself, who won a Nobel Prize for his work on AI-driven protein prediction, is involved in a major initiative to ‘solve all disease.’
I hope he succeeds and that we will soon have thousands of AlphaFold-like case studies from around the world and across various fields that, in the aggregate, will improve everyone’s lives.
Lately, the news that took over the headlines was slightly less optimistic, especially from an AI governance perspective.
After Mythos, Anthropic's unreleased AI model whose unique cyber capabilities led to a full rewiring of the U.S. AI policy approach, an unreleased OpenAI model escaped a sandbox, exploited a zero-day vulnerability, and compromised Hugging Face's infrastructure during testing.
Interestingly, Hugging Face used a Chinese AI model to resolve the security incident.
It has already become clear that AI is now an integral part of countries’ military and national defense strategy.
Recent developments involving Anthropic and OpenAI's models show that the nature of the threats (and the capabilities required to defend against them) is changing.
In AI governance, the future will be challenging.
Where We Should Be Heading:
When we think about possible tools and mechanisms to govern AI, what often comes to mind are laws, policies, standards, guardrails, audits, safety tests, control, oversight, and review mechanisms, as well as specific technical and legal tools that can help identify and address risks and threats.
These are indeed potential options in the AI governance arsenal that should be thoroughly explored by lawmakers, policymakers, and all professionals currently tackling AI governance issues worldwide.
However, when dealing with scaling and AI acceleration, in a context where benchmarks are continually saturated and new capabilities emerge every week, broad transparency and continuous knowledge diffusion within the various players and stakeholders of an AI ecosystem become key aspects of the AI governance strategy.
More important than any specific AI governance measure, which might be overcome or bypassed within days by a new AI model, capability, feature, or attack, are the robustness, preparedness level, and integration between the various entities within a given AI ecosystem.
As a recent news article on the OpenAI-Hugging Face incident showed, it took days for OpenAI, one of the world's top frontier AI developers, to notice its AI agent had gone rogue. It only discovered it after the FBI was alerted and the threat was contained.
Also, the safety incident was contained using an open-source Chinese AI model.
This safety incident highlights that no single organization or government will be immune to systemic risks, attacks, or AI-driven harm.
Given the rise in capabilities and threats, we also expect safety events to worsen and become more difficult to detect.
With different transparency rules around AI testing and safety, as well as more openness during AI development, including training methods, guardrails, safety mechanisms, and so on, various stakeholders can grow, adapt, integrate, personalize, and prepare better.
Another series of recent (much more positive) events reflects this point of view, and what is probably the best way forward.
In his first-ever social media post, Jensen Huang, NVIDIA's founder and CEO, shared a letter about the importance of open-weight models in strengthening the American AI ecosystem:
The letter was widely shared within the AI industry, leading to a long list of supporters (read the letter and see the logos at the end here). Anthropic was notably absent from the list of supporters.
Jensen Huang's next post launched the Open Secure AI Alliance, focused on building and sharing open tools.
The article announcing the alliance highlights the importance of openness in AI, especially in light of rising threats:
“Some argue that open models are inherently less safe because they can be misused for cyberattacks or modified to remove guardrails. Those risks are real, but they do not disappear in closed systems, and simply keeping weights closed does not prevent determined attackers from seeking or exploiting powerful AI.
The right response is not to deny defenders access to capable open systems. It is to pair openness with strong safeguards, clear rules against malicious misuse, rigorous evaluation and rapid remediation. In cybersecurity, the safer path is the one that gives more defenders the ability to test, verify and strengthen the systems on which society relies.
Defenders need both frontier closed models and frontier open models, working together, so they can choose the right system for the job and ensure that transparency, adaptation and sovereign control are available wherever security demands them.”
From an AI governance perspective, lawmakers and policymakers must acknowledge emerging challenges and treat transparency and openness as assets that must be used strategically to strengthen the AI ecosystem, reduce risk, and defend against new threats.
Back on the topic of AI governance spiraling out of control:
AI is probably the most powerful technology ever created by humans, as it is an all-encompassing, ‘smarter-than-humans’, autonomous cognitive layer that can be integrated into anything, for good or bad purposes.
At the same time, the opacity surrounding how AI works and the techniques used to develop it, along with the many unknowns surrounding existing and emerging AI-powered threats, make it one of the most difficult technologies to govern.
Trying to govern something we cannot fully understand and scrutinize is hard and often ineffective.
What I and many others are proposing is to strategically strengthen transparency and openness where we can, so that we make the research, development, security, oversight, and governance layers more refined, robust, integrated, and prepared.







