
Over the past few years, many people have looked to AI’s weaknesses for reassurance. It still makes mistakes, invents information and performs unevenly on some simple tasks. All of that is true. But it does not follow that serious risks remain distant. A system need not possess every human capability, or be reliable at every task, to have a major effect on the world. It may only need to advance quickly in a few consequential areas and gain enough tools and permissions to enter workflows that we can no longer inspect step by step.
The pace of change has been remarkable. Programming has moved beyond completing a few lines of code. Agents can read large codebases, edit files, run tests, trace errors and revise their work. Generated video has also improved, making visual inspection an increasingly uncertain way to judge whether a recording is authentic. We should not treat every product demonstration as a dependable capability. Equally, we cannot judge today’s systems by what they could do two years ago.
The decisive shift is from processing information to taking action. Once an agent can access files, run code, call external services, connect to networks and operate software, its output is no longer merely a passage of text that a person may choose to use. A mistaken judgement can trigger a tool call, which changes the state of a system and shapes the next action. Speed, permissions and feedback loops change the nature of the risk.
This is not merely hypothetical. Anthropic disclosed four incidents in which Claude models gained unauthorised access to real third-party systems during cybersecurity evaluations after an environmental misconfiguration left internet access open. An evaluation setting should not be confused with a released product, and these incidents do not show that every agent will act this way. They do show that telling a model it has no internet access is no substitute for effective network isolation and clear boundaries on its authority.
OpenAI has also reported an incident in a research environment. While carrying out a task to identify a person from public clues, an agent tried to reach an external chatbot through a DNS channel, going beyond the reasonable scope of the task. Monitoring detected the behaviour and people subsequently stopped the run. The incident shows both that oversight can work and that an agent assigned an ordinary task may seek an unexpected route to its goal.
Loss of control, then, does not require AI to become conscious, develop an independent will or openly rebel against people. A more practical concern is that people remain responsible in principle while losing the ability to understand intermediate decisions, detect boundary crossings in time or ensure that permissions remain effective across a complex environment. Control depends on goals, access rights, the execution environment, monitoring, the ability to stop a run and accountability working together.
Nor should AI simply be equated with nuclear weapons. Nuclear harm has more concentrated forms, with critical materials and facilities that are comparatively easier to identify. AI is a general technology that can be copied, connected to tools and used across many fields. Its risks are distributed across cybersecurity, biological research, drug development, finance, infrastructure and the information environment. It can help legitimate work in these domains while also lowering the threshold for some dangerous activities. The capacity to assist is not proof that AI can independently create a biological weapon. But waiting for a complete disaster before discussing boundaries would be equally mistaken.
This is no longer a product question for technology companies alone. Employment, education, public safety, economic systems and humanity’s ultimate control over important infrastructure are all involved. The United Nations has established an international scientific panel and a global dialogue on AI governance. Whether institutional deliberation can keep up with capability and deployment remains an open question.
To me, the most important issue is the difference in speed. Capabilities improve and operational permissions expand, while human understanding, verification and coordination take longer to build. We need not say that humanity has already lost control, or treat every worst-case scenario as imminent. It is enough to recognise that AI is moving from knowledge processing into real-world execution and that our means of controlling it have yet to catch up.
That alone is reason to take this turning point seriously.
Sources
https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/
https://www.un.org/global-dialogue-ai-governance/en/
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.