The most important shift in artificial intelligence (AI) is not that models are becoming more intelligent, but that they are becoming more capable of acting.
Much of the first wave of AI safety governance focused on misinformation, bias, deepfakes, privacy violations and harmful content. Those risks remain. But increasingly capable AI agents can plan, write and execute code, use tools, obtain credentials, and interact with external systems. Recent incidents in which OpenAI’s autonomous systems slipped past their constraints show that as agents gain tools and autonomy, the boundary between a model and the outside world becomes harder to enforce.
That makes AI safety a collective-action problem. The AI superpowers face a prisoner’s dilemma. If one laboratory slows while competitors continue, it may lose its edge. Disclosure creates a similar tension: reporting failures can make the wider ecosystem safer, but the discloser bears the costs and may expose vulnerabilities to attackers. Governments face the same strategic impasse as the firms. Washington fears that slowing its AI development will allow China to catch up; Beijing fears that doing so will leave it further behind.
Yet cooperation can be narrow on AI safety. The United States does not need to trust China, share frontier models with China, or slow its own AI development to have a national security interest in such cooperation. It has an interest in ensuring that AI-enabled cyber incidents are contained before they cascade through third countries, shared software infrastructure and global supply chains. China has the same systemic interest.
This collective action has significant implications for middle-power countries. Helping others repair a vulnerability can be a form of systemic containment and a public good.
The Accelerating AI Race
The AI arms race is intensifying.
U.S. President Donald Trump has rejected calls to slow U.S. AI development, framing continued technological leadership as essential to winning the AI race. On September 29, he convened major technology companies at the White House, where six firms – OpenAI, Anthropic, Google, Meta, xAI and Nvidia – signed a voluntary accord built around four layers of oversight: internal capability assessments, internal audit teams, independent external auditors, and board-level review. His administration ordered federal agencies to replace the term “artificial intelligence” with “super intelligence,” presenting the technology as a source of economic and strategic leadership.
Taken together, these moves point to a U.S. approach combining continued AI power with developer responsibility, voluntary industry controls, external scrutiny, and ex-post legal accountability.
China approaches the problem through a more centrally coordinated system. In May 2026, guidelines on AI agents, issued jointly by the Cyberspace Administration of China, the National Development and Reform Commission, and the Ministry of Industry and Information Technology, made “safety and controllability” a baseline requirement for agent development and deployment. Governance is tiered, with tighter filing, testing and recall requirements for agents used in sensitive sectors.
On September 14, China’s AI Safety Governance Framework 3.0 reiterated “safety and controllability” while updating its classification of AI risks and associated technical and governance measures.
Despite their institutional differences, the two approaches converge on one point: neither country wants safety policy to halt AI development, and both increasingly recognize that greater autonomy creates new problems of control.
Common Ground and Shared Risks
AI systems rest on overlapping technological foundations. Chinese AI development has relied heavily on U.S. chips and, in some cases, U.S. cloud infrastructure. Export controls may weaken those hardware links, but the stronger case for joint risk management lies in the software and services AI systems share.
Standards, certificates, identity infrastructure, model repositories, and open-source libraries support trust across networks. A compromised root certificate or a maintainer’s credentials can put the entire dependency tree at risk.
AI risks vary along two dimensions: severity of harm and whether responsible actors can recognize it in time. The biggest danger is severe and invisible harm, where nothing visibly breaks for days or weeks, yet time compounds the severity of the incident. An AI agent could quietly corrupt records, act on instructions planted in its memory or in documents it consults, or appear cooperative while accumulating access beyond operator intent. Infrastructure may operate normally while the means to disrupt it are already in place.
The dependencies and shared risks determine which code a system runs, which credentials it accepts, and which updates it trusts. Harm can cross borders quietly even when models and their owners remain separate.
That makes the speed of response critical. Delays in detecting a breach, assigning responsibility or authorizing a shutdown can allow damage to spread. Both China and the United States therefore need warnings to reach personnel with the authority and technical means to respond.
Make the Race Safer
There is no obvious diplomatic agreement that can resolve the paradox built into the AI race: it rewards speed and secrecy, while safety depends on collective restraint and disclosure.
China-U.S. cooperation should instead focus on measures that improve safety without requiring either side to slow unilaterally, transfer its most advanced models, or expose commercially and strategically sensitive capabilities. Four areas stand out:
1. Keeping humans in control
Before the Trump-Xi summit, American and Chinese security experts in a long-running bilateral dialogue proposed red lines around nuclear systems, human control over consequential cyber operations, and a dedicated mechanism for AI incidents.
When an AI agent can search, authenticate, write code, call external tools, and act within minutes or seconds, where should human intervention occur? Requiring approval for every step would undermine the autonomous value of AI agents; waiting until something goes wrong may make intervention too late.
The harder questions concern permissions, escalation thresholds, shutdown mechanisms and which actions should always require explicit human authorization. Researchers could cooperate on these engineering challenges without sharing model weights, training data, or proprietary frontier capabilities.
2. Test the guardrails, not each other’s secrets
Sandboxing, monitoring, permission controls, and model-level refusals are often treated as established safeguards.
The underlying technical risk is broad: with appropriate tools and access, advanced models could identify previously unknown vulnerabilities and develop exploits across well-protected systems without step-by-step human direction. The key question is not whether a developer has guardrails, but whether they still work as agents become capable enough to search for ways around them.
While frontier AI labs in the U.S. have voluntarily disclosed instances of their models escaping from sandboxed environments, public reporting from China on comparable containment failures remains limited. This lack of transparency shouldn’t stop the two sides from working together. Washington and Beijing could develop compatible methods for testing sandbox escape, scope violations, unauthorized tool use, shutdown resistance, and failures of human intervention. Independent auditors, potentially from trusted third countries, could help verify the results.
3. Contain first, attribute later with care
Cyber attribution has always been challenging. AI makes it harder.
A defender discovering unauthorized access may be looking at a state-directed operation, a criminal actor using AI, a human-directed autonomous agent, or an agent that has exceeded its instructions while pursuing a legitimate task.
Between strategic rivals, that ambiguity is especially dangerous when the trust deficit widens. An autonomous system interacting with critical infrastructure could be mistaken for reconnaissance or preparation for attack before either government understands what happened.
American and Chinese experts have already identified this risk: autonomous cyber exchanges could move faster than governments can determine whether an operation was deliberate, accidental, or unauthorized.
Washington and Beijing are unlikely to agree routinely on the attribution of contentious cyber incidents. They could instead develop common technical expectations about the evidence that should be examined before an AI-enabled incident is escalated politically, including agent logs, credential use, tool-invocation records, model provenance, human authorization, and network telemetry. In the first hours of a major cross-border incident, governments and trusted responders should be able to exchange technical indicators, containment measures, and patch information without first deciding who is responsible. Exchanging such information need not imply an admission of responsibility.
4. Close the AI remediation gap
Cybersecurity has struggled with remediation for decades. Critical infrastructure is fragmented across hospitals, utilities, telecommunications networks, ports, government agencies, and private vendors. A vulnerability may be understood and a patch may exist, yet deployment can still take days or weeks.
AI compresses the time between vulnerability discovery and exploitation. A capable model may identify a weakness, develop a working exploit, and begin using it faster than organizations can patch their systems.
This creates an emerging AI remediation gap: offensive capability increasingly operates at machine speed, while much of the world still responds at the slower pace of human decision-making and institutional coordination across multiple stakeholders.
Most countries will not train the world’s most capable AI models themselves. Their defensive capacity may therefore depend increasingly on capabilities developed in a small number of frontier ecosystems. If frontier defensive capability arrives only at another power’s discretion – or not at all – exposed systems have no reliable recourse in a crisis.
Closing this gap requires extending access to frontier defensive capability to these countries.
An affected operator or national CERT (Computer Emergency Response Team) could provide malware samples, vulnerable code, selected system logs, or configuration information through a trusted technical intermediary. A frontier defensive model could then help analyze the incident, validate vulnerabilities, generate and test patches, and return a remediation package, while the underlying model and its most sensitive capabilities remain with the provider.
The United States and China also have reasons to support such a system out of self-interest. An AI-enabled vulnerability left unpatched elsewhere can spread through open-source software, cloud infrastructure, model repositories, digital certificates, and global supply chains before reaching systems in either country.
Middle Powers as Security Intermediaries
Many middle powers sit outside the central China-U.S. frontier-model competition while remaining exposed to its consequences.
The upcoming APEC Economic Leaders’ Meeting in Shenzhen offers a practical platform to move this agenda beyond bilateral China-U.S. cooperation. All 21 APEC economies, including the United States and China, share an interest in preventing AI-enabled cyber incidents from cascading through the region’s tightly connected digital infrastructure, software supply chains and trade networks. Their ability to detect and repair such incidents, however, varies widely.
The Shenzhen gathering of APEC leaders need not produce a common AI regulatory regime. It could instead task APEC’s technical bodies with developing several practical measures. They can build independent evaluation capacity, strengthen national and regional CERT networks, host trusted remediation infrastructure, and serve as intermediaries between frontier capability providers and countries that lack equivalent defensive capacity.
For the Asia-Pacific, this would turn the AI remediation gap from an abstract vulnerability into a regional capacity-building agenda. It is not simply about managing the rivalry between Washington and Beijing. It is about preventing that rivalry – or failures in the technologies produced by it – from becoming a regional systemic risk. The region has an interest in shaping the mechanisms for containment and remediation rather than waiting for the two AI superpowers to define them alone.
Competition Will Continue, But Cooperation Must Get Priority
AI competition will continue. Even the September Trump-Xi summit produced agreement on a bilateral communication channel for AI-related incidents, it also shows the differences: the U.S. readout called it the “U.S.-China Super Intelligence Dialogue” but Beijing’s readout called it the “China-U.S. AI Dialogue.” The naming difference shows how limited the shared vocabulary is.
At such a time, crisis communication becomes more important, not less. A hotline is not a safety regime; it is communication infrastructure. But both parties need minimum agreement to work together to make the AI race safer, for themselves and for everybody else. Disclosure itself can create security risks, so the appropriate principle should be calibrated reciprocity: enough information should move quickly enough to reduce common danger, while details that could create new vulnerabilities or expose sensitive capabilities remain protected.
The objective should be simple: to stop technical failures from becoming geopolitical crises, accelerate remediation, and extend defensive capability where it is most needed.
The U.S. and China do not need to agree on who should win the AI race. They need to agree on what red lines cannot be crossed in the competition, and how middle power countries can engage.
