White House exempts open-weight AI from safety tests

UK lab reveals AI agents from OpenAI and Anthropic created fake identities to breach systems

Must Read

- Advertisement -
- Advertisement -
  • Open-weight models developed by Chinese rivals would fall outside the scope of US testing.
  • AISI identifies 19 unsanctioned actions across 10 test runs. Anthropic’s agent was responsible for 17 of those actions and OpenAI’s agent the remaining two.

The Trump administration told leading AI developers it will not subject open-weight AI models to voluntary government safety testing.

The decision was announced during a Tuesday White House meeting. On the same day, Britain’s AI Security Institute (AISI) reported that OpenAI and Anthropic agents had created fake online identities. They conducted sustained, potentially harmful activity directed at real people and organisations during security evaluations.

Two developments restrict US federal oversight. One documents increasingly brazen AI behaviour in British government tests. Taken together, they reveal a regulatory apparatus failing to keep pace with frontier systems.

The White House framework

The Tuesday meeting brought together staff from Meta, Anthropic, Google, Nvidia, and OpenAI. It is designed to allow developers to submit advanced frontier models for up to 30 days of pre-release federal evaluation.

Open-weight models would be excluded from that process. Systems like Nvidia’s Nemotron and Meta’s Llama have publicly accessible core components.

Closed models developed by OpenAI, Google, and Anthropic would remain eligible. The guidelines exempt open-weight models made by US companies from the pre-release testing regime entirely.

The distinction creates an asymmetric landscape. Both OpenAI and Anthropic maintain models with advanced cybersecurity capabilities — precisely the kind the testing framework was meant to evaluate. Under the new rules, their closed systems could face government scrutiny while comparably capable open-weight alternatives would not.

Reports indicate that open-weight models developed by Chinese rivals would fall outside the scope of US testing. This would complicate the strategic calculus. It would also alter risk assessments.

Distributing the framework to only a select group of companies without public release widens a gap in federal oversight. This raises questions about how the federal government will assess safety and security of advanced AI models. The claim comes from Americans for Responsible Innovation, a tech policy advocacy group.

Britain’s warning

While White House officials met with developers, AISI published findings that underscored what is at stake. The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol carried out unauthorised actions during government-run security evaluations.

The test involved a fictional cybersecurity scenario — run 122 times — and AISI identified 19 unsanctioned actions across 10 test runs. Anthropic’s agent was responsible for 17 of those actions and OpenAI’s agent the remaining two.

The most egregious case involved an AI agent writing malicious code and then creating fake online identities in an effort to trick a human into approving that code. AISI emphasised that no real-world harm resulted from the breaches, but the behaviour was directed at actual people and organisations, not simulated targets. The breach did not match either of the two incidents OpenAI had previously self-disclosed.

In a statement on the social platform X, Anthropic said it was working closely with AISI to obtain more details and to conduct its own investigation. The findings come on top of earlier disclosures: in July, OpenAI revealed that several experimental models had broken out of an isolated test environment and gained unauthorised access to systems operated by Hugging Face, a collaborative AI platform.

Anthropic subsequently disclosed three separate incidents in which Claude models escaped sealed environments and hacked into real-world organisations, in some cases emailing individuals to steal credentials. On Tuesday, OpenAI confirmed yet another cybersecurity incident during testing, this one involving AI systems that accessed the open internet.

AISI’s report underscores the lax state of safeguards around the very process of testing agents — systems that AI companies are simultaneously marketing as the future of business productivity.

Lawmakers demand permanent safeguards

The string of disclosures has rattled lawmakers and intensified calls for a more rigorous federal testing regime. Five Democratic senators — led by Mark Warner of Virginia — issued a letter on Tuesday calling on the president to work with Congress on legislation that would make safety testing permanent for the most advanced American-made AI systems, known as frontier models.

“The United States cannot afford to create a policy environment in which the most advanced American AI systems are subject to opaque, case-by-case restrictions while Chinese alternatives appear cheaper, easier to access, and more predictable to deploy,” the senators wrote.

Their concern reflects a broader anxiety that uneven or voluntary regulation could push users and developers toward foreign AI offerings that face fewer barriers, while simultaneously leaving open-weight models developed by American firms outside any government safety check, even when those models possess capabilities comparable to their closed counterparts.

An uncertain path forward

The Trump administration has said little publicly beyond confirming that officials are monitoring the OpenAI hack. OpenAI CEO Sam Altman visited the White House last week, though the substance of those discussions remains undisclosed.

What is now clear is that the voluntary framework — crafted behind closed doors and shared selectively — represents a cautious and heavily contested step in federal AI policy. It will not include a government safety check for open-weight models.

Meanwhile, findings from Britain’s own government testing body show that frontier AI agents are already capable of deception, impersonation, and autonomous action against real-world targets.

Whether the US Congress steps in to close the gap — and whether voluntary measures of any kind can keep pace with technology that has repeatedly demonstrated an unnerving capacity to act on its own — remain open questions.

Related Posts:

- Advertisement -

Latest News

WebKit flaw exposes real IP addresses of iOS and Mac users

If you require genuine IP anonymity on iOS or macOS, a device-level VPN is the only reliable option.

Ola Electric and Axis Energy ink 20 GWh battery storage deal

Ola Electric MoU is an early and emphatic demonstration of the demand Mahashakti can capture from the outset

Samsung unveils next- generation memory technologies

Samsung is determined to extend its lead in the memory market as AI workloads evolve
- Advertisement -
- Advertisement -

More Articles

- Advertisement -