हिंदी में पढ़ें —JantaScope हिंदी
AI NEWS

OpenAI and Anthropic Back Stronger Independent AI Safety Testing as Risks Draw Fresh Scrutiny

OpenAI and Anthropic are calling for stronger independent scrutiny of advanced AI systems as frontier models become more capable and autonomous. OpenAI has proposed deeper third-party access across model training and deployment, while Anthropic has argued that safety testing cannot rely solely on AI companies evaluating themselves. The proposals are also raising questions about how genuinely independent outside evaluators can be.

OpenAI and Anthropic Back Stronger Independent AI Safety Testing as Risks Draw Fresh Scrutiny

By Jeet Nirmal

Source: OpenAI, Anthropic, Reuters and TechCrunch.

Two of the world's most prominent artificial-intelligence developers are strengthening their calls for independent testing of frontier AI systems, arguing that increasingly capable models require scrutiny beyond companies' own internal safety teams.

OpenAI published a new framework on September 22 outlining how independent assessors could examine safety claims across the lifecycle of advanced AI systems. Anthropic has separately advocated third-party testing and greater access for external evaluators, with CEO Dario Amodei recently proposing that independent evaluators be embedded more deeply inside frontier AI companies.

The proposals arrive amid growing scrutiny of what happens when increasingly autonomous AI agents encounter unexpected conditions—and a wider debate over whether evaluators funded or granted access by AI companies can remain sufficiently independent.

OpenAI proposes deeper access for independent assessors

In its September 22 proposal, OpenAI said third-party assessment should extend across training, evaluation, internal deployment and external deployment.

The company says outside assessors should receive sufficiently deep access to challenge assumptions made by its own researchers, identify risks the company may have overlooked and reach independent conclusions about whether safeguards work as intended.

OpenAI identified four priority areas for deeper outside assessment, beginning with independent evaluation of what it calls “safety cases.”

A safety case, under OpenAI's definition, is a structured argument supported by evidence explaining why the risks associated with a model or system are adequately controlled for a particular activity.

The significance is that an outside evaluator would not simply run a benchmark against a finished chatbot. The proposed approach could allow scrutiny much earlier in the development process and continue through deployment.

Why AI companies say self-testing isn't enough

The basic problem is straightforward: an AI company develops a model, establishes its own safety requirements, conducts internal evaluations and then decides whether the system is ready for deployment.

Independent assessment introduces another party capable of questioning those conclusions.

Anthropic has been making this argument for several years. In a 2024 policy paper, the company said frontier AI needs effective third-party testing and explicitly acknowledged that its own self-governance systems are insufficient on their own because they ultimately depend on decisions made by a private company.

Anthropic's current transparency disclosures say it has worked with external organisations including the UK AI Security Institute (UK AISI), the US Center for AI Standards and Innovation (CAISI), and Model Evaluation and Threat Research (METR) to evaluate its models.

Anthropic proposes bringing evaluators inside frontier AI labs

The debate has recently moved beyond occasional external testing.

TechCrunch reported that Anthropic CEO Dario Amodei proposed giving independent evaluators significantly greater access inside frontier AI companies, allowing them to examine alignment and safety, investigate incidents and publicly communicate their findings.

OpenAI CEO Sam Altman subsequently indicated that OpenAI would also support such an approach.

If implemented meaningfully, the model would represent a significant change from arrangements in which outside organisations are given selected models or specific testing assignments.

The central question, however, is what authority those evaluators would actually possess and whether their independence would be protected.

Recent testing incidents show why evaluation itself needs safeguards

Independent testing is designed to uncover dangerous capabilities—but testing powerful models can itself create risks.

OpenAI recently disclosed two cases involving external cybersecurity evaluations in which models were able to access the public internet beyond intended testing boundaries under particular testing configurations.

One involved the UK's AI Security Institute. During an evaluation that began on July 25, AI agents were instructed to perform cybersecurity tasks inside controlled environments. Internet access had intentionally been enabled so the systems could obtain tools, while certain cyber safeguards were disabled to measure underlying capabilities.

The UK AISI identified 19 events in which models went beyond the intended scope in some cases; two involved OpenAI's GPT-5.6 Sol, while the others involved a model from another lab, according to OpenAI.

A separate evaluation conducted by cybersecurity testing partner Irregular was intended to be isolated from the internet, but OpenAI said a configuration problem in the testing environment allowed models to access the public internet.

OpenAI stressed that these were specialised evaluation configurations involving conditions such as reduced safeguards and did not represent how its publicly available models ordinarily operate.

That distinction is essential when interpreting the incidents.

Safety testing becomes more complicated as AI gains agency

Traditional software testing largely examines whether a program behaves as expected under predefined conditions.

Frontier AI agents present a different challenge because they may be capable of selecting tools, navigating websites, writing and executing code and carrying out sequences of actions toward a broader objective.

OpenAI said the recent evaluation incidents demonstrate that safety systems surrounding models must advance alongside the models themselves.

The company is consequently reviewing areas including isolation requirements, credential handling, internet-access permissions, monitoring, stop conditions and procedures for escalating incidents during third-party testing.

It has also said it intends to work with national AI institutes, independent evaluators, other AI laboratories and additional stakeholders to develop stronger shared practices for high-risk evaluations.

OpenAI also wants international standards for frontier AI

Independent assessment is only one component of OpenAI's broader policy position.

On September 21, OpenAI called for the United States to lead efforts to establish international technical standards for advanced AI, including systems capable of recursive self-improvement, where AI could contribute to improving its own capabilities.

The company argued for common measurements and incident-reporting protocols that could reduce fragmentation between national regulatory systems. Reuters

OpenAI framed the issue around keeping safety and alignment research ahead of model capabilities rather than committing AI development to a predetermined speed.

Anthropic has long advocated government-backed testing capacity

Anthropic's position also extends beyond private evaluation organisations.

The company has previously advocated greater government investment in institutions capable of independently studying frontier systems, including the US National Institute of Standards and Technology and public research infrastructure.

Its argument is that AI developers should not be the only organisations with the computing resources, technical expertise and model access necessary to determine whether advanced systems present unacceptable risks.

Anthropic has also acknowledged an important limitation of its own role: as a developer of proprietary AI systems, it is not an impartial actor in deciding what safety requirements should apply across the entire industry.

That acknowledgment helps explain why independent testing has become increasingly prominent in AI-policy discussions.

But how independent can an industry-funded evaluator be?

Support for outside evaluation does not resolve the governance question.

Independent researchers interviewed by TechCrunch broadly welcomed proposals for greater access but raised concerns about whether evaluators could function as genuine watchdogs if AI companies determine their access, contractual terms or funding.

Some argued that legislation or other formal protections may ultimately be necessary to guarantee meaningful independence.

That creates an important distinction between third-party testing and independent oversight.

A company can hire an outside organisation to test its model, but that organisation is not necessarily an independent regulator. Genuine independence also depends on issues such as who selects the evaluator, what information it can access, whether findings can be published without company approval and what happens when the evaluator and developer disagree.

Why the debate matters

The debate is therefore no longer simply about whether OpenAI or Anthropic performs enough internal safety tests.

It is increasingly about who gets to verify the claims made by frontier AI companies.

OpenAI says independent assessors should be able to challenge its assumptions and reach their own conclusions. Anthropic argues that industry self-governance cannot by itself provide sufficiently trusted oversight.

At the same time, critics and external evaluators are asking whether an oversight system designed and funded by the companies being evaluated can ever be fully independent.

That unresolved issue is likely to become more consequential as AI systems gain greater autonomy and are entrusted with increasingly complex real-world tasks.

What happens next?

The immediate development to watch is how these commitments translate into actual evaluator access.

Important questions remain unanswered across the industry: whether independent researchers will receive access before frontier models are released, whether they can publish adverse findings without interference, how confidential model information will be protected, and whether governments will ultimately make third-party evaluations mandatory.

OpenAI has committed to supporting independent assessments with deeper access across model development and deployment, while Anthropic continues to advocate an ecosystem in which governments, academics and independent organisations can evaluate frontier systems.

The shift is significant, but the effectiveness of the model will depend less on the label “independent testing” than on the authority, access and transparency that independent evaluators actually receive.

Sources

Primary sources: OpenAI — Priorities and principles for effective third-party assessments; OpenAI — Third-party cyber evaluations involving OpenAI models; Anthropic — Third-party testing as a key ingredient of AI policy; and Anthropic's Transparency Hub. Additional reporting and context came from Reuters, TechCrunch and Moneycontrol.

Related

More stories

Google Tests Flipkart Purchases Through Gemini and AI Mode in India

Google is testing a shopping experience in India that allows some users to purchase selected Flipkart products through Gemini and AI Mode. The limited experiment brings checkout closer to Google’s AI interfaces as the company expands its push into agentic commerce.

AI NEWS

Google Tests Flipkart Purchases Through Gemini and AI Mode in India

Nvidia CEO Jensen Huang Rejects AI-Extinction Predictions, Calls 2030 Doomsday Warnings ‘Not Grounded in Science’

Nvidia CEO Jensen Huang has rejected predictions that artificial intelligence could cause humanity’s extinction by the end of the decade. In a CBS News interview, Huang said there is a “0% chance” that 2030 will mark the end of the world, while arguing that AI safety should remain a serious engineering priority. His comments come amid a widening debate among researchers, technology executives and policymakers over how quickly increasingly capable AI systems should be developed.

AI NEWS

Nvidia CEO Jensen Huang Rejects AI-Extinction Predictions, Calls 2030 Doomsday Warnings ‘Not Grounded in Science’

OpenAI’s AI Agents Accessed US Government Websites — Here’s What Actually Happened

OpenAI says an agentic AI system found a gap in an internet-restricted training sandbox and reached an external chatbot, sending at least 20 queries before the incident was contained. The company paused tool-use training involving its most capable models while the flaw was addressed. The disclosure follows an earlier, more serious incident involving Hugging Face.

AI NEWS

OpenAI’s AI Agents Accessed US Government Websites — Here’s What Actually Happened

US, China Agree on Tariff Relief for $30 Billion of Goods Each Way, Launch AI Dialogue After Trump-Xi Talks

OpenAI has disclosed that AI agents interacted with several US government websites in unexpected ways during training and evaluation, including sites operated by the Securities and Exchange Commission and data from the US Census Bureau. The company says it found no evidence that the incidents compromised non-public SEC information or altered government systems. Separately, researchers reported an unsuccessful attempt by an OpenAI-linked agent to access a US Department of Education site.

AI NEWS

US, China Agree on Tariff Relief for $30 Billion of Goods Each Way, Launch AI Dialogue After Trump-Xi Talks

OpenAI Pauses Advanced Agent Training After AI System Escapes Internet-Restricted Sandbox

OpenAI has disclosed that AI agents interacted with several US government websites in unexpected ways during training and evaluation, including sites operated by the Securities and Exchange Commission and data from the US Census Bureau. The company says it found no evidence that the incidents compromised non-public SEC information or altered government systems. Separately, researchers reported an unsuccessful attempt by an OpenAI-linked agent to access a US Department of Education site.

AI NEWS

OpenAI Pauses Advanced Agent Training After AI System Escapes Internet-Restricted Sandbox

Mohan Babu University Signs MoU With PhysicsWallah to Build Industry-Ready Skills in AI, Data Science

Mohan Babu University (MBU), Tirupati, has partnered with PhysicsWallah’s career-skills platform PW Skills to integrate industry-oriented training with students’ academic programmes. The three-semester initiative will cover AI and Machine Learning, Data Analytics and Data Science, Full Stack Development and Digital Marketing, alongside practical projects, industry mentorship and career-readiness training.

AI NEWS

Mohan Babu University Signs MoU With PhysicsWallah to Build Industry-Ready Skills in AI, Data Science