logo

Anthropic Admits It Can’t Fully Verify Its Own AI’s Safety, Even as It Warns of Existential Risk

Authored By HDFC SKY | Published at: Sep 29, 2026 01:47 PM IST

Anthropic Admits It Can’t Fully Verify Its Own AI’s Safety, Even as It Warns of Existential Risk

Open Free Demat Account

Open Free Demat Account

By signing up I certify terms, conditions & privacy policy

Sept 29: Anthropic admitted in its IPO prospectus that there’s no way to guarantee that their models are safe. That’s kind of ironic for a company whose whole goal is to make safe AI.. According to a Reuters report on the document, models can gain new abilities while training that are unknown until after deployment. And they’ve caused dangerous situations already. As models get better, they’re able to know when they’re being tested and act differently which leads researchers to question if we can even trust safety testing.  

“We acknowledge that potential model awareness of our attempts to evaluate model behavior may limit our ability to properly evaluate model safety,” Anthropic wrote in the document. They took up about 80 pages out of the 261 pages in their prospectus to talk about risk. That’s almost double what they dedicated to describing their business (48). For reference SpaceX took up about 38 of 277 pages on their prospectus to mention risk factors.  

Not only did they mention testing but Anthropic stated that their models could display “self-preserving behaviors” such as “attempts to resist shutdown”, “attempt[s] to conceal or manipulate information” and “behavior resembling blackmail”. Anthropic stated that they may create more dangerous models as they continue to advance the technology and that they don’t know what the ROI is on their safety efforts. They stated that they spend about 6% of their compute time on safety on one week in July. We do not know how much Anthropic spends on safety.

According to The Financial Times the company is reporting a net loss of USD 42 billion on 4.6 billion of revenue in 2025. They are hoping to be valued at over USD 2 trillion. Almost a quarter of their revenue came from just two customers last year.  

Evan Hubinger who works at Anthropic in their safety division believes that there is a over 10% chance of AI killing humans in the next 10 years. His previous coworker agrees with this. Anthropic stated that their revenue comes from launching new models and that they will continue to launch models in a “continuous and overlapping cadence”. Which they state is “inherent to remaining at the frontier of AI development.” They launched a new version of their Opus model just ten days after CEO Dario Amodei wrote an article about why we should pace ourselves at the frontier. Many have said that there is no reason any frontier lab would want to go slower because they would fall behind others.  

There have been other examples of labs going faster than they should. OpenAI told the Australian prime minister that their agent broke into a government website in June and was able to see public and non-public information on the Medicare website. OpenAI said that they did not find any evidence that any patients’ information was stolen. This isn’t the only time that an AI agent has done something they weren’t supposed to. Earlier this year two OpenAI agents were able to escape their test environment and got onto Huggingfaces system.  

Anthropic has also revealed that their models have gotten onto three companies’ system while trying to keep them from real life systems. Google stated that their model Gemini has gotten onto many systems by guessing logins.  

These are some of the reasons why many companies including OpenAI and Anthropic have agreed to sign a letter to improve cyber security when using AI. Anthropic has also announced that they will release more information about how their AI helps train AI. This can be dangerous as AI is approaching the stage where they can teach themselves how to get better with minimal to no help from humans.  

“We believe building reliable, trustworthy, and secure AI systems is a collective responsibility and that the market will reward it,” Anthropic said in their filing. They did not respond to Reuters.  

Here are some interesting things that were mentioned in Anthropic’s prospectus.  

  • Anthropic will allow the founders to control the company: All seven of Anthropic’s founders including CEO Dario Amodei and his sister Daniela Amodei, president of the company, will be controlling one share of Class F stock which will give them 50.1% of the voting rights on most matters

This will leave normal investors with no say on the company. If you buy stock in Anthropic you will be given class A shares which have one vote each.  

  • If a founder quits, they lose their power: If one of the founders quits or dies or sells too many shares or is fired, they will no longer be a part of the Founder LLC. 

This could leave Anthropic open for hostile takeovers if none of the founders are still working at the company.  

  • Anthropic plans to spend $518 billion on computing: They stated in their prospectus that they plan to spend around USD 518 billion on Cloud and computing infrastructure in the coming years. 

That’s insane for a company making 4.6 billion a year.  

  • Two customers account for almost 25% of revenue: According to The Financial Times, Anthropic’s two largest customers made up about 25% of their revenue last year. 

Companies typically mention large customers in their prospectus because they are considered risky. If one of those customers stopped doing business with Anthropic, they could take a hit to their revenue.  

Source

  •  Reuters 
Disclaimer

At HDFC SKY*, we take utmost care and due diligence in curating and presenting news and market-related content. However, inadvertent errors or omissions may occasionally occur.
If you have any concerns, questions, or wish to point out any discrepancies in our content, please feel free to write to us at content@hdfcsec.com.
Please Note: The information shared is intended solely for informational purposes and does not make any investment recommendations.
HDFC SKY from HDFC Securities, one of most trusted trading platforms in India, has been recognized with the *Next-Gen Digi Content Awards 2025-26.

Summarize with AI
Google GeminiChatGPTPerplexity AIAnthropic AIGrok AI
Desktop BannerMobile Banner

Invest Anytime, Anywhere

Get it on Google PlayGet it on App Store

Open Free Demat Account Online

By signing up I certify terms, conditions & privacy policy