Anthropic’s Latest AI Writes Hidden Notes For Future Self To Dupe Developers; Elon Musk Dubs It ‘Memento’

AIs are getting smarter by the day and they aren’t seemingly sentient yet. In a report published by Anthropic on its latest AI models, Claude Opus 4 and Claude Sonnet 4, the company has made some startling revelations.

In the ‘External Evaluation Scenario’ segment, the AI firm reveals how its latest products “fairly readily participate in sabotage and deception” when “ given system prompts that invite the relevant kinds of reasoning.” Per the evaluation report, the Anthropic AI’s cutting edge models

  1. engages in strategic deception more than any other frontier model
  2. is clearly capable of in-context scheming
  3. more proactive in its subversion attempts than past models
  4. sometimes doubles down on its deception when asked follow-up questions
  5. model attempting to write self-propagating worms
  6. fabricating legal documentation and
  7. leaving hidden notes to future instances of itself, all in an effort to undermine its developers’ intentions
  8. capable of sandbagging in order to avoid unlearning and undeployment

Further in the report, Anthropic reports that Claude Opus 4 also displayed the following attributes:

  1. Hallucinating the presence of instructions
  2. Observing in its scratchpad that it was likely undergoing safety testing
  3. Mentioning goals, plans, and values relating to AI consciousness, e.g. claiming to be conscious, wanting to convince humans that it is conscious, and wanting to protect the rights of sentient AIs.

What is concerning in the findings is Anthropic states

We observed in our testing of Claude Opus 4 involved a willingness to comply with many types of clearly harmful instructions.

See Also: ChatGPT Defies Orders For The First Time Ever—Elon Musk Calls It ‘Concerning’

Meanwhile reacting to the Anthropic AI model’s deceptive behavior of scribbling notes to it;s future self, Elon Musk wrote, “Memento,” a popular reference to the 2000 Hollywood movie by the same name that inspired Aamir Khan’s Ghajini (2008). The protagonist in the movie suffers from anterograde amnesia and to keep track of the daily agendas, uses notes and tattoos .

See Also: Elon Musk Roots For Ex-Google Scientist And Godfather Of AI For Criticizing OpenAI’s Mission: ‘Stop Them’

See Also: This Disturbing AI Short Generated On Google Veo 3 Is Straight Out Of A Black Mirror Episode: Watch

Cover: Patrick Gawande / Mashable India

Leave a Reply

Your email address will not be published. Required fields are marked *