RADAR ·
Departing Anthropic researcher criticizes the race to self-improving AI
Jacob Coxon, a researcher who trained AI systems at Anthropic, announced on X that he had left the company, citing its lax approach to safety. Coxon, who previously trained systems for OpenAI, accused both companies of racing toward self-improving superintelligence.
Hours later, Evan Hubinger, who leads one of Anthropic's safety teams, said self-improving AI was arriving faster than expected and put his personal estimate that AI could kill all humans within the next decade at greater than one in ten. Hubinger added that the company does not yet have a plan for keeping advanced systems safe and aligned, and is not clearly on track to develop one. The exchange comes as the companies prepare for anticipated public offerings.
“We really do earnestly believe AI could kill all humans”Evan Hubinger, Anthropic
Source: The Verge · AI