Summary:
- The article examines the emergence of independent AI research organizations, such as METR (Model Evaluation and Threat Research), which are dedicated to rigorously testing frontier AI models for catastrophic risks.
- It highlights the shift toward third-party safety evaluations, moving away from self-policing by AI labs and toward standardized, empirical methodologies to assess capabilities like autonomous weaponization and cyberattack efficacy.