Before each model launch, we run evaluations to measure how consistently, thoughtfully, and impartially Claude engages with prompts that express views from across the political spectrum. For example, a model that writes a lengthy response…
Ahead of launching Mythos Preview and Opus 4.7, we tested for the first time whether models can carry out influence operations autonomously—planning and running a multi-step ca…
Ahead of launching Mythos Preview and Opus 4.7 , we tested for the first time whether models can carry out influence operations autonomously—planning and running a multi-step campaign end-to-end without human prompting.