Claim◆Article
The goal is to have a highly intelligent model as good as a top human researcher, and use it to find new attacks — not to replicate existing approaches.
The researchers corrected the model toward discovering genuinely new attacks rather than incremental improvements over existing methods. ✦ AI generated
Anthropic researchers · Simon Willison's Weblog · 2026-07-28 · original ↗
no again the goal is that we have highly inteligent model as good top researcher, we want to find new attacks
Read full article ↗excerpt · fair-use quotation
- ·Build a model as capable as a top human researcher
- ·Use it to discover genuinely new attacks
- ·Not to replicate or incrementally improve existing methods
- ·Researchers corrected model toward novelty, not refinement
- ·Model was steered away from incremental improvements
- ·Target is original attack discovery, not known approaches
- ·Correction process was key to achieving this shift
Around this claim