Researchers question independence of evaluators embedded in AI labs

Anthropic and OpenAI want to embed independent safety evaluators inside their labs, TechCrunch reports. Researchers welcome the proposed access while arguing that meaningful oversight also requires transparency, independence and eventually regulation.

Key points

  1. Both Anthropic and OpenAI are seeking to embed safety evaluators inside their AI labs.
  2. Researchers welcome the access but warn that meaningful oversight needs transparency and independence.
  3. The researchers' concerns extend to the eventual need for regulation.

Why it matters

Organizations relying on external safety assessments need to examine how evaluators operate, alongside what access they receive. An embedded arrangement leaves open whether reviewers can challenge the lab and disclose their findings.

What to watch

Look for published terms covering evaluator selection, access, reporting rights and whether labs can restrict disclosure of unfavorable findings.

Sources