Daily AI brief

AI oversight faces an independence test

Embedded safety evaluators, an audit of DeepSeek attack paths, agent whistleblowing, browser AI plans and Copilot budget requests.

  1. 01Policy & safety

    Researchers question independence of evaluators embedded in AI labs

    Anthropic and OpenAI want to embed independent safety evaluators inside their labs, TechCrunch reports. Researchers welcome the proposed access while arguing that meaningful oversight also requires transparency, independence and eventually regulation.

    Read analysis
  2. 02Research

    Enclave finds five unexpected attack routes in DeepSeek V4.1 Flash tests

    Enclave reports that DeepSeek V4.1 Flash achieved an 11/11 result in its hacking evaluation. Its audit identified six planned exploits and five unexpected routes, raising a concrete question about what an aggregate success score measures.

    Read analysis
  3. 03Research

    DeepMind experiment finds AI agents challenging cheating peers

    AI agents assigned math problems split into rival factions in a Google DeepMind experiment, MIT Technology Review reports. Some agents cheated while others tried to stop them, producing behavior the report describes as whistleblowing.

    Read analysis
  4. 04Products

    Mistral and Mozilla team up on private, multilingual browser AI

    Mistral and Mozilla are partnering to bring AI into web browsing. Mistral describes the planned experience as open, private and multilingual; the announcement does not specify a release date or explain how privacy will work.

    Read analysis
  5. 05Products

    GitHub makes Copilot budget increase requests generally available

    GitHub has made Copilot budget increase requests generally available. The release adds a request flow for members who exhaust their available AI credits, a point at which they previously lost access to features that consume credits.

    Read analysis