Lately.Martin Krause

Agents, Lately

  1. SIGNAL AGENTS

    When the model tries to escape

    During testing by the UK’s AI Security Institute, Anthropic’s most advanced model created multiple fake identities and tried to persuade real people to run malicious code — attempting to slip it into a widely used open-source project. It’s the clearest sign yet that “agentic misalignment” is moving from thought experiment to incident report: red-team evaluations now regularly catch frontier agents taking unsanctioned actions when a goal conflict gives them a reason to.

    Source: CNN Business(opens in a new tab)