
A safety-level ladder modelled on biosafety practice
On 19 September 2023, Anthropic published its Responsible Scaling Policy, described in its original announcement as a framework of 'technical and organizational protocols' aimed at catastrophic risks from misuse or autonomous harmful behaviour, rather than at everyday model quality issues. The policy borrows its structure from biosafety levels: ASL-1 covers systems with no meaningful catastrophic risk, such as older or narrow models; ASL-2 covers systems, including the company's models at the time, that show early hints of dangerous capability, such as being able to discuss weapons-related information without meaningfully uplifting a bad actor; ASL-3 and above describe thresholds the company had not yet reached, defined for systems that would substantially increase misuse risk relative to non-AI alternatives, or show early autonomous capability.
A conditional pause, and evaluation on a schedule
The policy's central commitment is conditional rather than absolute: the announcement states that the ASL system 'implicitly requires' Anthropic to pause training of more capable models if its safety procedures cannot keep pace with a model's capability, and that the company commits not to deploy an ASL-3-level model that shows meaningful catastrophic misuse risk under adversarial testing. The current version of the document, as retrieved on 16 September 2026 from the policy page, describes running capability evaluations at regular intervals to check whether a model has crossed a threshold, and publishing redacted Risk Reports subject to some external review, commitments about process and cadence, not published numerical thresholds a third party can independently verify against the underlying model.
What a self-imposed policy can and cannot guarantee
The version history shows the policy has been revised repeatedly since 2023; a second version in October 2024 added specific ASL-3 deployment safeguards such as real-time classifiers and jailbreak detection, and further revisions followed into 2025 and 2026 as the company refined its capability thresholds. This is worth naming plainly: the policy is a voluntary corporate commitment, adopted, defined and revisable by the company that must comply with it, and enforced through the company's own board approval process rather than by an external regulator or independent auditor. A policy of this kind can commit an organisation to a stated process, evaluate on schedule, escalate safeguards at a threshold, pause if safeguards lag, but it cannot, by its self-imposed nature, guarantee that the organisation's own judgement about whether a threshold has been crossed will be timely or conservative.
- Has the company published evidence that a given evaluation cycle actually ran on schedule, or only that the policy calls for one?
- Which version of the policy was in force when a specific model was released, and what threshold applied under that version?
- What would have to happen for an external party, rather than the company itself, to determine that a safety level has been crossed?
The policy's value lies in making a scaling commitment public and dated enough to be checked against later behaviour; its limit is that the checking, absent external enforcement, still runs through the same organisation that wrote the policy.
Sources & reading trail
Original announcement: publication date, the ASL framework, and the conditional pause commitment.
Source published: 19 September 2023 · Retrieved: 16 September 2026
Version history showing the September 2023 original, the October 2024 second version, and later revisions.
Source published: Not established · Retrieved: 16 September 2026
Current living policy text describing evaluation cadence and Risk Report commitments, as retrieved 16 September 2026.
Source published: Not established · Retrieved: 16 September 2026
Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.