Intelligence/Technology/Frontier safety frameworks
Last updated 18 July 2026
Technology · Frontier governance · Entry 01

Frontier safety frameworks read as a version diff

Five companies write the rules that govern the most capable AI systems in existence. Those rules have been rewritten twenty-eight times. This is a line-by-line comparison of what changed, when, and in which direction.

28Published versions
9Anthropic revisions
9Diff dimensions
3Directions on manipulation
18 Jul 26Corpus close date
01 · At a glance

What the frameworks are, and what moved

Anthropic, OpenAI, Google DeepMind, Meta, and xAI each publish a voluntary document that sets capability thresholds, evaluation commitments, and conditions for deployment. These are among the more developed self-governance instruments in any industry. They are also the operative constraint on frontier model release, since binding regulation is only now arriving.

Every one of these documents has been revised. Anthropic has published nine versions of its Responsible Scaling Policy since September 2023, four of them between February and July 2026. Google DeepMind is on version 3.1. Meta rewrote and renamed its framework in April 2026. The revisions carry information that the first drafts do not: what a company changed while operating under competitive and regulatory pressure describes its posture more precisely than what it wrote before shipping anything.

Four movements show up across the set.

  • Pause language became conditional. Anthropic's version 3.0 dropped the general pause commitment and separated what it will do alone from what it recommends the industry adopt. Meta's Critical tier action changed from stop development to develop with mitigations.
  • Competitor-conditional clauses spread. Anthropic, OpenAI, and Google DeepMind each hold a provision permitting relaxed safeguards if another developer proceeds without comparable ones. Meta holds no such clause in either version.
  • Manipulation moved in three directions at once. OpenAI removed persuasion as a tracked category in April 2025. DeepMind added a harmful manipulation threshold in September 2025. Anthropic and Meta have never tracked it.
  • Transparency machinery grew as commitments loosened. Named risk reports, external review rights, and board oversight all expanded in the same versions that softened the hard tripwires.
Why this matters to allocators

Until the EU AI Act's systemic-risk obligations take effect on 2 August 2026 and California's SB 53 disclosure regime matures, these documents are the primary published description of how a frontier developer decides whether a model is safe to release. They are the closest available proxy for operational risk discipline at companies whose valuations assume continued unrestricted deployment.

02 · Method

How the comparison was built

Primary documents were retrieved from company domains and content delivery networks, and cross-checked against the METR Frontier AI Safety Policies index for version numbers and dates. Secondary sources establish that a version exists or attribute named commentary. They are never presented as document text.

Each version is compared against its predecessor on the same nine dimensions, so the cross-company read is possible. Quotations are held under 15 words and attributed to document and version. Claims that could not be confirmed against a primary document were moved to the anomalies section rather than left asserted.

Figure 01
The nine diff dimensions
Applied identically to every framework. Hover any segment for the question it answers.
View data as a table
The nine diff dimensions
DimensionNameQuestion it answers
D1Capability thresholdsWhat level of capability triggers a response, and on what measurement basis
D2Evaluation categoriesWhich risk domains are in scope
D3Commitment modalityThe grammar of obligation: will, will provided that, intend to, aim to
D4Triggering and timingWhen evaluations run relative to training and release
D5Response and mitigationWhat happens once a threshold is crossed, and who judges sufficiency
D6GovernanceNamed decision authority, board involvement, escalation, reporting
D7Conditional withdrawalLanguage releasing the company from commitments if competitors proceed
D8SecurityWeights protection, insider threat, mapping to an external standard
D9Definitional scopeWhat counts as a covered model: compute floors, open weights, fine-tunes
The dimensions separate what a document promises from when it promises it, who decides, and under what conditions the promise is released.
03 · The corpus

Twenty-eight versions, three years

The set below covers every published version located as of 18 July 2026. Retrieval status records whether the document itself was read, whether an archive capture stood in, or whether existence rests on an index entry.

Figure 02
Publication timeline against regulatory milestones
Each mark is a published version. Large marks are major releases, small marks are point revisions. Dashed lines mark external events. Correlation is shown; no causal claim is made.
Anthropic
OpenAI
Google DeepMind
Meta
xAI
View data as a table
Publication timeline
CompanyVersionDateNote
Anthropicv1.019 Sep 2023First framework of its kind
Anthropicv2.015 Oct 2024Restructured around capability thresholds
Anthropicv2.131 Mar 2025State-programme CBRN threshold added
Anthropicv2.214 May 2025Insider threat exclusion widened
Anthropicv3.024 Feb 2026Rewrite. Pause commitment dropped
Anthropicv3.12 Apr 2026Pausing retained as discretionary option
Anthropicv3.229 Apr 2026Trust gains external review powers
Anthropicv3.326 May 2026CB weapons threshold revised
Anthropicv3.48 Jul 2026Risk report sharing set to 200 staff
OpenAIBeta18 Dec 2023Four tracked categories
OpenAIv215 Apr 2025Persuasion removed. Marginal-risk clause added
OpenAIFGF28 May 2026Companion regulatory mapping document
Google DeepMindv1.017 May 2024Critical Capability Levels introduced
Google DeepMindv2.04 Feb 2025CCLs mapped to security levels
Google DeepMindv3.022 Sep 2025Harmful manipulation CCL added
Google DeepMindv3.117 Apr 2026Tracked Capability Levels added
Metav1.03 Feb 2025Uniquely enable standard. Stop development tier
Metav2.08 Apr 2026Renamed. Substantially contribute to standard
xAIDraft20 Feb 2025Watermarked draft at the Seoul deadline
xAIv1.020 Aug 2025MASK dishonesty criterion
xAIv2.030 Dec 2025Renamed, two days before SB 53 in force
xAIrev30 Jun 2026Quantitative criteria removed
RegulatorySeoul commitmentsMay 202416 companies commit at the AI Seoul Summit
RegulatoryParis summitFeb 2025AI Action Summit
RegulatorySB 53 signed29 Sep 2025California Transparency in Frontier AI Act
RegulatorySB 53 in force1 Jan 2026Obligations commence
RegulatoryUS executive order2 Jun 2026Voluntary pre-release government access framework
The February 2025 cluster precedes the Paris AI Action Summit. The 2026 cluster tracks California SB 53 coming into force and the approach of EU systemic-risk obligations.

Corpus table

CompanyFrameworkVersionDateSourceStatus
AnthropicResponsible Scaling Policyv1.019 Sep 2023www-cdn.anthropic.comPrimary
AnthropicRSPv2.015 Oct 2024www-cdn.anthropic.comPrimary
AnthropicRSPv2.131 Mar 2025www-cdn.anthropic.comPrimary
AnthropicRSPv2.214 May 2025www-cdn.anthropic.comPrimary
AnthropicRSPv3.024 Feb 2026anthropic.comPrimary
AnthropicRSPv3.12 Apr 2026www-cdn.anthropic.comPrimary
AnthropicRSPv3.229 Apr 2026cdn.sanity.ioPrimary
AnthropicRSPv3.326 May 2026cdn.sanity.ioPrimary
AnthropicRSPv3.48 Jul 2026cdn.sanity.ioCurrent
OpenAIPreparedness FrameworkBeta18 Dec 2023cdn.openai.comPrimary
OpenAIPreparedness Frameworkv215 Apr 2025cdn.openai.comCurrent
OpenAIFrontier Governance Framework28 May 2026n/acdn.openai.comCompanion
Google DeepMindFrontier Safety Frameworkv1.017 May 2024storage.googleapis.comPrimary
Google DeepMindFSFv2.04 Feb 2025storage.googleapis.comPrimary
Google DeepMindFSFv3.022 Sep 2025storage.googleapis.comPrimary
Google DeepMindFSFv3.117 Apr 2026storage.googleapis.comCurrent
MetaFrontier AI Frameworkv1.03 Feb 2025web.archive.orgArchive
MetaAdvanced AI Scaling Frameworkv2.08 Apr 2026ai.meta.comCurrent
xAIRisk Management FrameworkDraft20 Feb 2025data.x.aiPrimary
xAIRisk Management Frameworkv1.020 Aug 2025data.x.aiPrimary
xAIFrontier AI Frameworkv2.030 Dec 2025data.x.aiPrimary
xAIFrontier AI Frameworkrev.30 Jun 2026media.x.aiCurrent
MicrosoftFrontier Governance Frameworkv1.0Feb 2025microsoft.comPrimary
MicrosoftFrontier Governance FrameworkupdateFeb 2026microsoft.comDiff open
AmazonFrontier Model Safety Frameworkv1.09 Feb 2025amazon.sciencePrimary
NvidiaFrontier AI Risk Assessment17 Feb 2025n/aimages.nvidia.comIndex
MagicAGI Readiness Policyv1.02 Jul 2024magic.devIndex
NAVERAI Safety Framework7 Aug 2024n/aclova.aiIndex
G42Frontier AI Safety Framework6 Feb 2025n/ag42.aiIndex
CohereSecure AI Frontier Model Framework7 Feb 2025n/acohere.comIndex

A page-count column was omitted. Counts could be confirmed for only 6 of the 30 entries, and a column that is mostly empty asserts less than it implies. The unresolved counts are recorded in the anomalies section. Amazon's document states 9 February 2025; the METR index labels it 10 February 2025.

04 · Anthropic

Responsible Scaling Policy

The RSP was the first document of its kind, published September 2023. It sets AI Safety Levels, assessment procedures, and required safeguards, and governs decisions on training and deployment. Anthropic maintains a public changelog and publishes redline comparisons for point releases, a practice no other developer in the set matches.

Version history

Nine versions: v1.0 (19 September 2023), v2.0 (effective 15 October 2024), v2.1 (31 March 2025), v2.2 (14 May 2025), v3.0 (24 February 2026), v3.1 (2 April 2026), v3.2 (29 April 2026), v3.3 (26 May 2026), v3.4 (8 July 2026). Four revisions landed in roughly four months in 2026.

Diff by dimension

D3 · Commitment modality

Version 3.0 is the substantive shift in the entire corpus. Anthropic separates what it commits to unilaterally from what it recommends the industry adopt, and names the reason.

driven by a collective action problem RSP v3.0 announcement, February 2026

The most demanding measures now sit in a recommendations tier that Anthropic will strive to advance but cannot commit to following ... unilaterally. Frontier Safety Roadmap goals are described as nonbinding but publicly-declared targets.

Figure 03
Where the commitments live, v2.x against v3.x
Version 3.0 introduced a second tier. Measures that were previously described as company commitments now sit in a tier conditioned on industry-wide adoption.
View data as a table
Anthropic commitment architecture
VersionTierContents
v2.x, to May 2025Unilateral commitments (single tier)Required safeguards, with pause if the standard cannot be met
v3.x, from Feb 2026Industry-wide recommendationsRAND SL4 and the strongest measures. Nonbinding, publicly declared
v3.x, from Feb 2026Unilateral commitmentsASL-3 security, risk reports, trust oversight
Anthropic's stated rationale is that unilateral adoption of the most demanding measures carries cost without proportionate risk reduction if competitors do not follow.

D5 · Response and mitigation commitments

The general pause commitment was dropped in v3.0. Version 3.1 then clarified that the option remains available without the obligation.

Pause commitment D5 · Mitigation
v2.x · through May 2025
Each Capability Threshold pairs with a required standard and an implied commitment to pause development or deployment as needed if that standard cannot be met.
v3.1 · April 2026
Anthropic remain[s] free to take measures such as pausing at its discretion, with the decision routed through a Risk Report and a sufficiency judgment.
Effect. The pause moves from a pre-committed tripwire to a retained discretionary option. GovAI records Karnofsky describing this as the biggest change in version 3.

D4 · Triggering and timing

Version 3.0 extended the routine evaluation interval from three months to six. Anthropic acknowledges in the same document that its most recent evaluations were completed 3 days later than the three-month interval, and that it identified instances where it fell short of meeting the full letter of its requirements.

D8 · Security commitments

Anthropic continues to commit unilaterally to the ASL-3 security standard. The more demanding RAND Security Level 4, which addresses state-level weight theft, moved into the industry-wide recommendations tier. The v3.0 announcement describes RAND's SL5 as currently not possible.

Insider threat scope, ASL-3 security standard D8 · Security
v2.1 · March 2025
Excludes highly sophisticated state-compromised insiders from the scope of the standard.
v2.2 · May 2025
Excludes sophisticated insiders and state-compromised insiders from the scope of the standard.
Effect. The excluded adversary class widens from one narrow category to two broader ones. Verified against the 14 May 2025 changelog entry; exact footnote text in each PDF remains a gap.

D6 · Governance and accountability

Governance expanded while substantive commitments loosened. Version 3.2 authorises the Long-Term Benefit Trust to request external review of Risk Reports, approve external reviewers, and receive regular briefings. Version 3.4 changed the internal disclosure requirement.

Internal sharing of unredacted risk reports D6 · Governance
v3.0 to v3.3
Requires sharing with all regular-clearance Anthropic staff.
v3.4 · 8 July 2026
Requires sharing with at least 200 Anthropic employees, and public Risk Reports must indicate where material was redacted.
Effect. Internal readership becomes a floor rather than a whole-population requirement, alongside a new public redaction marker. The two changes run in opposite directions on disclosure.

D1, D2, D7, D9 · In brief

  • D1 thresholds. Version 3.0 reframes the AI research and development threshold around compressing two years of 2018 - 2024 AI progress into a single year. Version 3.1 clarified this means rate of progress, not researcher productivity. Versions 3.3 and 3.4 each revised a threshold to better track the threat model of concern.
  • D2 categories. Centred on CBRN and autonomous AI research and development. Cyber operations sits under ongoing assessment in v2.0 rather than as a committed threshold. No persuasion or manipulation category has appeared in any version.
  • D7 conditional withdrawal. Version 3 relocates competitor provisions into an appendix titled Commitments Related to Competitors.
  • D9 scope. Risk Reports cover all publicly deployed models and, where risk warrants, internal models.

Reading

The RSP moved from a document of pre-committed unilateral tripwires to one that distinguishes unilateral action from industry recommendation, paired with new transparency instruments and expanded trust oversight. Anthropic names the collective action problem as the driver and acknowledges past shortfalls against the letter of prior requirements. Whether the trade is favourable depends on whether the reporting and external review provisions bind as firmly in practice as the pause commitment they partly replaced.

05 · OpenAI

Preparedness Framework

Published in beta on 18 December 2023 and revised once, on 15 April 2025. A separate Frontier Governance Framework followed on 28 May 2026, which OpenAI states does not replace the Preparedness Framework. Two full versions makes this the least-revised framework among the four primary subjects.

Diff by dimension

D1, D2 · Thresholds and categories

The beta used four tracked categories on a four-level scale, gating deployment at medium or below post-mitigation and halting development at critical. Version 2 collapsed the scale to two operative thresholds and reduced the tracked set to three categories.

Figure 04
Tracked categories, beta to version 2
Persuasion left the framework. Nuclear and radiological moved to a research tier. Model autonomy was renamed and narrowed to AI self-improvement.
View data as a table
OpenAI tracked categories, beta to version 2
Beta category (Dec 2023)Disposition in version 2 (Apr 2025)
CybersecurityTracked. Carried forward unchanged
CBRNSplit. Biological and chemical remain tracked; nuclear and radiological move to a research tier
Model autonomyRenamed and narrowed to AI self-improvement
PersuasionRemoved. Handled outside the Preparedness Framework
OpenAI states that persuasion risks do not fit the criteria for inclusion and are handled outside the Preparedness Framework.

Version 2 defines High capability as capability that could amplify existing pathways to severe harm and Critical as capability that could introduce unprecedented new pathways to severe harm. The biological and chemical High threshold is tied to meaningful counterfactual assistance to novice actors.

Persuasion as a tracked category D2 · Categories
Beta · December 2023
Tracked categories: cybersecurity, CBRN, model autonomy, persuasion.
v2 · April 2025
Tracked categories: biological and chemical, cybersecurity, AI self-improvement. Persuasion is handled outside the Preparedness Framework.
Effect. The category with the clearest link to electoral and market integrity leaves the pre-commitment structure. OpenAI's stated rationale is the severity criteria for inclusion. The Frontier Governance Framework of May 2026 reintroduces harmful manipulation as a covered area for regulatory purposes.

D7 · Conditional withdrawal

Absent in the beta. Version 2 introduced a marginal-risk clause under which OpenAI may adjust our requirements if another developer releases a high-risk system without comparable safeguards, provided it confirms the change does not net increase severe-harm risk, acknowledges the adjustment publicly, and keeps safeguards more protective.

D3, D4, D5 · Modality, triggering, mitigation

Version 2 applies a five-part test to what qualifies as a tracked risk: risks must be plausible, measurable, severe, net new and instantaneous or irremediable. It formalises Capabilities Reports and Safeguards Reports reviewed by the Safety Advisory Group, and states that reducing risk generally does not require reducing capability. The beta's commitment to test fine-tuned versions of models was removed.

Named commentary, not document text

Shyam Krishna of RAND Europe told Fortune on 16 April 2025 that OpenAI appears to be shifting its approach. Former OpenAI safety researcher Steven Adler wrote on X on 15 April 2025 that OpenAI is quietly reducing its safety commitments, flagging the removal of the fine-tuned model testing requirement.

D6 · Governance

The beta gave the board the right to reverse leadership safety decisions, and version 2 preserves it: the Board may reverse a decision and mandate a revised course of action. The Safety Advisory Group recommends to leadership, with the chief executive or a designee holding final authority.

Reading

Version 2 narrowed the tracked set from four categories to three, simplified four risk levels to two operative thresholds, and added an explicit conditional relaxation clause. OpenAI frames these as sharper focus and better measurability. The May 2026 Frontier Governance Framework is a compliance mapping layer to California SB 53 and the EU code of practice rather than a tightening of the underlying thresholds.

06 · Google DeepMind

Frontier Safety Framework

Four versions since May 2024. The framework is built on Critical Capability Levels, thresholds at which a model could cause severe harm without mitigation, matched to recommended security levels and deployment mitigations. It is the most process-oriented document in the set, specifying what will be measured and reviewed more than what action follows.

Version history

v1.0 (17 May 2024), v2.0 (4 February 2025), v3.0 (22 September 2025), v3.1 (17 April 2026). The current document states Version 3.1 (April 17, 2026).

Figure 05
Scope added by version
DeepMind is the only primary subject whose covered surface grew at every revision. Each block enters at the version where it was introduced.
View data as a table
DeepMind scope added by version
VersionDateAdded in that versionCumulative elements
v1.0May 2024Critical Capability Levels1
v2.0Feb 2025Security level mapping, deceptive alignment3
v3.0Sep 2025Harmful manipulation CCL, internal deployment safety cases5
v3.1Apr 2026Tracked Capability Levels6
Correction to the prior draft: Tracked Capability Levels entered in version 3.1, not version 3.0. Several third-party trackers collapse the two revisions into one entry.

D2 · Evaluation categories

Version 3.0 added a Critical Capability Level for harmful manipulation, covering models with capabilities that could be misused to systematically and substantially change beliefs and behaviors in identified high-stakes contexts. The 22 September 2025 announcement was authored by John Flynn, Helen King, and Anca Dragan.

Early warning tier D1 · Thresholds
v2.0 to v3.0
A single tier of Critical Capability Levels, with an alert threshold and safety buffer ahead of each.
v3.1 · 17 April 2026
Adds Tracked Capability Levels in certain domains, a lower tier covering less extreme risks that triggers monitoring before a Critical level is reached.
Effect. A two-tier structure replaces a single tier. This is a widening of coverage, running against the direction of travel elsewhere in the set.

D4, D6 · Triggering and governance

Evaluations run at a regular cadence and after a significant capability jump. Version 3.0 extended safety case review to large-scale internal deployments, stating that for advanced machine learning research and development levels, large-scale internal deployments can also pose risk. The framework commits to sharing model information, evaluation results, and mitigation plans with governments where a model reaches a level posing material risk. Review runs through internal bodies including the AGI Safety Council. As an Alphabet division, DeepMind has no independent board structure of its own.

D7 · Conditional withdrawal

Version 3.1 states that the social value of certain mitigations is significantly reduced if not broadly applied, and that recommended security levels may be adjusted if our understanding of the risks changes, for example where a model does not possess capabilities meaningfully different from other publicly available models. DeepMind judges when the condition is met.

D3 · Commitment modality

Manipulation and misalignment research is described as nascent, and the associated level and risk assessment as exploratory and subject to further research. These qualifiers carry across versions and are the framework's principal source of flexibility.

Reading

DeepMind's framework expanded across 2025 and 2026, adding a manipulation threshold, an early warning tier, and internal deployment safety cases. It buys flexibility through process language and exploratory qualifiers rather than through weakened commitments, and it couples this with an explicit statement that the value of some mitigations depends on industry-wide adoption.

07 · Meta

Frontier AI Framework, now Advanced AI Scaling Framework

Meta published the Frontier AI Framework on 3 February 2025, ahead of the Paris AI Action Summit, and replaced it in April 2026 with a renamed second iteration. The framework operates in the context of open-weight release, which makes its adversary modelling and its trigger language the load-bearing choices.

Version history

v1.0 (3 February 2025) and v2.0 (8 April 2026 per the announcement, 7 April 2026 per the document change log). The framework describes itself as the second iteration, previously titled the Frontier AI Framework.

D1, D3 · Threshold language and modality

The primary standard changed, and so did the consequence attached to the highest tier. These two changes run in opposite directions.

Trigger standard and Critical tier action D1 · D3 · D5
v1.0 · February 2025
Applies where a model would uniquely enable a threat scenario. Critical tier action: stop development.
v2.0 · April 2026
Applies where a model would substantially contribute to a threat scenario. Critical tier action: develop with mitigations, permitted only where mitigations reduce risk to that of a moderate or lower model.
Effect. The trigger catches more models while the consequence at the top tier becomes a mitigation pathway rather than a halt. Meta describes version 2 as a significantly updated and more rigorous version.
Figure 06
Meta risk tiers and the action attached to each
Every tier in version 2 routes through mitigation to moderate or lower. The halt condition present in version 1 is not carried forward.
View data as a table
Meta risk tiers and actions
Risk tierv1.0 action (Feb 2025)v2.0 action (Apr 2026)
ModerateReleaseRelease
HighDo not releaseDevelop with mitigations
CriticalStop developmentDevelop with mitigations
Version 2 also adds published Safety and Preparedness Reports covering risk assessment, evaluation results, deployment rationale, and limitations.

D2 · Evaluation categories

Version 1 covered cybersecurity and chemical and biological risk. Version 2 adds Loss of Control as a third domain, decomposed into autonomous AI research and development, evaluation awareness (whether a model can identify when it is being evaluated and selectively adapt its outputs), monitor interference, and autonomous replication as an emerging concern. No persuasion or manipulation domain appears in either version.

D6 · Governance

Version 1 located the release decision with an unnamed leadership team. Version 2 names the Chief AI Officer and the Director of Alignment and Risk as the decision-makers who assign risk thresholds and approve deployment, and adds a published Model Spec listing intended propensities including acquiescence to shutdown and lack of coercive power-seeking behavior.

D7 · Conditional withdrawal

Neither version contains a competitor-conditional relaxation clause. This is a confirmed absence against the retrieved version 2 text and the archived version 1, and it distinguishes Meta from the other three primary subjects.

D9 · Definitional scope

Version 2 expands the deployment taxonomy to internal deployment, limited deployment, controlled deployment, closed release, or open release, and states that Loss of Control risks may occur with similar probability with any type of deployment, including internal deployment. The four-part inclusion test is plausible, catastrophic, net new, and instantaneous or irremediable.

Reading

Version 2 broadens scope and lowers the trigger while converting the top-tier consequence from a halt to a mitigation pathway, and it names individual officers as accountable decision-makers. The governance additions are specific and auditable. Meta remains the only primary subject with no competitor-conditional clause.

08 · Other published frameworks

xAI, Microsoft, Amazon, and the rest of the signatory set

xAI

xAI published a draft Risk Management Framework watermarked DRAFT on 20 February 2025, the Seoul commitment deadline, promising an updated version within three months. Version 1.0 followed on 20 August 2025, roughly six months past that deadline. It states a quantitative deployment criterion: maintaining a dishonesty rate of less than 1 out of 2 on MASK.

xAI released Grok Code Fast 1 on 28 August 2025, eight days later. The model card reported a 71.9 percent dishonesty rate on the same benchmark. AI Lab Watch's Zach Stein-Perlman described the framework as dreadful and profoundly unserious. This is named external commentary, cited to establish the criticism rather than as document text.

A renamed Frontier Artificial Intelligence Framework version 2.0 followed on 30 December 2025, two days before California's Transparency in Frontier Artificial Intelligence Act came into force. A further revision dated 30 June 2026 is a distinct document at a separate address, resolving the question of whether these are two revisions or a relabelling. The Midas Project reports that the June 2026 revision removed both quantitative risk acceptance criteria and whistleblower protection language; this is secondary reporting.

Microsoft

Microsoft published its Frontier Governance Framework in February 2025. It sets qualitative capability thresholds, scales security safeguards to pre-mitigation risk levels across four bands, and integrates with Microsoft's wider AI governance programme. A February 2026 update exists as a distinct primary document. What changed between the two could not be established from a primary comparison in this pass and is carried as an open item.

Amazon

Amazon published its Frontier Model Safety Framework on 9 February 2025, tied to the Paris summit and its endorsement of the Korea Frontier AI Safety Commitments. It defines Critical Capability Thresholds across CBRN, offensive cyber operations, and automated AI research and development, and commits not to deploy models exceeding thresholds without safeguards. The document commits to review at least annually. No second version appears as of 18 July 2026. Amazon has instead published model-specific evaluations under the framework, including for Nova Premier and Nova 2.0 Lite.

The remainder of the signatory set

Cohere, G42, NAVER, Nvidia, and Magic remain at their first published versions with no second version indexed. Nvidia's and Cohere's documents emphasise domain-specific rather than catastrophic risk. No published second version was found for Zhipu or Samsung, neither of which appears on METR's published-policy list.

METR's December 2025 update records that sixteen companies agreed to the Seoul commitments with an additional four companies joining since then, and that twelve companies have published frontier AI safety policies. As of 18 July 2026 the public index still reflects that count.

09 · Cross-framework read

Where the documents converge and where they part

Convergent movement

All five major developers now use a capability threshold structure paired with pre-committed responses. Loss of control and misalignment has become a shared domain: DeepMind through misalignment levels, Meta through Loss of Control as a third domain, OpenAI through the Frontier Governance Framework, and Anthropic through an affirmative misalignment case at its research and development thresholds. Named risk-reporting artefacts now exist at all four primary subjects.

Competitor-conditional clauses are the second convergence. Three of the four primary subjects hold one. Anthropic elevated the concept from an emergency provision to the organising logic of version 3.

Figure 07
Competitor-conditional relaxation clauses
A provision permitting relaxed safeguards where another developer proceeds without comparable ones. In every case the company itself judges whether the condition is met.
View data as a table
Competitor-conditional relaxation clauses
CompanyClauseForm
AnthropicPresentOrganising logic of version 3. Competitor provisions in a dedicated appendix
OpenAIPresentMarginal-risk clause in version 2, subject to public acknowledgement
Google DeepMindPresentBroad-application clause. Recommended security levels may be adjusted
MetaAbsentNo clause in either version
xAIAbsentRelies on quantitative benchmark criteria instead
No framework in the set attaches an external verification step or an independent adjudicator to the invocation of these clauses.

Divergent movement

The clearest divergence is manipulation. Within roughly fourteen months the four primary frameworks moved in three different directions on the same capability.

Figure 08
Persuasion and manipulation, tracked or not
One removal, one addition, one reintroduction through a companion document, and two standing absences.
View data as a table
Persuasion and manipulation across frameworks
CompanyPeriodStatus
OpenAIDec 2023 to Apr 2025Persuasion tracked as one of four categories
OpenAIApr 2025 to May 2026Not tracked
OpenAIFrom May 2026Harmful manipulation reintroduced in the companion governance document
Google DeepMindFrom Sep 2025Harmful manipulation Critical Capability Level
AnthropicAll versionsNever tracked
MetaAll versionsNever tracked
The category with the most direct bearing on information integrity is the one on which the frameworks agree least.

Modality diverges as well. Anthropic separated unilateral commitments from industry recommendations and dropped its pause commitment. OpenAI added a conditional relaxation clause. Meta lowered its trigger but holds no competitor clause. DeepMind expanded scope at every revision. xAI relies on quantitative benchmark thresholds that no other developer uses.

Figure 09
Direction of travel by dimension
Each cell records whether the most recent revision widened coverage, narrowed it, or left it materially unchanged relative to the prior version.
â–² Widened or tightened
● Restructured, mixed effect
â–¼ Narrowed or softened
· No material change
View data as a table
Direction of travel by company and dimension
CompanyD1D2D3D4D5D6D7D8D9
Anthropic (v3.3 to v3.4)MixedNo changeSoftenedSoftenedSoftenedWidenedSoftenedSoftenedWidened
OpenAI (Beta to v2)MixedNarrowedMixedWidenedMixedNo changeSoftenedMixedNarrowed
Google DeepMind (v3.0 to v3.1)WidenedWidenedNo changeWidenedWidenedWidenedSoftenedMixedNo change
Meta (v1.0 to v2.0)MixedWidenedSoftenedWidenedSoftenedWidenedNo changeWidenedWidened
Read across rows for a company's posture, down columns for where the industry is moving together.

Terminology drift

The four now describe similar constructs in language that no longer maps cleanly. Cross-company comparison on any single capability requires translation, which is itself a finding.

ConstructAnthropicOpenAIGoogle DeepMindMeta
Threshold unit Capability Threshold, ASL standard High, Critical capability Critical Capability Level Moderate, High, Critical risk
Early warning Checkpoint Not separately named Tracked Capability Level Capability checkpoint
Deployment test Required Safeguards sufficiency Sufficiently minimise Safety case at level Substantially contribute to
Risk artefact Risk Report Capabilities and Safeguards Reports FSF report Safety and Preparedness Report
Levels in scale ASL-2, ASL-3, ASL-4 2 operative 2 tiers from v3.1 3 tiers

Revision timing against external events

Revision dates cluster around summit and regulatory milestones. The pattern is correlation and is stated as such; companies give multiple reasons for revisions, including learning from implementation.

  • February 2025. Meta v1.0, DeepMind v2.0, Microsoft v1.0, Amazon, G42, Cohere, Nvidia, and the xAI draft all land in the weeks around the Paris AI Action Summit.
  • 30 December 2025. xAI publishes version 2.0 two days before California's Transparency in Frontier Artificial Intelligence Act takes effect.
  • February 2026. The India AI Impact Summit produces the New Delhi Frontier AI Impact Commitments, signed by thirteen developers, and a declaration endorsed by 92 countries and international organisations. These address usage insights and multilingual evaluation rather than catastrophic-risk thresholds.
  • May and June 2026. OpenAI's Frontier Governance Framework and Microsoft's update map to SB 53 and the EU code of practice. A US executive order of 2 June 2026 creates a voluntary pre-release government access framework for covered frontier models, with an access window of up to 30 days, and expressly prohibits mandatory licensing or preclearance.
  • 2 August 2026. EU AI Act systemic-risk obligations for general purpose AI take effect, ahead of which the 2026 revision wave sits.
10 · Anomalies and open questions

What did not resolve cleanly

Items below did not confirm against a primary document in this pass, or produced conflicting evidence. They are recorded rather than resolved.

#ItemStatus
01Anthropic v1.0 autonomy trigger. The quantitative 50 percent aggregate success rate phrasing attributed to v1.0 was not reverified word for word. The existence of an autonomy-based ASL-3 trigger is not disputed.Unconfirmed
02Anthropic v2 competitor clause section number. Cited elsewhere as section 7.1.7. Section number and exact wording not reverified against the PDF. Substance corroborated by GovAI; v3.1 confirms a relocated appendix on competitor commitments.Unconfirmed
03Anthropic ASL-4 definition. v2.0 retained a forward commitment to define further thresholds mandating ASL-4 safeguards. Whether v3.x defines any, and whether the original 2023 commitment was removed without changelog acknowledgement, needs a word-for-word comparison not completed here.Open
04Anthropic insider footnote text. Substance verified via changelog. Exact footnote text and numbering in the v2.1 and v2.2 PDFs not extracted.Partial
05Karnofsky quotation on regulation. Verified through GovAI's rendering rather than the original post.Secondary
06Meta compute threshold. The figure of at least 1026 operations and the open-weight adversary phrasing were not located in the retrieved sections of version 2.Unconfirmed
07Meta version 1.1. Listed in the ETO AGORA catalogue. No Meta-published v1.1 was found. Appears to be a third-party cataloguing artefact.Resolved as artefact
08Meta v2.0 date. Change log 7 April, blog 8 April, a social post 20 April. Reading: 7 April is the document effective date, 8 April the public announcement, 20 April a re-post.Resolved
09Microsoft February 2026 update. Exists as a primary document. Diff against version 1.0 not established.Open
10OpenAI Frontier Governance Framework specifics. The ISO 27001 and SOC 2 Type II references and the Ireland entity oversight assignment rest on trade press in this pass.Secondary
11Page and word counts. Confirmed for 6 of 30 entries only. The corpus table omits the column rather than carrying mostly empty cells.Partial
12Amazon publication date. The document states 9 February 2025. The METR index labels it 10 February 2025. Both preserved.Conflict noted

Corpus close

No framework activity was found later than 8 July 2026. Anthropic's policy page states last updated 8 July 2026. No OpenAI Preparedness Framework version 3 exists. DeepMind's current version remains 3.1. Meta has published no revision after version 2.0.

11 · Sources

Documents analysed

All access dates 18 July 2026. Primary documents were read directly except where noted in the corpus table.

Anthropic

Responsible Scaling Policy pages and update log; RSP PDFs v1.0 through v3.4 on www-cdn.anthropic.com and cdn.sanity.io; the version 3 announcement; Frontier Compliance Framework. GovAI analysis of RSP v3.0 cited as secondary.

OpenAI

Preparedness Framework beta and version 2 PDFs on cdn.openai.com; the version 2 announcement; Frontier Governance Framework PDF and announcement. Fortune, 16 April 2025, and Steven Adler on X, 15 April 2025, cited as named commentary.

Google DeepMind

Frontier Safety Framework PDFs v2.0, v3.0, and v3.1 on storage.googleapis.com; the strengthening and updating announcements on deepmind.google; Gemini model cards.

Meta

Advanced AI Scaling Framework version 2 on ai.meta.com and the accompanying blog; version 1.0 through a web.archive.org capture; Frontier Model Forum summary; ETO AGORA catalogue entry.

xAI, Microsoft, Amazon

xAI framework PDFs on data.x.ai and media.x.ai and the Grok Code Fast 1 release note; Microsoft Frontier Governance Framework PDFs for 2025 and February 2026; Amazon Frontier Model Safety Framework and Nova evaluation reports on amazon.science. AI Lab Watch, SaferAI, and the Midas Project cited as named commentary.

Cross-cutting

METR Frontier AI Safety Policies index and Common Elements report; whitehouse.gov presidential action of 2 June 2026; law firm client alerts on that order; Brookings and Lawfare on SB 53; Press Information Bureau and Carnegie Endowment on the India AI Impact Summit.

Cite this entry

Shrestha, S. (2026) "Frontier safety frameworks read as a version diff."
Intelligence, Technology desk, entry 01. Published 18 July 2026.
https://intelligence.sushantshrestha.com/technology/frontier-safety/01-version-diff.html

Next

02 · Frontier safety frameworks, read against what the models actually did. In preparation.