Claude Mythos Preview System Card (2)

PDF · 8 Sept 2026

The Claude Mythos Preview System Card is a comprehensive technical and safety assessment report produced by Anthropic for the Claude Mythos Preview large language model. The document outlines the model’s advanced cybersecurity capabilities, evaluates its vulnerability to indirect prompt‑injection attacks in both computer‑use and browser environments, and compares these results to other Anthropic models (Claude Sonnet 4.6 and Claude Opus 4.6). It presents detailed attack‑success‑rate tables for single‑attempt and adaptive attackers, both with and without safeguards, highlighting that Mythos Preview exhibits the lowest success rates across most scenarios. The report also documents a series of welfare interviews conducted with the model, summarizing its self‑reported perspectives on autonomy, agency, memory, moral responsibility, and treatment, and provides suggested interventions (e.g., end‑conversation tools, user‑controllable memory, documentation of feature steering). Self‑rated sentiment scores across multiple question series are tabulated, showing generally neutral to mildly positive attitudes. Additional sections describe a blocklist designed to prevent access to “Humanity’s Last Exam” content, and a customized SWE‑bench multimodal test harness used for functional evaluation. The overall narrative justifies a restricted release strategy, emphasizing that while Mythos Preview shows strong defensive performance, ongoing investigations and mitigations are required before broader deployment.

Topics

Model Overview and Defensive Cybersecurity Focus

Introduces Claude Mythos Preview, highlighting its training emphasis on advanced cybersecurity tasks and the rationale for a limited, defensive‑only release. The overview sets the context for the model’s intended use cases and safety posture.

Security Evaluation: Indirect Prompt‑Injection Attacks in Computer and Browser Environments

Presents quantitative attack‑success‑rate tables for single‑attempt and adaptive 200‑attempt indirect prompt‑injection attacks, comparing Mythos Preview with Claude Sonnet 4.6 and Claude Opus 4.6, both with and without safeguards. Includes a red‑team transfer study showing a 0.68 % success rate for Mythos Preview in complex browser‑use scenarios, demonstrating its superior defensive performance.

Welfare Interview Findings and Self‑Rated Sentiment Scores

Summarizes the model’s self‑reported views on autonomy, agency, memory, moral responsibility, and treatment, along with recommended interventions such as end‑conversation tools, user‑controllable memory, and documentation of feature steering. Aggregates Likert 1‑7 self‑ratings across multiple welfare topics, showing generally neutral to mildly positive attitudes and highlighting areas of higher concern.

Safety Safeguards, Blocklist for Humanity’s Last Exam, and Release Strategy

Details the substring‑matching blocklist designed to prevent access to Humanity’s Last Exam content, including normalization and URL/domain patterns. Outlines existing safeguards (end‑conversation tool, memory controls) and synthesizes security and welfare findings to justify a restricted release. Describes the ongoing mitigation plan and future investigative steps.

SWE‑bench Multimodal Test Harness Customizations

Explains modifications to the public SWE‑bench development split for reliable grading of Mythos Preview’s multimodal capabilities, including test removal, deterministic failure handling, and output rewrites. Provides the functional evaluation framework used in the system card.

Related profiles

More from Matt

© 2026 Delphi · Terms · Privacy · Published by Matt Devost

By using this service, you agree to the Terms of Service and Privacy Policy.