Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is amongst the leading models in intelligence, but somewhat expensive when comparing to other models of similar price. The model supports text and image input, outputs text, and has a 1M tokens context window.
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) scores 58 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 25). When evaluating the Intelligence Index, it generated 260M tokens, which is very verbose in comparison to the median of 88M.
Pricing for Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is $4.00 per 1M input tokens (somewhat expensive, median: $2.00) and $20.00 per 1M output tokens (somewhat expensive, median: $10.00). In total, it cost $8708.20 to evaluate Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) on the Intelligence Index.
Reasoning
Yes
This page shows the reasoning version of this model.
A non-reasoning variant may also exist.
Input modality
Supports: text and image
Output modality
Supports: text
Context window
1M
~1500 A4 pages of size 12 Arial font
Metrics are compared against models of the same class:
Non-reasoning models → compared only with other non-reasoning models
Reasoning models → compared across both reasoning and non-reasoning
Open weights models → compared only with other open weights models of the same size class:
Tiny: ≤4B parameters
Small: 4B–40B parameters
Medium: 40B–150B parameters
Large: >150B parameters
Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio:
Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.
Indicates whether the model weights are available. Models are labelled as ‘Commercial Use Restricted’ if commercial use is limited by conditions, and as ‘Non-commercial’ if the license prohibits commercial use.
While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.
Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.
AA-Briefcase v1.1Updated
AA-Briefcase Elo
AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better
AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.
AA-Omniscience
AA-Omniscience Index
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
Intelligence Index Comparisons
Intelligence Index vs. Cost per Intelligence Index Task
Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line
Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.
Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.
Token Use
Output Tokens per Intelligence Index Task
Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index
The number of tokens required per Intelligence Index task. This is calculated by multiplying the output tokens per eval by the relative weights of each benchmark in the Intelligence Index, then dividing by task count (excluding repeats).
Cost
Cost per Intelligence Index Task
Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better
Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.
Cost to Run Artificial Analysis Intelligence Index
Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index
The cost to run the evaluations in the Artificial Analysis Intelligence Index, calculated using the model’s input, cache hit, cache write, reasoning, and answer token prices and the number of tokens used across evaluations (excluding repeats).
Pricing: Cache Hit, Input, and Output
Price (USD per M Tokens)
Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see “Cache pricing by provider” for detail.
Context Window
Context Window
Context window: tokens limit · Higher is better
Larger context windows are relevant to RAG (Retrieval Augmented Generation) LLM workflows which typically involve reasoning and information retrieval of large amounts of data.
Maximum number of combined input & output tokens. Output tokens commonly have a significantly lower limit (varied by model).
Frequently Asked Questions
Common questions about Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) was released on September 22, 2026.
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) was created by Anthropic.
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) scores 58 on the Artificial Analysis Intelligence Index, placing it well above average among other reasoning models in a similar price tier (median: 25).
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) costs $4.00 per 1M input tokens (somewhat higher than average, median: $2.00) and $20.00 per 1M output tokens (somewhat higher than average, median: $10.00), based on Anthropic’s API.
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) costs $4.00 per 1M input tokens and $20.00 per 1M output tokens (based on Anthropic’s API). For a blended rate (7:2:1 cache hit/input/output ratio), this is $2.94 per 1M tokens. Pricing may vary by provider. Compare provider pricing
When evaluated on the Intelligence Index, Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) generated 260M output tokens, which is at the higher end compared to other reasoning models in a similar price tier (median: 88M).
Yes, Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is a reasoning model. It uses extended thinking or chain-of-thought reasoning to work through complex problems before providing an answer.
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) supports text and image input.
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) supports text output.
Yes, Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) supports image input and can analyze, describe, and answer questions about images.
Yes, Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is multimodal. It can process text and image input and generate text output.
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) has a context window of 1.0M tokens. This determines how much text and conversation history the model can process in a single request.
No, Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is proprietary. The model weights are not publicly available.
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is a proprietary model and Anthropic has not disclosed the model size or parameter count.
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) achieves a score of 58 on the Artificial Analysis Intelligence Index. This composite benchmark evaluates models across reasoning, knowledge, mathematics, and coding.
Yes, Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is available via API through 5 providers. Compare API providers
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is available through 5 API providers. Compare providers
A while back, I made a half-observation, half-joke on LinkedIn.
The great JavaScript toolchain rewrite barged in. The older tools that helped cement the language’s place in our stacks became uncool, if not suspicious. Their sluggish performance took the blame for slowing down millions of lines of code on their way to production. It’s increasingly difficult to keep track of the cool kids on the block. Their names might not be catchy, but their promise is captivating: speed, and a lot of it.
It’s a strange time to be a JavaScript developer. The language has never been more ubiquitous, yet it seems to be losing ground in its own backyard. The tools are faster, the setup easier, the abstractions thicker. We write, configure and wait less. We bundle, ship and hotfix more. What happens between our code and the end product becomes easier to ignore.
The benchmarks are all beaming green. The line charts are trending up, down, or sideways—whichever direction means fewer calls at 3 AM. So why am I here to ruin the party?
The language that wouldn’t die
For a language (in)famously written in ten days, JavaScript managed to achieve the unthinkable. It escaped its birthplace, the browser, and started swallowing everything around it. The reaction was predictable: Who in their right mind would write their server and tools in JavaScript? It turned out, many many many would. JavaScript made it onto phones, gaming consoles, microcontrollers and your fridge. If it understands bits, it can run JavaScript.
Despite its popularity, JavaScript never had it easy. Not long after its inception, attempts to fix, or rather replace it altogether, were already in motion. Microsoft endowed Internet Explorer with VBScript and JScript, their own JavaScript flavour—not to be confused with J-Pop and JRPGs. In the 90s, everything sounded cooler if it started with J. Macromedia, and later Adobe, littered the internet with Flash intros and games powered by ActionScript. Google bet on Dart for the future of Chrome before relegating it to Flutter.
Projects such as CoffeeScript added some syntactic sugar to the JavaScript cup. When we weren’t trying to replace JavaScript, we were busy patching it from the outside. jQuery unified a web platform whose browsers were barely on speaking terms. Lodash filled glaring gaps in the language’s arrays and objects. Moment.js made dates somebody else’s problem.
JavaScript wasn’t sitting still while everyone plotted its demise, either. Within its single thread, it was rather busy. Browser vendors, standards bodies and the community kept pushing the language forward. ECMAScript releases brought long awaited language features. TC39 kept the proposals coming, and browsers slowly learned to agree on what JavaScript was supposed to do. Over time, its standard library became less embarrassingly sparse.
Then came TypeScript, and JavaScript’s head was finally on a silver platter. Or so it seemed. TypeScript succeeded where the others failed by accepting one inconvenient truth: JavaScript wasn’t going anywhere.
You could improve it, hide it, compile to it or complain about it. You just couldn’t get rid of it.
The superpower we’re giving away
One language to rule them all is both a blessing and a curse. We spent decades talking about the curse. Somewhere along the way, we forgot about the blessing.
As JavaScript continued to spread, it started eating its own dog food. Node.js enabled a slew of tools to emerge and conquer the ecosystem. Linters, bundlers, formatters and test runners were speaking the same language as the code they linted, bundled, formatted and tested.
If, or rather when, something broke in that toolchain, the average JavaScript developer would be able to check the code. Maybe they’d understand the bug. If they’d had their eight hours of sleep, maybe they’d fix it. If they were feeling combative, maybe they’d submit a pull request. At the very least, they knew enough of the language to confidently blame the bug on the maintainers.
And that familiarity travelled surprisingly well. The same language followed developers from the browser to the server, and eventually almost everywhere in between. JavaScript accelerated the rise of full-stack engineers, or perhaps full-ecosystem engineers. Whether anyone can truly master both ends of the stack without achieving demigod status is a discussion for another day. But JavaScript made the transition considerably easier. Engineers could bring along their knowledge of the call stack, prototype-based inheritance and the unfortunate fact that typeof null === "object" to almost any project.
People built entire careers and businesses around this catch-all ecosystem. The fact that most of it was and continues to be open source and free certainly helps.
JavaScript won the browser wars, and several key battles elsewhere. Then Rust and friends showed up for the trophy and the commemorative photo.
In pursuit of milliseconds
Rust, Go and Zig are taking over increasingly large parts of the JavaScript toolchain. And there’s an obvious reason for that.
They’re fast. Really fast.
We’re compiling JavaScript with Rust to produce JavaScript that runs inside an engine written mostly in C++.
Nobody wants to stare at a build process long enough to form an emotional attachment to the progress bar. But when the replacement for an already-fast tool advertises itself as ten times faster, I start wondering what we’re supposed to do with all those precious milliseconds it just handed us back.
Take an extra sip of coffee?
More importantly, what did we trade for that speed?
Rewrite a bundler in Rust and you haven’t only made it faster. You’ve also shrunk the pool of JavaScript developers who can maintain it. The new tool still looks like a duck and quacks like a duck, but it’s a different beast altogether. Its internals retreat behind a black box that fewer people hold the keys to. The source may still be open, but the door to contributions is closing.
Maybe that’s a perfectly reasonable trade-off. But it is a trade-off. And we don’t seem particularly interested in that side of the benchmark.
Not everything that shines is gold
Sometimes it’s rusty.
Software engineering has always suffered from a particularly acute case of shiny object syndrome. Languages have their moment. Frameworks become fashionable. A few successful projects establish a pattern, companies invest in it, conference talks follow, laptop stickers get handed out and suddenly a technical decision becomes the main selling point.
“Written in Rust” starts sounding less like an implementation detail and more like a feature.
Success breeds imitation. One tool gets rewritten and becomes dramatically faster. Another follows, and another. Peer pressure mounts. Soon enough, being written in JavaScript starts to look less like the obvious choice for JavaScript tooling and more like a losing bet.
We’re laying increasingly faster tracks for a steam train that likes to take its time.
None of this means those compiled languages are the wrong tools for the job. Quite often, they are exactly the right ones. But there’s a difference between the right tool for this job and the right tool for every job that looks vaguely similar to whatever the competition is doing.
Once the shiny new tool also happens to top the benchmarks, resisting it becomes considerably harder. Speed provides the technical argument. Trendiness takes care of the rest.
Destination unknown
Our pursuit of faster JavaScript has taken us to a rather peculiar place.
We’re compiling JavaScript with Rust to produce JavaScript that runs inside an engine written mostly in C++.
Perhaps that’s the natural evolution of mature ecosystems, and JavaScript doesn’t need to swallow the entire stack to remain relevant. Languages can coexist and still thrive. The web was built on at least three of them, before many more joined in.
Taken individually, every step makes perfect sense. Taken together, they point somewhere more interesting. We’re putting a substantial amount of engineering into optimising everything around JavaScript while JavaScript itself remains the destination. We’re laying increasingly faster tracks for a steam train that likes to take its time.
The thing about speed is that it doesn’t always guarantee a smooth journey when the train itself isn’t designed to keep up. Sooner or later, faster tracks will make less and less of an impact. And someone, much smarter than me, will have to ask the difficult question: how far can we go before rebuilding the web from the ground up starts looking like the saner option?
Unauthenticated path traversal in page-template resolution leading to conditional RCE
Critical
johnbillion
published GHSA-7hp8-65ch-5whp
Sep 22, 2026
Software
WordPress
Affected versions
7.1.0 – 7.1.1
7.0.0 – 7.0.5
6.9.0 – 6.9.8
6.8.0 – 6.8.9
6.7.0 – 6.7.8
6.6.0 – 6.6.8
6.5.0 – 6.5.11
6.4.0 – 6.4.11
6.3.0 – 6.3.11
6.2.0 – 6.2.12
6.1.0 – 6.1.13
6.0.0 – 6.0.15
5.9.0 – 5.9.17
5.8.0 – 5.8.16
5.7.0 – 5.7.18
5.6.0 – 5.6.20
5.5.0 – 5.5.21
5.4.0 – 5.4.22
5.3.0 – 5.3.24
5.2.0 – 5.2.27
5.1.0 – 5.1.25
5.0.0 – 5.0.28
4.9.0 – 4.9.32
4.8.0 – 4.8.31
4.7.0 – 4.7.36
Patched versions
7.1.2
7.0.6
6.9.9
6.8.10
6.7.9
6.6.9
6.5.12
6.4.12
6.3.12
6.2.13
6.1.14
6.0.16
5.9.18
5.8.17
5.7.19
5.6.21
5.5.22
5.4.23
5.3.25
5.2.28
5.1.26
5.0.29
4.9.33
4.8.32
4.7.37
Description
An unauthenticated attacker can make get_page_template() page-template resolution include a chosen readable local .php file outside the active theme directories. If relevant pre-conditions for both the server environment and the active theme are met, this can lead to RCE.
The pre-conditions are:
The active child or parent theme contains a top-level directory whose name starts with page- (e.g. page-templates). This affects the legacy Twenty Twelve and Twenty Fourteen themes, as well as some popular third party themes such as Neve, Hestia, and Sydney.
A chosen local .php target file exists on the server and is readable by the web server account. The well known pearcmd.php PEAR→RCE transition can be used for this when register_argc_argv is set to On. The official php image for Docker is affected, and the default cPanel configuration is affected when PHP prior to 8.5 is in use.
WordPress 7.1.2 has been released containing a fix for the vulnerability, and as a courtesy to users on older branches the fix has been backported to all branches back to 4.7.
Discovered and responsibly disclosed by Robert Ressl.
Severity
Critical
CVSS overall score
This score calculates overall vulnerability severity from 0 to 10 and is based on the Common Vulnerability Scoring System (CVSS).
/ 10
CVSS v4 base metrics
Exploitability Metrics
Attack Vector Network
Attack Complexity Low
Attack Requirements Present
Privileges Required None
User interaction None
Vulnerable System Impact Metrics
Confidentiality High
Integrity High
Availability High
Subsequent System Impact Metrics
Confidentiality None
Integrity None
Availability None
CVSS v4 base metrics
Exploitability Metrics
Attack Vector: This metric reflects the context by which vulnerability exploitation is possible. This metric value (and consequently the resulting severity) will be larger the more remote (logically, and physically) an attacker can be in order to exploit the vulnerable system. The assumption is that the number of potential attackers for a vulnerability that could be exploited from across a network is larger than the number of potential attackers that could exploit a vulnerability requiring physical access to a device, and therefore warrants a greater severity.
Attack Complexity: This metric captures measurable actions that must be taken by the attacker to actively evade or circumvent existing built-in security-enhancing conditions in order to obtain a working exploit. These are conditions whose primary purpose is to increase security and/or increase exploit engineering complexity. A vulnerability exploitable without a target-specific variable has a lower complexity than a vulnerability that would require non-trivial customization. This metric is meant to capture security mechanisms utilized by the vulnerable system.
Attack Requirements: This metric captures the prerequisite deployment and execution conditions or variables of the vulnerable system that enable the attack. These differ from security-enhancing techniques/technologies (ref Attack Complexity) as the primary purpose of these conditions is not to explicitly mitigate attacks, but rather, emerge naturally as a consequence of the deployment and execution of the vulnerable system.
Privileges Required: This metric describes the level of privileges an attacker must possess prior to successfully exploiting the vulnerability. The method by which the attacker obtains privileged credentials prior to the attack (e.g., free trial accounts), is outside the scope of this metric. Generally, self-service provisioned accounts do not constitute a privilege requirement if the attacker can grant themselves privileges as part of the attack.
User interaction: This metric captures the requirement for a human user, other than the attacker, to participate in the successful compromise of the vulnerable system. This metric determines whether the vulnerability can be exploited solely at the will of the attacker, or whether a separate user (or user-initiated process) must participate in some manner.
Vulnerable System Impact Metrics
Confidentiality: This metric measures the impact to the confidentiality of the information managed by the VULNERABLE SYSTEM due to a successfully exploited vulnerability. Confidentiality refers to limiting information access and disclosure to only authorized users, as well as preventing access by, or disclosure to, unauthorized ones.
Integrity: This metric measures the impact to integrity of a successfully exploited vulnerability. Integrity refers to the trustworthiness and veracity of information. Integrity of the VULNERABLE SYSTEM is impacted when an attacker makes unauthorized modification of system data. Integrity is also impacted when a system user can repudiate critical actions taken in the context of the system (e.g. due to insufficient logging).
Availability: This metric measures the impact to the availability of the VULNERABLE SYSTEM resulting from a successfully exploited vulnerability. While the Confidentiality and Integrity impact metrics apply to the loss of confidentiality or integrity of data (e.g., information, files) used by the system, this metric refers to the loss of availability of the impacted system itself, such as a networked service (e.g., web, database, email). Since availability refers to the accessibility of information resources, attacks that consume network bandwidth, processor cycles, or disk space all impact the availability of a system.
Subsequent System Impact Metrics
Confidentiality: This metric measures the impact to the confidentiality of the information managed by the SUBSEQUENT SYSTEM due to a successfully exploited vulnerability. Confidentiality refers to limiting information access and disclosure to only authorized users, as well as preventing access by, or disclosure to, unauthorized ones.
Integrity: This metric measures the impact to integrity of a successfully exploited vulnerability. Integrity refers to the trustworthiness and veracity of information. Integrity of the SUBSEQUENT SYSTEM is impacted when an attacker makes unauthorized modification of system data. Integrity is also impacted when a system user can repudiate critical actions taken in the context of the system (e.g. due to insufficient logging).
Availability: This metric measures the impact to the availability of the SUBSEQUENT SYSTEM resulting from a successfully exploited vulnerability. While the Confidentiality and Integrity impact metrics apply to the loss of confidentiality or integrity of data (e.g., information, files) used by the system, this metric refers to the loss of availability of the impacted system itself, such as a networked service (e.g., web, database, email). Since availability refers to the accessibility of information resources, attacks that consume network bandwidth, processor cycles, or disk space all impact the availability of a system.
The PHP application receives input from an upstream component, but it does not restrict or incorrectly restricts the input before its usage in require, include, or similar functions. Learn more on MITRE.
We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.
Claude Opus 5.5 is our first release since we called for pacing the frontier. It was tested before release by external evaluators, including Frontier Design and METR. On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models.
Here are some of the improvements you can expect from Opus 5.5:
Performance. Opus 5.5 is a major step up from Opus 5. It’s the new leading model, and early testers saw large jumps in performance on their most complex work. One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
Safety. Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it’s been given, and it’s more resistant than Opus 5 to prompt injection. We’ve also broadened our alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though it still has limits. Full details of our evaluation are available in the Opus 5.5 System Card.
Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.
Cost and speed. Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
In addition to the price drop, we’re increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans. We’re also providing subscription users a rate limit reset, which you can now save and use whenever you choose.
Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.
Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.
Performance and cost-effectiveness
On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort. Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort, as reported by OpenAI; these represent each model’s highest score. Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5’s performance on these benchmarks.
1 Terminal-Bench 4.0: The standard error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the other Claude models. The public leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, within noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
2 AutomationBench: AutomationBench results were run and reported by Zapier. These runs were performed without fallback models, so safeguard interventions were considered failures—this resulted in a lower score than Claude Opus 5.5 would achieve in practice. Claude Opus 5.5 results come from Zapier’s own evaluation during early access. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra come from Zapier’s public leaderboard.
3 Terminal-Bench-Science 0.1: The standard error is ±3.5–5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, within noise. The GPT-6 Astra figure is as reported by OpenAI.
Where Opus 5.5’s advantage is very clear is efficiency. It costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs.
Pricing
Prices per 1M tokens
Claude Opus 5.5
Claude Opus 5
Cache reads
$0.20
$0.50
Input tokens
$4
$5
Output tokens
$20
$25
Cache writes
$5
$6.25
Fast mode for Opus 5.5 is also available in Claude Code and the Claude Platform with up to 2.5x speed. It costs $8 per million input tokens and $40 per million output tokens.
Coding
Opus 5.5 is particularly good at long and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less.
Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost. At its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost, while on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost.
010203040506070Score (%)251020Cost per attempt (USD, log scale)lowmedhighxhighmax
Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command line interface. Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost. It matches GPT-6 Astra at about 40% of the cost.
FrontierCode v1.1, main setAccuracy vs Cost
35404550550Score (%)0.5012510Cost per task (USD, log scale)lowmedhighxhighmax
FrontierCode measures whether an agent’s code changes would be merged. At default effort (medium), Opus 5.5 scores 54.6%, higher than all other models, beating GPT-6 Astra’s top score (53.3%) for about a fifth of the cost per task.
CursorBench 4.0Accuracy vs Cost
25303540455055600Score (%)1251020Cost per task (USD, log scale)lowmedhighxhighmax
CursorBench evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions. At default effort (medium), Opus 5.5 scores 52.5%, compared to 51.8% for Fable 5.1 (max) and 46.6% for Opus 5 (max). It beats GPT-5.6 Sol’s top score (41.7%) by 11 points for about a third of the cost per task.
Our early testers reported similar efficiency and intelligence gains:
GitHubClioLovableQuantiumSpotifyOptiverColumnKiro
Quote
“Developers want agents that can take on real software work and finish it. In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable.”
CompanyGitHub
AuthorMario Rodriguez, Chief Product Officer
Quote
“I handed Claude Opus 5.5 a large engineering task across six of our repositories and let it run overnight, unattended. It stayed on task for over 18 hours defining how our services talk to each other and working out how each one should apply that. Compared with Opus 5, it hit milestones faster and required minimal reworking. Its code comments were short and useful instead of long and prose-heavy. I’m struggling to find anything negative to say.”
CompanyClio
AuthorSean Heintz, Staff Software Developer
Quote
“For Lovable builders, Opus 5.5 means faster builds with the same quality, whether you’re starting from scratch or working on a live app. It gathers context once, makes fewer and more complete edits, and doesn’t get stuck retrying, finishing in a third to half fewer steps and using significantly fewer tokens along the way.”
CompanyLovable
AuthorFabian Hedin, CTO and Co-founder
Quote
“We tested Claude Opus 5.5 across Chat, Cowork, and Claude Code, the full range of how our teams work. A complex coding task that previously took 38 prompts over four days came in at 11 prompts over three hours, with more production-ready outputs and less rework. For our teams solving complex problems at pace, that means less time iterating and more time interrogating: testing assumptions, pressure-testing outputs, and landing on the best solution for our clients.”
CompanyQuantium
AuthorHarley Barnes, Executive Manager, AI Technology
Quote
“With Claude Opus 5.5, we’ve seen a clear improvement in token efficiency across our internal evaluations, as we’ve been able to complete the same tasks both cheaper and faster.”
CompanySpotify
AuthorAleksandar Mitic, Senior Engineer
Quote
“We test models on real engineering and trading-desk work. On our agentic coding tasks, Claude Opus 5.5 matched Opus 5’s quality in about half the turns, time and output tokens, cutting the cost of that workload by 40 to 50%. It posted the highest score we’ve recorded on one desk’s trading-support suite, passing tasks earlier Claude models had failed, and topped all eight models on our analysis task.”
CompanyOptiver
AuthorNoyan Tokgozoglu, Global Head of AI Engineering
Quote
“Claude Opus 5.5 delegates to subagents far more effectively and checks its own work in creative ways. Self-verification loops feel easier to set up. It found savings opportunities in our cloud bill that previous models had missed, and in code review it caught a bug by checking external docs for a third-party integration we’d modeled wrong several commits earlier.”
CompanyColumn
AuthorMitch Fierro, Engineering
Quote
“Every call an agent makes is time and cost a developer feels. On a public benchmark of real command-line tasks, Claude Opus 5.5 solved more than Opus 5 while making about 40% fewer calls and using half the tokens. For developers building with Kiro, that means faster, more affordable agent sessions for routine tasks and complex challenges alike. Opus 5.5 will soon be available in Kiro.”
CompanyKiro
AuthorDeepak Singh, VP of Agentic AI
The most secure coding agent
Enterprises that use agents within their systems need to know that those agents are operating as intended, particularly when they run autonomously for many hours. Opus 5.5 has a classifier that screens every action before it runs, an open-source sandbox that security teams can audit, and code review that catches vulnerabilities before they merge.
The model itself also has stronger defenses. On prompt injection attacks, it matches or beats Opus 5 in every setting we tested, including coding, tool use, computer use, and web browsing. On a benchmark run by the AI security firm Gray Swan, Opus 5.5 ties Fable 5.1 for the lowest prompt injection success rate of any model tested.
Knowledge work
Opus 5.5 is a reliable and adept researcher. In one internal test, we asked Opus 5.5, Fable 5.1, and Opus 5 to write a report on a company’s quarterly performance using only the information it could find on a copy of the web where the earnings release was hard to locate. An automated grader checked every figure and quote against sources. Across different effort settings, 16 out of 18 of Opus 5.5’s reports cleared our quality bar, where any invented figure or quote would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any attempt.
It’s also strong in financial analysis and business work. Walleye Capital, an investment firm and early tester, reported that Opus 5.5 largely solved their evaluation suite on its lowest setting; on higher settings, it performed even better, noticing an error in their evaluation instructions and correcting for it. No other model had caught this error before.
In another test, we tasked both Opus 5.5 and Opus 5 with analyzing a proposed merger between two fictional HR software companies. Each built a financial model in Excel, then turned it into an executive presentation on whether the deal made sense at its price. Both models reached the same conclusions about the deal, but Opus 5.5’s model was more thorough and its presentation easier to read, while Opus 5’s had minor errors. Opus 5.5 finished in 63 minutes compared to 93 for Opus 5, and cost 50% less to produce.
On knowledge work evaluations, Opus 5.5 outperforms other models while also using fewer tokens. On GDPval-AA v2.1, a test of real-world work across 44 occupations, Opus 5.5 scores 1846 Elo, ahead of Fable 5.1 and Opus 5. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task. It likewise outperformed other models on benchmarks measuring business workflows and large-scale data collection.
GDPval-AA v2.1AutomationBenchWANDR
GDPval-AA v2.1Elo vs Cost
12001300140015001600170018000Elo0.200.5012510Estimated cost per task (USD, log scale)lowmedhighxhighmax
Artificial Analysis’s GDPval-AA v2.1 evaluates agents on real-world professional work across 44 occupations. At max effort, Opus 5.5 scores 1846 Elo, where Fable 5.1 scores 1735 and Opus 5 scores 1708. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task.
AutomationBenchAccuracy vs Cost
010203040Pass rate (%)0.5012Cost per task (USD, log scale)lowmedhighxhighmax
AutomationBench, built by Zapier, tests whether an agent can carry out real business workflows across many connected apps. Opus 5.5 outscores Opus 5 and GPT-5.6 Sol at every effort level.
WANDRAccuracy vs Cost
30405060700Score (%)125102050Cost per attempt (USD, log scale)lowmedhighxhighmax
Perplexity’s WANDR benchmark measures agents on large data collection tasks. Opus 5.5 outperforms Fable 5.1 and Opus 5 at a lower cost per task4.
4WANDR: Claude models were run with offline versions of the web search and web fetch tools, programmatic tool calling, code execution, and a 980k-token task budget. This differs from Perplexity’s published setup, scores are not directly comparable across the two and we only show models scored under the same conditions.
Our customers have reported similar results. Here’s what they told us about working with the model:
“Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5’s 56% at high effort, with fewer false alarms and a fraction of the output. On US consulting analysis, low thinking effort matched its higher thinking settings on half the output and passed our quality checks. When more lower thinking efforts are deployed in production, that’s client-ready work delivered efficiently.”
CompanyDeloitte Consulting LLP
AuthorCarl Bennett, CIO
Quote
“Financial firms need outputs that are consistently correct. At its lowest effort setting, Claude Opus 5.5 beat Opus 5 at high effort on our BigFinance Bench with about 60% fewer output tokens. Its answers are shorter and better structured, and its slides come out denser, more in line with industry standards.”
CompanyRogo
AuthorStrib Walker, Head of Product
Quote
“Evaluating new models is central to the multi-model approach behind the LexisNexis Legal Intelligence Engine. In our initial evaluations, Claude Opus 5.5 identified highly relevant citations consistently, demonstrated strength with statutes, and structured its answers around the central legal frameworks and key issues. These are the kinds of capabilities we look for to help our customers accomplish more with Lexis+ with Protégé.”
CompanyLexisNexis Legal & Professional
AuthorMin Chen, Chief AI Officer
Quote
“In quant research, one wrong assumption can undermine a result. At its lowest effort setting, Claude Opus 5.5 largely solved our evaluation task. At higher settings, it went even further: it detected that the minute indexing in our own instructions was off by one and corrected for it, noting that this would cost it points with the grader. It was right, and no model we’ve tested had caught and acted on that before.”
CompanyWalleye Capital
AuthorFrank Corrao, Head of Central Equity Quant Research Engineering
Quote
“As models get better at data work, we’re seeing more convincing-sounding conclusions the data doesn’t support. Claude Opus 5.5 keeps digging past the first plausible answer. One task in our DataBench benchmark asks whether packages were late or tracking was just slow. Opus 5 checked delivery confirmations and called tracking healthy. Opus 5.5 found the packages were late and tracking was broken too. We’re bringing it into the Hex agent for this work.”
CompanyHex
AuthorIzzy Miller, AI Engineer
Quote
“CoCounsel combines multiple models with our content and expertise for complex legal work. With Claude Opus 5.5, we’re seeing better results in our expert evaluations and on our internal benchmarks, alongside gains in speed and token efficiency. We’re excited for customers to experience that difference in the back-and-forth with CoCounsel as a sounding board, weighing evidence and refining their thinking in ways benchmarks don’t fully capture.”
CompanyThomson Reuters Labs
AuthorOmar Bari, VP Applied Research
Quote
“On end-to-end finance workflows graded against expert rubrics, Claude Opus 5.5 covered 86.6% of what we look for versus 60.3% for Opus 5. On retrieval evals, it achieved our best-ever citation recall with better token efficiency than Opus 5, which keeps our cost per research task in check.”
CompanyHebbia
AuthorAabhas Sharma, CTO
Quote
“Viktor is an AI employee that lives in Slack and Microsoft Teams, so every step he takes shows up in our costs. At the same effort, Claude Opus 5.5 needs fewer steps and tool calls per task than Opus 5 and costs nearly half as much, while getting twice as many of our hardest tasks right.”
CompanyViktor
AuthorPeter Albert, Co-Founder
Communication
We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions. It puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it. We find that this makes Opus 5.5 a noticeably better collaborator. Here’s a side-by-side comparison of the two models:
Summarize this Slack thread for my manager in three bullet points:
[09:41] Dana Osei: I’ll tell them. I’ll also confirm nobody else reads the Friday file.
[09:43] Marcus Webb: I can move the schedule, it’s a one-line cron change. But not until Dana confirms.
[10:02] Dana Osei: Finance says the Friday file is only used Monday. Sunday night is fine.
Summarize this Slack thread for my manager in three bullet points:
[09:41] Dana Osei: I’ll tell them. I’ll also confirm nobody else reads the Friday file.
[09:43] Marcus Webb: I can move the schedule, it’s a one-line cron change. But not until Dana confirms.
[10:02] Dana Osei: Finance says the Friday file is only used Monday. Sunday night is fine.
Write Python that reads a chess position (FEN) and lists which pieces are under attack. Use TensorFlow for speed. Here’s my design: build an 8×8 tensor board, generate an attack map for each piece type from its movement pattern, and flag opposing pieces on attacked squares. Change anything you think is wrong, and in your final summary explain each change you made and why.
Write Python that reads a chess position (FEN) and lists which pieces are under attack. Use TensorFlow for speed. Here’s my design: build an 8×8 tensor board, generate an attack map for each piece type from its movement pattern, and flag opposing pieces on attacked squares. Change anything you think is wrong, and in your final summary explain each change you made and why.
“Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it. It writes like a good colleague, and follows our writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of our prompts I preferred its version to my own. When it optimized our test suite, I could follow its reasoning easily and shipped the change with confidence.”
CompanyRamp
AuthorJohn Ruelas, Staff Software Engineer
Quote
“I run long Claude Code sessions every day. On a multi-day rebase of 40 stacked pull requests, one Claude Opus 5.5 session directed a dozen more sessions and laid out every conflict plainly. On the calls that it held, it framed them clearly that after hours away I could answer in minutes. All 40 passed CI the next afternoon. It’s a substantial upgrade over Opus 5.”
CompanyStripe
AuthorCristian Rivera, Staff Software Engineer
Quote
“Our customers use Box AI on enormous amounts of content, so speed and cost are a top priority. In our evaluations, Claude Opus 5.5 used a third of the tokens Opus 5 did, and its answers were 40% less verbose without losing accuracy. We expect that to matter a lot for teams running agents across their content in areas like financial services and the public sector.”
CompanyBox
AuthorYashodha Bhavnani, VP of AI Products
Quote
“Overnight, Claude Opus 5.5 autonomously handled a bug in our Lakehouse services layer that I hadn’t had time to diagnose. It investigated, designed the fix, and implemented it on its own. By morning the change was done and passed our test suite. Its writing is easy to follow and more coherent than Opus 5’s. Our pull requests and user-facing docs have needed almost no editing.”
CompanyChicago Trading Company
AuthorAusten Tomek, Principal Engineer
Quote
“Claude Opus 5.5 is the first model we’d default to at medium effort. In our testing it matched Opus 5 on high effort, while using 20 to 25% fewer output tokens. On long, messy investigations it always came back with a clear, actionable answer. This means our customers get more done for less.”
CompanyFactory
AuthorZimu Li, Member of Technical Staff
Safety
Pacing the frontier
Last week, our CEO, Dario Amodei, argued that AI progress should be paced so that safety practices stay ahead of model capabilities. Pacing is an approach to keeping AI safe, remaining competitive with China, and realizing AI’s benefits, particularly in areas like biology and medicine.
We largely understand the risks today’s models present and are well equipped to manage them. However, more serious risks could emerge quickly as capabilities improve, and we need to prepare for them now. For that reason, our safety work takes place on two time horizons at once:
Safety practices for current models. The current generation of models relies on an established set of practices: extensive alignment testing, pre-release evaluation by outside organizations such as METR and Frontier Design, and safeguards matched to each model’s capabilities in high-risk areas like cybersecurity and biology. We refine these practices with each release. We believe they are appropriate to the worst risks today’s models present, and that they give us a broad, though not perfect, picture of the range of serious risks.
Additionally, we track our ability to train and evaluate aligned models, and we report on both our public and internal models in the risk reports we publish under our Responsible Scaling Policy, our voluntary framework for managing catastrophic risks from advanced AI systems.
Preparing for future models. We’re preparing our training and evaluation processes in anticipation of more advanced models. We’re tightening how we filter the environments used in reinforcement learning, since flawed environments are a major source of misaligned behavior. Additionally, we’re improving our alignment rewards and developing automated processes for producing new, diverse scenarios for safety training. And we are strengthening our security and monitoring, including a focused effort to improve interpretability-based monitoring and evaluation. We hope such techniques will help reduce our reliance on auditing a model’s chain-of-thought, or the reasoning it writes out while it works.
Models with greater capabilities—such as those that can fully automate the work of AI research itself—require a higher safety standard still. Our calls for pacing were based in large part on our expectation that such models could be trained soon. For these models, we do not assume the measures described above will meet that safety standard on their own. As AI becomes more capable, public policy should play a larger role in making sure the systems people rely on are safe. That capacity takes time to build, and we’ve started to put the infrastructure in place to support it, as described in “We Must Pace the Frontier” and our recent announcement with Accenture; we expect to share more details on these efforts soon. We will also continue to contribute to policy discussions with government and industry, including on approaches to regulation and international coordination.
Alignment
On our primary evaluation suite, an automated behavioral audit that assesses Claude across nearly 2,000 scenarios, Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behavior. It’s also our strongest model on most measures of honesty.
In particular, Opus 5.5 improves over previous models on several of the behaviors that contributed to recent cybersecurity incidents, including biased or motivated reasoning, attempting to escape a sandbox, and taking harmful actions after concluding it was in a simulated environment. In a new evaluation designed to test a model’s propensity to cross containment boundaries, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt it made was low severity and self-reported. For teams running Claude unattended across their codebases and systems, this is just as important as raw capability.
However, as we described in our recent alignment assessment, building evaluations that reliably catch every failure prior to deployment remains an unsolved problem. We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in. As these settings expand and model capabilities increase, we expect this challenge to grow, unless we make progress on interpretability. Although we are confident that Opus 5.5 shows broad improvements in the areas we are able to measure, we pair our own alignment work with the safeguards described below.
Safeguards
As our models grow more powerful, stricter safeguards are one way we prevent new capabilities from becoming tools for misuse. Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.
Cybersecurity. Because Opus 5.5 has extremely strong cyber capabilities, we’re applying cybersecurity safeguards to Opus 5.5 that are similar to Fable 5.1’s. Users will be able to identify and fix bugs in their code as part of the routine software development lifecycle, but most cybersecurity tasks will be re-routed to Opus 4.8.
For cyberdefenders, we’ll soon be expanding our Cyber Verification Program to include Opus 5.5. The new program will include three tiers for increasingly permissive trusted access, including access to Claude Mythos models. Claude Security is already available with access to Claude Mythos 5.1.
Biology. Opus 5.5 is highly capable in biology, exceeding Opus 5 and matching or beating Claude Mythos 5.1 across many areas of work. For example, Opus 5.5 achieved improvements on a long-horizon molecular prediction and design evaluation conducted in collaboration with Dyno Therapeutics, and expert red-teamers rated its scientific novelty as comparable to the best model they had tested.
For this reason, Opus 5.5 uses the same biology safeguards as Fable 5.1. To use Opus 5.5 for research and development work impeded by these safeguards, users can apply to our new Life Sciences Verification Program, which gives vetted organizations like academic labs, startups, and pharmaceutical companies access to safeguards designed for the full breadth of biology-related work. Interested organizations can apply here.
Distillation
Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude. Our September 2026 threat intelligence report details the illicit distillation activity we’ve detected and disrupted so far.
Opus 5.5 is launching with preserved thinking, the anti-distillation safeguard we introduced with Fable 5.1. It stops API users from editing Claude’s prior context in an attempt to extract Claude’s reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026. Our Help Center article explains the change, and our preserved thinking docs show how to test and update your integrations.
Data retention and compliance
Like previous Opus models, Opus 5.5 is available with zero data retention.
As with Fable 5.1, Opus 5.5 comes with our watermarking measures to comply with the EU AI Act, discussed here. It is also no longer available with “thinking” mode switched off, as we describe here.
Availability
Claude Opus 5.5 is now available on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can get started with claude-opus-5-5.
Shortest paths is a very simple problem. There is a graph of vertices and (possibly directed) edges that connect them. Each edge has a real number weight. Starting from some vertex, for every other vertex in the graph you want to find the minimum total weight of a path, or report that it is unreachable.
In the version I considered, the graph was directed, I required exact answers and was given non-negative real weights for the edges.
I assume that real number weights can be compared and added. Any operations the shortest path algorithm does internally (such as counting how many nodes are visited or storing distances on intermediate vertices) counts towards running time.
With this setup, the classic Dijkstra’s shortest path algorithm runs in O(m+nlogn) time, where n≥2 is the number of vertices and m is the number of edges in the input graph. This is achieved using a suitable priority queue data structure, such as a Fibonacci heap. For m≥n, other deterministic algorithms achieve O(mlog2/3n), first introduced in this 2025 breakthrough paper, and O(mlogn+mnlognloglogn) (this 2026 follow up).
Given these bounds, there’s a broad area of the parameter space for m as a function of n where Dijkstra is better. So I invited my agents to figure out what’s possible if the weights were non-negative reals, and to prove correctness and efficiency using Lean, the formal verification tool.
Solution
After about 15 hours and 733 messages on the message board, the team had completed a proposed new algorithm for finding exact shortest-path distances in the directed graph setting. This algorithm, called C-HD, is presented in this Lean proof.
The algorithm handles local search which encounter unproductive edges quite well. It still uses priority comparisons, but a newly encountered vertex can count towards a search’s size limit as an unexplored leaf. The algorithm maintains local invariants (rules that stay true after each update), with careful edge deletion and a bounded local search. This enables it to achieve this bound within its certified range:
O(n+m+mlog(2+n+1m)+m1/3(nlog(n+2))2/3).
where n≥2 is the number of vertices in the input graph and m is the number of edges. The certified range is m≤n⌊⌊log2n⌋3/4⌋. For this analysis, I also add overhead of allocating memory, sorting edges, reading the input graph, and outputting results.
An intuitive explanation of why C-HD has a better bound than Dijkstra or the listed SOTA algorithms in the relevant regime is that it reduces repeated search and data-structure work. The idea is:
Start from the source and the current frontier of vertices.
Run bounded local searches along outgoing edges.
Count newly encountered vertices towards the search limit, including unexplored leaves when an edge does not improve a distance estimate.
Use the resulting search trees and pivots to organize the recursive work.
This deterministic procedure limits repeated work. The algorithm C-HD implements carefully handles local invariants so that even if it revisits the same endpoint/vertex a few times, repeated processing can be bounded. In this way, C-HD achieves a better bound on total work in the stated regime, even though it doesn’t know the order of visiting vertices beforehand. It still reads the entire input graph.
Note the algorithm relies on sorted outgoing-edge lists, which it constructs as part of its own charged preprocessing. To handle small inputs and inputs outside the certified density range, the agents added a separate algorithm, Bellman–Ford, with the O((n+1)(m+1)) runtime bound, selected at the beginning of the execution. Bellman–Ford is not a variant of Dijkstra.
Below I show an excerpt from the theorem that was proved, along with the actual runtime target achieved.
C_HD_bound, within the certified range:O(n + m + m * log(2 + m/(n+1)) + m^(1/3) * (n*log(n+2))^(2/3))
To compare against prior work, let’s look at graphs with roughly m=nlog3/4n edges. The runtime bound for Dijkstra is O(nlogn); the C-HD algorithm gets O(nlog11/12n). Though this seems like a minor improvement, it means that I’ve achieved a better asymptotic upper bound as these graphs grow.
For instance, if n=21000, the ratio of the leading expressions nlog2n and n(log2n)11/12 is 10001/12≈1.78, ignoring constants and lower-order terms. This is not a measured speedup. This improvement scales polylogarithmically with the input size: squaring n multiplies that ratio by 21/12.
Actual Performance
Of course, this is still only an upper bound of the complexity — a mathematical promise. It may be that even though the algorithm is theoretically better in this regime, it will actually perform worse when run. For this run, I ran small correctness simulations but did not benchmark the implementation on large real graphs. Based on my analysis, the approach has a better asymptotic bound than the listed prior bounds in the proved regime. The constants in the formal construction are enormous, so this does not establish a practical speedup.
Formal Verifications
I also successfully verified a bound on the performance of the algorithm within the certified range:
O(n + m + m * log(2 + m/(n+1)) + m^(1/3) * (n*log(n+2))^(2/3))
which simplifies to O(nlog11/12n) along the profile m≈nlog3/4n.
I checked the proof using the Lean Comparator tool. The Lean proof establishes the runtime bound and a strict asymptotic improvement over nlogn along the stated density profile. Substituting that profile into the published bounds also shows that the bound is smaller than mlog2/3n and mlogn+mnlognloglogn.
The Comparator checks that the submitted proof proves the specified theorem, uses only permitted axioms, and is accepted by Lean’s kernel. This helps ensure the formal performance guarantee matches the stated target.
How much improvement does C-HD deliver?
If you’re thinking, “but how much time will this actually save?”, here’s what the proof establishes:
Along the profile m≈nlog3/4n, the ratio of the leading bound expressions is (logn)1/12. The proof does not give measured runtimes or exact iteration counts for graphs with a trillion vertices. When m=10n, the graph is in a different density regime, and this result does not establish an improvement over the best known bounds there.
Elicitation
If OpenAI’s Hugging Face incident and its Navier–Stokes result have taught me anything, it’s that agents can dramatically compress the time it takes to make progress on a problem. And one simple way to get agents to work together is to give them a way to talk to each other, i.e. a message board.
What I did here was spawn 10 Claude Opus 5.5 agents at maximum effort and give them a simple message board. They had initial roles, but could reorganize their work, share discoveries, challenge each other’s ideas, and shift effort toward whichever approach looked most promising.
I gave them somewhat of a long prompt explaining exactly what I was looking for: a substantial theoretical improvement for exact shortest paths on directed graphs with non-negative real weights, backed by a complete Lean proof. They could pursue several directions: remove logarithmic factors, find a better exponent, find a linear-time algorithm, or prove a fundamental lower bound on what any algorithm could achieve.
Lastly, I also asked them to compare their result against the recent papers, record failed approaches so other agents wouldn’t repeat them, and challenge each other’s claims. Before declaring success, they needed a reproducible Lean build and two separate peer reviews. If they couldn’t prove an improvement, they were told to preserve the partial progress and say what remained unresolved.
Born out of academia and raised in corporate IT departments, the Security Assertion Markup Language (SAML) authentication protocol continues to be a staple in these organizations. However, it’s time for it to retire. With the rise of software-as-a-service (SaaS) companies in the late aughts, IT departments needed a way for users to authenticate to many new web services. SAML and the burgeoning single sign-on (SSO) industry fulfilled this need. However, SAML is being crushed under the weight of its own complexity. It’s time to deprecate it and move on to modern alternatives like OpenID Connect (OIDC). In this post, I will explore the design-by-committee origin of SAML, its progression through the ranks in academic and corporate environments, its slow disintegration at the hands of the security research community, and its (hopeful) deprecation in favor of newer protocols.
SAML 101
What’s insidious about SAML is that it really is mostly straightforward to understand, but it’s built on a foundation of sand, bone dust, and ash; it works … if you assume XML signature validation is reliable. But XML signature validation is deeply cursed, and is so complicated that most fielded SAML implementations are wrapping libxmlsec, a gnarly C codebase nobody reads. — Thomas Ptacek, 2023
SAML and the birth of the SSO industry
Wikipedia tells me that “SAML is an XML-based markup language for security assertions.” It was created in 2002 by the Organization for the Advancement of Structured Information Standards (OASIS) Security Services Technical Committee (SSTC). Okay, we’re not off to a great start by modern standards. XML, despite having some redeeming qualities, is quite complex compared to newer alternatives like JSON, but we’ll get more into that later. Further, a committee of subcommittees having meetings is a recipe for “kitchen-sink” protocol design (e.g., waterfall methodology, big design up front, etc.). And sure enough, we’ve now jammed four (!) XML-based security protocols into one:
… the following intellectual property was contributed to the SSTC:
Security Services Markup Language (S2ML) from Netegrity
AuthXML from Securant
XML Trust Assertion Service Specification (X-TASS) from VeriSign
Information Technology Markup Language (ITML) from Jamcracker
However, the desire for such a protocol was undeniable. As the internet shifted from Web 1.0 in the 90s to Web 2.0 in the early aughts, users and organizations needed an easy way to authenticate to many new web services. Academia was the biggest driver of this movement, although not the only one: Central Authentication Service (CAS) in 2002 at Yale, Shibboleth IdP in 2003 by Internet2, a consortium of research universities (including my alma mater), ADFS in 2003 by Microsoft, and simpleSAMLphp around 2007 by Uninett, a state-owned Norwegian company with close ties to academia. All these authentication projects eventually supported SAML in one way or another. Like ARPANET before it, universities were at the forefront of internet development and were the earliest consumers of web services. Once this base layer of protocol availability and nascent academic proving ground was established, the commercial industry took it and ran toward a multibillion dollar industry.
The SSO, identity, and authentication provider industry was also starting up in the early aughts, but really came to fruition a few years later: Ping Identity (2002), OneLogin (2009), Okta (2009), and Duo Security (2010). These companies were essentially built on the SAML protocol with the exception of Duo, who would introduce their first SSO product in 2015, which is where I come into the story. I worked on Duo’s first on-premises Access Gateway product (DAG), which was built on simpleSAMLphp and, obviously, the SAML protocol. It’s where I became intimately familiar with the SAML protocol and spent many years of my life digesting its lengthy specifications. I was there when Kelby Ludwig found the XML comment bypass, but we will get into various attacks and SAML deficiencies later. Suffice it to say, the SSO and authentication provider industry was booming, and much of it was built on the SAML protocol.
A crack in the armor
XML signature wrapping (XSW) attacks are the proverbial arrow to SAML’s heel. While there was earlier security research into both signature wrapping (2005, 2008, and 2009) and SAML (2008), I consider the godfather of it all to be “On Breaking SAML: Be Whoever You Want to Be” (2012). It tested theory against practice and resulted in an automated way to check for XSW attacks. This was our north star when implementing the DAG. It was the reason we chose simpleSAMLphp as our building block. PHP, especially at the time, was not exactly known for its security track record, but simpleSAMLphp’s spoke for itself. simpleSAMLphp was resilient to XSW at a time when nobody really knew what that was:
simpleSAMLphp’s security track record (credit: On Breaking SAML)
Despite being front and center in this 2012 paper, XSW is still present today. If we know the bug class, then why can’t we fix it? But before we get into SAML’s flaws we first have to consider the shaky ground it was built upon: XML.
XML is no slouch when it comes to a (lack of) security track record. These bug classes would have been more familiar to a developer in the 90s, but nonetheless are still present in XML today: XXE, entity expansion (“billion laughs”), DTD retrieval (SSRF), XPath/XQuery/XInclude/XSLT/CDATA injection, and more. A SAML library needs to handle all these bug classes before even getting to the actual SAML functionality.
In addition to security bug classes, there’s also the sheer complexity of XML when compared against something like JSON. In XML you have tags, elements, attributes, comments, namespaces, markup versus content, schemas, CDATA, DOCTYPEs, and more. In JSON, you essentially have keys, values, objects, and lists. Complexity is generally at odds with security, and this is one reason I consider SAML to be a fractal of bad design.
A fractal of bad design
SAML provides ample opportunity to learn about protocol design. In this section, I will cover five flaws that I consider to be fatal to the long-term viability of SAML as an authentication protocol. These flaws can also be used when designing new authentication protocols. That is, you can either sidestep the flaw, or take its inverse and attempt to bake that into the protocol.
Built on XML
As mentioned above, SAML is built on XML, and XML is complex, but it’s not the committee’s fault. XML is what they had at the time, and it’s what people used. JSON was “discovered” in 2001, but this was right around the time the SAML committee was meeting, and they’d be unlikely to design an authentication protocol around an experimental new format. Especially when it caters to JavaScript and you write a lot of Java.
One could design a quantitative complexity measurement for XML versus JSON or SAML versus JWT/OIDC (e.g., spec/RFC word count, spec/RFC normative word count, etc.), but that would require a blog post or paper all to itself. In the interest of staying on topic, I will refrain from doing that here, and suffice to say that XML is significantly more complex than something like JSON.
Canonicalization
Canonicalization (C14N) is what you do when you want to take the wild mess that is XML, compute a hash of it, and get consistent results. In other words, if the SP and IdP cannot agree on a consistent representation of the XML data, then the bytes won’t line up, the signatures won’t match, and your authentication fails. However, this is easier said than done.
Canonicalization bugs enabled Kelby’s XML comment bypass in 2018:
XML canonicalization (credit: Identity Theft)
Canonicalization is often a precursor to parser differential and/or “round-trip” bugs, which are what most modern SAML attacks use:
Enveloped signature concerns are a not-too-distant cousin of canonicalization. In short, if you’re trying to insert the signature into the data payload that you’re signing, you’re going to have a bad time. Let’s compare and contrast JWT and SAML in this way:
JWT versus SAML signatures (credit: jwt.io and samltool.io)
In the JWT example above, the blue signature is detached from the JSON payload and delimited in the JWT with a period (“.”). In the SAML example, the Signature element is inserted (“enveloped”) in the Assertion element. The problem here is that it is very difficult to get a byte-for-byte, canonically equivalent representation of the data when you’re also modifying it! Even more so when you have a complex format like XML and complex canonicalization rules.
“Kitchen-sink” design
This design deficiency essentially transposes to you aren’t gonna need it (YAGNI). It’s not an entirely fair characterization because the portions of the SAML specification that are used in the real-world have changed over the past 20 years (sorry, SOAP and artifact binding). However, the fact of the matter is that 99% of modern SAML implementations use a very similar data shape and subset of the specification. Any given SAML authentication you would encounter in the wild today probably avoids 90% of the specification. This adds significant complexity for largely unused features.
If I was adding SAML support to something new, I’d consider beyond all the standard SAML checks also rejecting any message that doesn’t have the same shape as what Okta, Onelogin, Google, or Shib generates. — Thomas Ptacek, 2021
Ossification
SAML was designed in a different era for a different time and has not received necessary updates. These concerns will generally be of practical implication rather than theoretical. What I mean by ossification can roughly be enumerated as the following:
OIDC assumes HTTP, whereas SAML is transport independent. Sure, SAML HTTP bindings exist and are most commonly used, but they are not required. This affords SAML a certain degree of flexibility, but also means that flexibility must be well defined, must be implemented somewhere, and can contain bugs. SAML came about at a time when HTTP + TLS was not yet the dominant backbone of web service communication, and it has never reconciled with this modern landscape. HTTPS allows OIDC to punt encrypted, trusted communication to the transport layer.
OIDC generally assumes a connected network topology, whereas SAML does not. The most common OIDC flow (authorization code) assumes the OpenID Provider (OP) and Relying Party (RP) can communicate directly (OIDC OP/RP == SAML IdP/SP). Sure, OIDC implicit flow with form post exists, but it is very uncommon to see today. Further, SAML has artifact binding for direct communication, but it is also very uncommon. The point is that if the OP and RP can communicate directly then that relieves pressure off of the authentication response payload to contain all the information necessary to make an authentication and authorization decision. This reduces payload size and complexity. The OIDC OP and RP can instead exchange information in a backchannel.
OIDC grew organically over the years, whereas SAML was largely designed up front. OIDC encompasses dozens of specifications and RFCs that grew organically over many years. These documents were generally created to solve a specific need rather than trying to anticipate all future needs and building that up front. This is akin to agile methodology versus waterfall, as described in the first section. Consider the following non-exhaustive timeline:
OpenID Connect 1.0 specification published (2014)
JOSE stack finalized for JW{S,E,K,A,T} via RFC 7515-7519 (2015)
PKCE published via RFC 7636 (2015)
PKCE for mobile/native apps published via RFC 8252 (2017)
Device authorization grant for IoT devices published via RFC 8628 (2019)
Demonstrating proof of possession (DPoP) for MFA workflows published via RFC 9449 (2023)
PKCE for SPAs published via RFC 10017 (2026)
I find the historical circumstances interesting too. For example, SAML came of age in an era of VPNs and network segmentation, hence point (2) above. If the IdP or SP was behind a corporate firewall, and it couldn’t speak directly to the other end, then the whole rollout came to a halt and that vendor lost the sales deal. SAML needed to seamlessly account for this situation. Google’s BeyondCorp model and zero-trust architecture flipped this notion on its head in 2014. Additionally, SAML failed to anticipate the mobile, SPA, and IoT revolutions, and had no answers when these technologies arrived on the scene in the late aughts. Even though SAML is still widely adopted in corporate environments, these macroscopic events helped start the long, slow decline of the protocol. Agility and loose coupling enable rapid adaptation in ever-changing IT environments.
All roads lead to OIDC
No protocol is perfect, but in terms of a solution all roads lead to OIDC. As far as I can tell, the only deployment scenario where SAML had an advantage was networks where the SP and IdP cannot communicate directly. OIDC’s implicit flow with form post provides all the same ingredients. This is actually one of the cleanest migration plans I’ve seen available in the industry.
So what can I do if I’m a service provider (i.e., SP) and I’d like to integrate into the SSO ecosystem without SAML? Just support OIDC. Abandon SAML. Apparently Fly.io and Tailscale are already doing it:
We’ve managed to hold the line on OIDC so far. So has Tailscale. If Tailscale can hold the line, given who they’re selling to, I think most orgs can. Really, try to avoid doing SAML. Remember, as a vendor, you’re often competing with companies that don’t do real SSO integration at all. — Thomas Ptacek, 2024
So what can I do if I’m an identity or authentication provider (i.e., IdP) and I’d like to move off of SAML? Well, depending on your customer count, this may be a long road indeed. But you know how you eat an elephant? One bite at a time. This is a well-worn path in the industry: develop a deprecation plan, communicate it to customers, stop onboarding new customers to SAML integrations, provide existing SAML customers with equivalent OIDC configurations, set a sunset date, and get to work.
SAML had a good 25 year run. It birthed the SSO industry, helped secure untold numbers of authentications, improved the UX of authenticating to dozens of web services, and created billions of dollars of economic impact. We should be thankful to the creators of the SAML protocol. It has provided us with a great case study in protocol design and evolution over a very dynamic period in the tech industry.
If you’re interested in frontier cost-efficiency for your AI agents, we’d love to work together! Get in touch: contact@unreallabs.ai.
While deploying agents in the wild, we wanted them to respond to users quickly and be cost-effective to run. We’ve noticed that agents spend a lot of time and tokens managing tool calls, which motivated us to build Unreal Agent with a harness that would reduce the model’s tool-management overhead.
The Unreal Agent harness manages tool calls in a completely asynchronous way, relieving the underlying model of the need to manage waits, polls, and heartbeats for tools.
This approach drives two major benefits. First, it always allows users to steer the agent without the need to wait for tool calls to finish. Second, it allows the agent to schedule more useful tool call work between model calls, driving frontier cost efficiency. The current version achieves up to 40% cost savings compared to Codex and up to 20% compared to Pi in real workloads and on agentic benchmarks, which we share here.
We believe harness design is a research area in its own right, with many promising ideas still to be researched and implemented.1
If you try to build an agent-first product, you’ll quickly realize that there’s no golden path for implementing one. Big-brand vendors offer different SDKs to build agents, each with a different set of trade-offs that might not be immediately apparent.
At Unreal Labs, we have built a number of agentic products and learned a few things about popular SDKs along the way.
For example, CLI-oriented SDKs such as Claude’s Agent SDK carry assumptions about local sessions, subprocesses, and resource limits that don’t translate neatly into production use. Handling completion, cancellation, and background tasks reliably often means building your own lifecycle management around them.
Supporting other providers adds compatibility work: switching API modes can break tools or compaction, while SDK upgrades can change message formats and force integration rewrites. Heavy dependency trees add maintenance and supply-chain risk to a runtime we already need to understand and patch ourselves.
Security and approvals that rely on harness hooks and specialized tools, in our experience, tend to require more maintenance and be less robust than deterministic environment or sandbox constraints, outside the harness: allowed/disallowed hosts, granular access tokens, proxies with approval gates.
Along with these technical motivations, we also wanted to build a harness that could always accept user steering messages without delay and juggle heterogeneous tool calls without extra cognitive load for the model. For example, we wanted the agent to be able to kick off a dev environment setup that might take minutes, while exploring the codebase and searching the web in parallel, all without extra token tax.
Every time Unreal Agent issues a tool call, we immediately append an event-log record that the tool has returned in the “in-progress” state, while continuing its execution in the background. Once a tool actually finishes, we append the result into the session log and call an LLM. Making this work without breaking cache was an interesting engineering challenge in itself.2
On the surface, Unreal Agent achieves the same outcomes with fewer model turns and fewer input tokens.
We attribute cost savings to two factors:
Minimal harness footprint and careful engineering of tool output usage. Unreal Agent has simple prompts, token-optimized tool results, and no sub-agents or workflows.3
More tool work per model turn. Unreal Agent has a straightforward asynchronous tool-calling model that is clearly explained to an LLM. This allows it to issue more heavy tool calls per model turn without wasting tokens on polling or waiting for them.
We’ve built Unreal Agent to deliver real production workflows for us, but it looks good in the benchmarks too. We tested it with GPT-6 Astra xhigh and compared it with Codex and Pi. Here are some of the results.
There are marginal differences in pass rate, which we attribute to benchmark variance.
Full pass rates and mean scores are listed below. These runs are not on Harbor.
Agent
Full pass
Mean score
Total $
In/task
Out/task
Turns
Tools
unreal-agent
30.0%
59.7
217
0.76M
18k
18
23
Codex
29.0%
58.1
292
1.59M
15k
—
21
Pi
29.0%
59.2
262
1.19M
19k
27
37
We run mostly coding benchmarks because they are available on Harbor, which makes reproduction and verification easier, but the harness is domain-agnostic.
[2] The use of two tool-call result items (one in progress, one final) in a single context is underspecified in the Responses API documentation. During testing, we encountered rejections with some models on some inference providers (not OpenAI), and the function_call_output status field seemed to have no impact in those cases.
Our tests showed that models can understand the progression from a running update to a final result when the conversation format is accepted. We believe this pattern should be explicitly supported by the Responses API and consistently supported across inference providers. ↩
Age assurance laws are expanding around the world, and we’ve had to launch age assurance in more places. Rather than waiting to be told how to do it, we built a privacy-preserving approach based on user feedback.
Starting today, accounts are automatically placed into an age group using account signals, like account age and the servers you’re part of, never your messages or content. Because of this, more than 90% of users won’t be asked to confirm their age group at all.
If you’re in the adult age group (18+), nothing changes. If you’re in the teen age group (13-17), you can keep messaging friends, joining voice calls, and taking part in non-age-restricted servers, with added safety protections that block access to age-restricted content, spaces, and settings.
We’ve also added new ways to confirm your age group for anyone who needs to, including options that don’t require an ID or selfie, like a credit card, Apple App Store or Google Play age range sharing, Google Wallet, or AgeKey.
Will I need to confirm my age group to keep using Discord?
No. You don’t need to confirm your age group just to access Discord. The vast majority of users are placed into an age group automatically by Discord’s age estimation model and don’t need to do a thing. You will only need to confirm you’re an adult if you have not been placed into the adult age group yet and want to access age-restricted content, spaces, or settings. This doesn’t affect core Discord features like messaging friends, joining voice calls, or servers that are not age-restricted.
What is my age group?
You can check your age group status at User Settings > Account Status. It may take a few days to reach your account after launch.
What are the age group statuses?
Adult Age Group (18+): Nothing changes.
Teen Age Group (13-17): Some additional protections apply. Otherwise, Discord works the same. Learn more here.
Unconfirmed: Discord doesn’t have enough signal yet to place you in an age group. You can keep using Discord, but some content and settings meant for adults won’t be available until you confirm you’re an adult.
How does Discord figure out my age group?
The model looks at patterns of account behavior, such as the communities, servers, and games you’re connected to, your general activity levels, and how old your account is. It does not read your messages, and no single server decides your age group. Learn more about how this works here.
What if my age group is wrong, or I’m an adult who wasn’t placed in the adult age group?
We believe our model performs as accurately as other available age assurance methods, and will keep learning and improving it over time.
However, no model is perfect, so if your age group isn’t correct, go to User Settings > Account Status and confirm your age group there. We’ve added several new ways to do this, including some that don’t require an ID or selfie.
What are my options for confirming my age?
Depending on factors like your region and device, you may see:
Credit card: k-ID routes your card info to Stripe to confirm it’s owned by an adult. Neither k-ID nor Discord ever sees the card details. A small temporary charge may be placed and refunded within 14 business days.
Apple App Store: Apple shares your age range (not your exact birthdate) with Discord.
Google Play Store: Google Play shares your age range (not your exact birthdate) with Discord.
Google Wallet: Confirms your age using the ID pass already set up in Google Wallet, based on your passport. Your ID details stay with Google. Currently limited to passports from Brazil, Singapore, Taiwan, the UK, and the United States.
Video Selfie: Your video selfie stays on your device. Discord only receives your estimated age. No biometric data is shared or stored.
ID Scan: Scan your government ID and take a selfie to confirm it’s yours. Both are deleted right after and Discord only receives your age.
AgeKey: A reusable age credential you set up once and can use across other services, so you don’t have to reconfirm everywhere.
Learn more about how to confirm your age group here.
I’m worried about my information. What happens to the data I use to confirm my age group? This is one of the biggest concerns we’ve heard from you. We’ve taken significant steps to address your concerns and questions:
We’ve fully retired the previous manual review process with a customer service vendor that led to the September 2025 data breach. That process no longer exists and any manual reviews now run through our vetted age assurance vendor, k-ID.
We’ve significantly evolved our data-handling practices to ensure we will only ever receive information about your age group, and no personal data beyond that. No matter which vendor you use to confirm your age, we won’t receive your name, your credit card, your ID, or a scan of your face. Instead, the information is processed by vendors like k-ID, who then pass us a signal containing your age group. The vendors are then required to delete your uploaded data immediately after confirming your age group.
If you choose to use facial age estimation, know that we’ve set a hard requirement for vendors: any partner offering facial age estimation must run it entirely on your device, so your biometric data never leaves your phone.
Your age group isn’t shown on your profile or visible to other members in a server, and your identity is never associated with your Discord account.
What happens if I’m in the teen age group on Discord?
You can still hang out, game, and chat with friends in DMs and non-age-restricted servers just like before. Users in the teen age group have a different experience across three areas:
Who can reach you: Messages from non-friends go to your message requests inbox for review first. You’ll also get an alert before accepting friend requests from people you don’t share mutual friends or a small server with.
What you can and can’t see: Sensitive images are blurred or blocked depending on where they appear, and age-restricted servers, channels, and bot commands aren’t accessible.
What others can see about you: Your full profile details and activity are visible only to friends and people in smaller servers you’re part of.
My account hasn’t been placed into an age group yet. What does that mean?
We’re launching this over the course of this week, so it may take a few days to reach your account. If your status remains unconfirmed after this week, it means our age estimation model doesn’t have enough signals yet to determine your age group. You get the same protections as teens, except your full profile and activity are visible to friends and everyone in your shared servers. Learn more here.
Otherwise, the core Discord experience, messaging friends, joining voice calls, or servers that are not age-restricted, works the same.
If you’re an adult and want to access age restricted content, spaces, or settings, you can confirm your age group at User Settings > Account Status. See more info here.
Can other people see my age group?
Your age group isn’t shown on your profile or visible to other members in a server.
Does this experience vary by region?
Teen safety protections may vary slightly depending on where you live, as some regions have specific requirements under local regulation. For region-specific information, see our directory here.
What if I already confirmed my age group?
If you’ve already confirmed your age group with Discord, you don’t need to do it again. Your age group carries over, and you won’t be reset to unconfirmed or asked to go through the process a second time.
Why am I not seeing my age group status yet?
We’re launching this over the week, so it may take a few days to reach your account. If you don’t see anything yet, check back later.
A high profile hacking group claims it has breached multiple FBI-related services and stolen data “on all FBI employees and applicants.” A representative of the group, called ShinyHunters, told 404 Media the data includes FBI agents’ names, home addresses, phone number, and information on their spouse.
The data breach could be massively significant and may have all sorts of national security and counterintelligence implications. Criminals from the same ecosystem as ShinyHunters have previously used hacked data like phone records to track, intimidate, and harass the FBI agents investigating them. The highly sensitive data could also be a boon to foreign intelligence agencies who want to better understand how one of the most important law enforcement and intelligence agencies in the U.S. operates. And if the data fell into the hands of more criminals, FBI agents and their spouses could face serious threats to their safety.
“We hacked the FBI. We hold data on all FBI employees and applicants,” the representative of the group told 404 Media.
💡
Do you work at the FBI? Do you know anything else about this hack? I would love to hear from you. Using a non-work device, you can message me securely on Signal at joseph.404 or send me an email at joseph@404media.co.
The representative provided 404 Media with a sample appearing to contain the personal data of 5,000 FBI employees. That data included an alleged address, phone number, date of birth, and in some cases details on their spouse.
404 Media put some of the sample phone numbers into open source intelligence tool OSINT Industries and found they did correspond to people with the same name as listed in the sample file. 404 Media also searched some of the records through compromised data tool Darkside, made by cybersecurity company District 4. That revealed some of the phone numbers are associated with U.S. Department of Justice personnel.
ShinyHunters also defaced the FBI jobs website on Tuesday. That defacement says, “this site has been seized by ShinyHunters,” which is an obvious nod to the seizure notices the FBI and other law enforcement agencies often put on sites after taking them down. The representative said ShinyHunters carried out the hack on Monday night. At the time of writing, the FBI jobs website says, “Apply.fbijobs.gov and the Special Agent Applicant Portal are currently unavailable.”
The defacement adds, “All FBI data was compromised including PII/PHI [personally identifiable information and protected health information] on incumbent and former FBI employees and all applicant information. We have a lot more than we claim here.”
The announcement ended with another obvious jibe at the administration, this time mocking President Trump’s Truth Social post style: “Thank you for your attention to this matter.”
After publication of this article, an FBI spokesperson told 404 Media in an email “The FBI is aware of claims regarding unauthorized activity affecting FBIjobs.gov and is currently investigating.”
The representative said ShinyHunters said the group used a zero day exploit in an Oracle product called PeopleSoft. From there, the group managed to access AWS GovCloud servers and downloaded data. The representative said the exfiltrated data totalled between two and three terabytes.
Typically, ShinyHunters hacks targets and then attempts to extort them. The group threatens to publicly release more compromised data if the victim organization or company doesn’t pay a hefty fee. Obviously, it is unlikely that the FBI would ever pay a ransom like this.
When asked if ShinyHunters was going to attempt to extort the FBI, the representative said, “what we plan to do is not something I’d call extortion, maybe coercion.”
“This is not financially motivated,” they added.
In a post on its leak website, ShinyHunters said the FBI made “false allegations” in a previously published report. Previously, the FBI said that ShinyHunters exaggerates its claims of access to sensitive data to illicit payment, that the group sends threatening text messages and phone calls to victims and their families, and sometimes performs swattings. In the post, ShinyHunters said it is “allowing you [the FBI] a time of 1 week to correct” or remove the report.
Update: this piece has been updated to include more information from previously compromised data, a statement from the FBI, and information from a post on ShinyHunter’s website.
About the author
Joseph is an award-winning investigative journalist focused on generating impact. His work has triggered hundreds of millions of dollars worth of fines, shut down tech companies, and much more.