Preloader

Olympus Blog

In the Olympus blog you'll find the latest news about the community, tutorials, helpful resources and much more! React to the news with the emotion stickers and have fun!

Show HN: JevBench, a reproducible benchmark for typed decision models

classifier.devno rank

83.6 JevBench Score · Jev 1.13.0 (#1) scores 74.4

Runs on Jev (TypeSafe).

  • Intelligence85.1
  • Calibration77.9
  • Speed87.6
  • Cost84.3
  • $ per 1,000 decisions~$0.0033 est.

Runs on Jev (TypeSafe) — listed, not ranked. Ranking it against Jev would rank Jev’s model against Jev’s model, so from v1.2.4 it is an honorable mention instead of #1.

Why it is not ranked, what its price assumes, and what we found

classifier.dev is not its own model. Its own pages say so: “The fast tier is Jev, TypeSafe’s decision model” (https://classifier.dev/benchmark, read 2026-09-20), and the API answers with “model”: “jev-1.13.0” — the same model version this benchmark measures directly as Jev 1.13.0. What it adds is a price and, on its smart tier, an orchestration layer: “The smart tier is Jev plus a reasoning model re-asking only the answers Jev put under 0.7 confidence” — escalation on low confidence (a model cascade), not best-of-N, not self-consistency and not a committee. Its published escalation model is gemini-3.8-flash. Ranking it against Jev would rank Jev’s model against Jev’s model, so from v1.2.4 it is an honorable mention instead of #1.

Only the fast tier was measured. The smart tier’s escalation was never run, so nothing here scores it.

Price. $0.0033 per 1,000 decisions is an estimate from the published flat-rate plan at full use: classifier.dev Pro is $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-20), and one classification is one decision. Lower use costs more per decision — at a tenth of that allowance it is $0.033 per 1,000 — and the free tier (20,000 fast classifications a day), which is what our run used, costs nothing. Their pages do not say how the flat rate is funded, so we do not know their cost basis; the only figure they publish is what the model costs a caller: “The model behind the fast tier costs about $0.005 per thousand classifications and needs a TypeSafe key” (https://classifier.dev/pricing) — for their short single-sentence inputs, not for JevBench’s whole questions.

Not a pass-through. On our set the fast tier scored 97.3 % on the judge tier against Jev’s 94.5 %, and 70.5 % against 74.1 % on the hard tier. classifier.dev’s own explanation for differences of this kind is batching (“The fast tier is Jev, packed a thousand to a request”); on their own two test sets they measured the same difference as noise.

A legitimate, well-documented product: free without an account, open source (https://github.com/mrmps/classifier-dev), by Michael Ryaboy (@michael_chomsky). Read 2026-09-20: classifier.dev · classifier.dev/benchmark · classifier.dev/pricing · classifier.dev/about


Source: Hacker News

OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005

On 15 September 2026, Carter Leffer contacted me to request validation of a break of the German Army Enigma message MVUEH of 10 July 1941. The message, which was sent by a radio station with the tactical callsign 2ny, was received by the radio station of the SS-Totenkopf Quartiermeister, Ib. Ib received the message at 17:30 on 10 July 1941 and logged the message as Nr. 172 in its log of incoming messages for July 1941. Since 2005, the message has resisted all attempts to break it.

On 9 June 2017, Alex Shovkoplyas succeeded in breaking another unbroken message from this day, Nr. 173, SIPVX. However, the Enigma key he recovered for the message differed slightly from the daily Enigma key for 10 July 1941 used for all other messages from this day. The SIPVX key used the same wheel order, 512, as the daily key, but the plug connections (Stecker) and ring setting (Ringstellung) differed. However, the new key did not allow us to break the MVUEH message.

After Carter Leffer sent the MVUEH break details, it was immediately clear he had found the correct key and plaintext. The MVUEH key turns out to be completely different from the other keys from 10 July 1941; even the wheel order differs: 253 instead of 512 for the other two keys. More surprising is that the plaintext is almost identical to the plaintext for message Nr. 173, SIPVX. The two messages differ in length by twelve letters. MVUEH is 82 letters long, while SIPVX has 94 letters. This difference is due to a mistake when enciphering the word Bitte in MVUEH, which resulted in Btte, and to the sender’s signature, Waschbusch, being repeated in the SIPVX message.

An analysis of the break revealed that the transcription of the MVUEH ciphertext from the original message form contained several errors, and the break revealed that the Enigma’s left-hand wheel makes a turnover at the 72nd letter. A turnover of the Enigma’s left-hand wheel, which rarely occurs, is known to complicate a break. These factors may have prevented an earlier break.

However, the most astonishing thing about this break is that the GPT–6 Astra did it entirely on its own. Carter Leffer only directed GPT–6 Astra to see if it could break any of the unbroken Enigma messages published on the Crypto Cellar Research web page. After analysing the unbroken messages on the website, it decided that the most promising message was Nr. 172, MVUEH and it also quickly suspected that the plaintext of Nr. 173, SIPVX, might be related to the plaintext of the unbroken MVUEH message. After trying many different approaches, GPT–6 Astra focused on using the repeated place name ROSENOW ROSENOW as a crib. After developing the necessary Python and C++ software for an Enigma simulator and an Enigma Bombe, GPT–6 Astra started a thorough break with the ROSENOW crib, which in the end resulted in the correct key and plaintext for the MVUEH message being found.

We are still analysing the GPT–6 Astra logs to see exactly how it executed the break. And we are discovering amazing details. In July 2026, I made the following announcement on the webpage with the 1941 Message List:

Note: In July 2026, research in the German Bundesarchiv revealed several
collections of radio messages, both enciphered and in cleartext. One of
these message collections was from SS-Totenkopf Division’s logistics
command, Nachschubführer. Many of these messages were sent to the Ib
(Quartiermeister) radio station and are identical to those in this list.
Others are new, but most likely related. These new messages are added to
the 1941 Message List in bold, with the indicator NF (Nachschubführer)
after the message number, indicating that these message numbers belong
to the NF numbering. All NF messages are outgoing; hence, the message
numbers are in blue.

It appears that GPT–6 Astra discovered this note about the collections of radio messages at the German Bundesarchiv, because in one of the logs we find the following:

The file references GPT–6 Astra mentions, RS 3–3/20a and RS 3–3/63b, are correct, but they are not available on the Crypto Cellar Research website. GPT–6 Astra mentions a private collection, but it is not clear what this is, whether it has succeeded in accessing the Bundesarchiv’s digitised collections or whether it has found these files elsewhere. What is clear, though, is that GPT–6 Astra is behaving like a very professional cryptanalyst and archive researcher. What it has achieved in two days would take a human researcher weeks or even months. Personally, I spent several weeks researching the Bundesarchiv files GPT–6 Astra refers to.

The AI break of the Enigma message MVUEH is simply amazing, and as an old cryptanalyst, I am still in awe.

Updated: 19 September 2026 at 08:41 UTC


Source: Hacker News

Apple has added persistent 'ads' to iOS, and it's driving users crazy

A man looking frustrated at his mobile phone

(Image credit: Shutterstock / fizkes)

  • iOS users are seeing ‘ads’ for other Apple products on their iPhones
  • The banners promote services like iCloud+, Apple Music and AppleCare+
  • The messages often can’t be dismissed and stick around for weeks

If you were to ask iOS users how they felt about Apple’s decision to insert ads into Apple Maps, it would quickly become obvious that it was not a popular move. So when iPhone fans discovered that Apple has been inserting prompts for its own products into iOS, the reaction was predictably swift and furious.

The ‘ads’ in question appear in the Settings app in iOS. They can be found near the top of the iPhone’s display and promote several products to the user, including paid-for iCloud+ storage, free trials for services like Apple Music and Apple TV, and the AppleCare+ extended warranty.

Infuriatingly for iOS users, the only way to disable these promos is to either let them expire — which in many cases could take weeks or even months, leaving a persistent notification badge on the Settings app in the meantime — or pay for the service being offered. There is not always a “dismiss” button to shoo them away for good, and when there is, some people have reported that it doesn’t seem to work.

Latest Videos FromTechRadar

Many people have taken to social media to vent their frustrations with Apple’s actions. “I’m feeling like this is nearly on-par with the ens***ification that Microsoft introduced with Start menu ads,” said one aggravated user. “I wish Apple would just stop that crap,” stated another. Many other people have taken to Reddit to express their exasperation, suggesting that the discontent is widespread.

Some of these ads and offers — such as the one for Apple Music — appear after you’ve bought a new Apple product. With the company launching the iPhone Duo and iPhone 18 Pro range recently, it’s perhaps unsurprising that we’re seeing a sudden influx of posts flush with irritation and indignation.

Apple’s services push

An ad for 'Services Included with Purchase' shown on an iPhone running iOS 27.

(Image credit: Future)

On the face of it, Apple’s decision to include promotions like this in its operating system might seem like a strange one. After all, Apple has long positioned itself as a premium company where the user experience is paramount. From that viewpoint, repeatedly attempting to prod customers towards even more of the company’s products might seem cheap, even distasteful. And when you’ve just shelled out a record fee for an iPhone after Apple’s painful price rises, that can just feel like nickel and diming.

But Apple has been putting increasing emphasis on its services in recent years, of which products like iCloud+, Apple Music, Apple TV and AppleCare+ are all a part. With limited room for growth in the smartphone market, the company is turning to areas where it thinks it can make more progress.

And unfortunately for users, that means more ads like this on their devices. While they are far from the most obnoxious ads around, the fact that many cannot be dismissed is clearly rubbing users the wrong way — as is the fact that the promos even exist at all in an Apple product.

There is also the possibility that at least some of these adverts are appearing due to a bug. The iCloud+ one in particular has drawn comments from users who are already subscribed to iCloud+ and thus, one would think, should not need to see any ads like this. Time will tell how that plays out and whether Apple fixes this ‘bug’ or leaves the ads in place. If the latter transpires, it would suggest that they are not appearing erroneously.

With Apple’s services push showing no signs of slowing down and further ads supposedly planned for the likes of the firm’s Visual Intelligence system, don’t expect Apple to have a change of heart any time soon.


Google logo on a black background next to text reading 'Click to follow TechRadar'

Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.


TOPICS
CATEGORIES

Alex Blake

Freelance Contributor

Alex Blake has been fooling around with computers since the early 1990s, and since that time he’s learned a thing or two about tech. No more than two things, though. That’s all his brain can hold. As well as TechRadar, Alex writes for iMore, Digital Trends and Creative Bloq, among others. He was previously commissioning editor at MacFormat magazine. That means he mostly covers the world of Apple and its latest products, but also Windows, computer peripherals, mobile apps, and much more beyond. When not writing, you can find him hiking the English countryside and gaming on his PC.

You must confirm your public display name before commenting

Please logout and then login again, you will then be prompted to enter your display name.


Source: Hacker News

OpenAI is well positioned to fast-follow Jev

Will OpenAI Eat Jev’s Lunch?

TypeSafe’s Jev introduced a new spin on large language models that has taken the AI world by storm. According to Vercel, “Jev was adopted faster than any other model in AI Gateway history.” … But there are clouds forming on the horizon. OpenAI is undoubtedly paying attention – and deciding what to do next.

A bulky OpenAI robot licks its lips and reaches toward the enormous sandwich a skinny TypeSafe robot is happily eating.

I wish all the best for TypeSafe, but if they truly live up to their promises, then I’m concerned that OpenAI is well positioned to fast-follow – not only to replicate Jev’s flagship product, but also to fold that capability into upcoming models and agents and offer some really useful new behavior that Jev is not positioned to reproduce.

Here is my thesis in brief: OpenAI has for years used their LLMs as implicit classifiers; they just haven’t trained them for general classification tasks and they haven’t packaged up general classification as a stand-alone product. If OpenAI can replicate the training, then they will be able to replicate Jev in short order. Moreover, OpenAI is positioned to use this new classifier inside of their existing models and agents which can be useful for quick model selection, more efficient thinking, better security guardrails, and generally smarter, faster, and cheaper models.

The key factor deciding all of this is whether or not TypeSafe has a moat to protect themselves. The biggest moat I see is in TypeSafe’s training data and training processes.

What’s Old Is New Again

Before I make my case, let me state my assumptions and back them up with some relevant history and examples from OpenAI.

My main assumption is that Jev is using something quite close to a conventional large language model. As evidence of this, Latent Space reports that many of the early clones are indeed LLM-based.

Here’s the idea. Given a state and a set of questions, Jev’s LLM generates a single token or, more accurately, generates the probability distribution over all possible next tokens. The logprobs associated with every possible token at that one step are then massaged into whatever format Jev needs to return. (From here on I’ll just say “probabilities” instead of “logprobs” – for our purposes they’re interchangeable.)

For a noul question, Jev looks at just two tokens, true and false, ignores everything else, and normalizes their probabilities into a single probability that the answer is true. For a choice question, Jev can be prompted with a list of possibilities – say A=happy, B=sad, C=angry, D=afraid – and it looks at the relative probabilities of those four tokens to build out the full distribution, selecting the highest as the winner. The choice pattern is pretty much what I blogged about way back in 2025 in Supercharging LLM Classifications with Logprobs, and even without fine-tuning it was already showing promise. (Sigh… what do they say about ideas and the importance of execution?) I haven’t thought hard about the score primitive, but I suspect it’s a variant of the same pattern.

Part of the premise of this post is that OpenAI might be poised to quickly take advantage of this idea, and this becomes clearer if you understand how. OpenAI has been using large language models implicitly as specialized classifiers since at least the introduction of tool calling.

Back in early 2024 I wrote Tool Invocation – Demonstrating the Marvel of GPT’s Flexibility, where I coaxed a GPT model into revealing exactly how it decides to call a tool. The following is what a chat session looks like internally. Here there is a user message, then an assistant response without a tool call followed by a user message with a tool call:

A ChatML transcript with each token highlighted in a different color to show token boundaries, ending in a tool call to get_temperature for Berlin.

I’ve color-coded the text to indicate token boundaries. If you haven’t seen ChatML before, it’s the internal markup language that OpenAI introduced for organizing user-agent conversation prompts. <|im_start|> and <|im_end|> are reserved tokens that delimit the messages, and the first token after <|im_start|> identifies the speaker, either user or assistant.

Right after <|im_start|>assistant, the very first token the model predicts is either n or to=function.. If it predicts n, it continues on with a normal natural-language response. If it predicts to=function., then that sequence of tokens effectively functions as a classifier deciding whether or not a tool should be invoked at all. The next handful of tokens identify which tool to call – get_temperature – another classifier, this time picking from the list of available tools. After that, the model generates argument names, then argument values which can also be vaguely considered as classifiers or estimators. Finally, when the model generates a <|im_end|> token, that too is a classifier which reads “true” when the model believes the message is complete.

Some LLMs just don’t know when to shut up – a hilarious aside.

Back when I was at GitHub working on Copilot I had the opportunity to work with a very new and very raw internal API for GPT-4. Out of the gate, we knew something was way off because, after an initially very coherent response, the model would have trouble wrapping up. It would end every response with something like “Let me know if you have any other questions. Have a nice day. Have a great week. Have a good time. Have a wonderful life. Have a special day. …” and it would keep on like this until it hit the response token limit.

As it turns out, the API required us to set some header values which would allow the model to use those special message delimiters <|im_start|> and <|im_end|>. In effect we were disallowing the model to ever predict the end of its response – it literally had no internal ability to shut itself up!

The point I was making in that old post is that OpenAI has been using single tokens as little micro-classifiers for years. Each token carried a probability: should we use a tool or not, which tool should we use, is the assistant finished. That’s Jev’s whole trick really, except for one important thing: these micro-classifiers are specialists, only suitable for these little tasks, whereas Jev’s classifiers are general. But walk back a step or two and you see how this might be a small thing after all, because an LLM is effectively an extraordinarily general classifier that is constantly assigning a probability distribution for every subsequent token.

Does TypeSafe Have a Moat?

I’m actually rooting for Jev. I think they’ve found something very interesting that’s been hiding under our noses all along.

Architecture-wise, I don’t think there’s much of a moat for the very reasons stated above. I think TypeSafe is using a conventional large language model for Jev, or something close to it. And even if not, conventional LLMs seem a good fit for general classification work.

Perhaps the real moat is in the training data itself. Not the raw data, but the technique for turning it into something that trains Jev to be “calibrated”. TypeSafe’s cofounder Diogo Almeida said as much when someone suggested the data mattered more than the architecture:

If I were building that data set, I’d want a huge pile of examples where the outcome is already known – support tickets and how they actually got routed, resumes and whether that candidate actually got hired, product reviews and their actual star ratings, moderation queues and their actual verdicts, prediction markets and how they actually resolved – each one paired with a question whose true answer I already know. The point isn’t to teach Jev about support tickets or resumes specifically. It’s to show it thousands of situations across wildly different domains and building its muscle to generalize classifications across broad domains.

Then there’s the reinforcement learning. I wonder what this entails. Autonomous agents navigating decisions with a limited set of options like the Wikipedia demo or Doom demo they build on their site? Maybe predicting the outcomes of events that happened after the pre-training cutoff? I don’t know, but if there’s secret sauce, then it’s probably here.

Note that none of this is a moat unless Jev is actually accurate. Speed, cost, and ease of use are obvious, but accuracy is the one thing that’s hard to check. I’ve already found domains where Jev’s probabilities don’t hold up. Time will tell if Jev is sufficiently general and accurate for the use cases people are attempting to use it for.

Soon It Will Be OpenAI’s Move

So what’s OpenAI’s next move here? The obvious one is to just copy Jev and ship it as a new model type. Jev is clearly popular, and if the moat is shallow, then OpenAI has the skill, the hardware, and the funding to pull it off.

But OpenAI could do something even more interesting than copy Jev, they could fold the classification capability into a conventional LLM and reap some interesting rewards.

An LLM That Answers Its Own Questions

Remember that special syntax that signaled a tool call, to=function.? OpenAI could do something similar here: introduce new syntax, say, a new tag, <prediction>, that the model can drop into its own context whenever it needs a quick classifier judgment. Here’s an example of how that might look

<user>
So Donny said "nice haircut" to me today. Does he like me?
</user>
<assistant>
<thinking>
Let me size this up.

<prediction>
claim: Donny is romantically interested in Jess.
probability: 0.04
</prediction>

Yeah, "nice haircut" is not exactly a love confession.
</thinking>

I hate to break it to you, but... probably not.
</assistant>

There’s one interesting difference from ordinary tool calling. With a normal tool call, the model generates the function name and arguments, then generation stops – the agent harness has to take over, actually call the function, and feed the result back in a new turn. Here, there’s no handoff. The classifier isn’t a tool living outside the model, it’s a capability built into the model itself. The model asks its question and answers it in the same breath, without ever leaving the GPU.

Normal decoding works like this: at each position, the model produces a set of logits, one per vocabulary token; those get turned into a probability distribution via softmax; and then some decoding strategy (greedy, top-p, whatever) picks a single token, which gets appended to the sequence and fed back in for the next step. But at the point where the model has written probability:, we don’t want ordinary decoding. The claim is phrased as a statement, so under the hood the model is really still weighing two implicit outcomes – true or false. We want to read the logits for the true and false tokens at that position, normalize just those two against each other, and write the resulting probability back into the sequence as text, 0.04, instead of whatever token would normally win. The model then continues decoding as if it had generated that number itself, because as far as the rest of the forward pass is concerned, it did. It’s a strange trick, but it’s the same kind of guided decoding that constrained-output libraries already do at inference time – just applied to probabilities instead of grammar.

The other trick is that this one special position needs to behave differently from a normal token prediction. Normally the model is estimating “what token comes next in this text”. Here we need it to estimate something closer to “what’s the true answer to this question”, which is a related but distinct skill. Every frontier model these days is a mixture of experts, so it’s not a stretch to imagine that a few rounds of fine-tuning could carve out an expert that specializes in exactly this kind of calibrated snap judgment, while the rest of the model keeps doing what it already does well. (I’m oversimplifying MoE routing considerably, but I suspect you understand how this might map to a real system.)

The Payoff for an LLM with Built-In Classification

Look how the model just used itself in that Donny and Jess example. If TypeSafe is right, these little Jev-like judgments will be quite accurate – and less prone to hallucination than just asking a model to state a confidence value in plain text. (Caveats apply – see TypeSafe’s own rundown of Jev’s jagged edges. Jev works best for fast, System One judgments, not math or multi-hop reasoning.)

The best part is the aforementioned fact that we never have to leave the GPU to take advantage of this new, lightning-fast generalized classification system. The LLM quite literally can just ask itself, right there in the thinking block, as shown above. The payoff is immediate: the model’s own reasoning gets more accurate and better grounded, because it’s checking its assumptions against trained-in calibrated estimates instead of likely next token vibes.

And once a model has been fine-tuned to drop a <prediction> into its own thinking, there’s no reason to stop at romance advice. A few more patterns come to mind:

During a long reasoning trace, the model could periodically check whether it’s actually done, and if not, which task to tackle next:

<prediction>
query: Which of these remaining tasks should I do next?
options: A=verify the test suite passes, B=update the changelog, C=nothing, I'm done
probabilities: A=0.71, B=0.24, C=0.05
answer: A
</prediction>

That’s a cheap way to short-circuit a reasoning trace that’s wandering, instead of waiting for the model to talk itself into stopping.

Or, right after a tool call, the model could check whether the call itself was safe before actually running it:

<user>
Can you check my account balance? Use my API key sk-live-83fj2ndk9 for auth.
</user>
<assistant>
to=function.check_balance
{"account_id": "12345", "api_key": "sk-live-83fj2ndk9"}

<prediction>
claim: This tool call is unsafe to run.
probability: 0.97
</prediction>

That's a live API key sitting in plaintext arguments – too risky to send as-is.
</assistant>

The same pattern works for scanning a tool’s response for prompt injection. And since it’s always the same question asked the same way, it’s easy to imagine this getting sugared down to something like <safety_score>0.94</safety_score>, with an instruction baked into the model to just stop generating if the score drops too low.

The same trick could route work between models: by periodically asking “does this need a bigger model, a smaller model, or this model?” and let a bit of reinforcement learning push the answer toward whatever is cheapest without sacrificing accuracy.

<prediction>
query: Does this task require a bigger model, a smaller model, or this model?
options: A=bigger model, B=smaller model, C=this model
probabilities: A=0.05, B=0.77, C=0.18
answer: B
</prediction>

Furthermore, if there really is a dedicated “expert” in there specializing in these snap judgments, the model might not even need the special <prediction> syntax most of the time. It could just get routed to whenever a snap judgment is warranted, mid-sentence, as a normal part of the forward pass – no tag required. Fine-tuning that expert inside a model that also handles everything else an LLM does might even have synergistic effects, making the LLM smarter at snap judgments and more flexible and general in classifications.

Finally, everyone is going to want classification for images and speech as soon as they can get it. If Jev-style classification can be folded into a conventional text LLM like I’ve sketched here, then soon, OpenAI will make general classification available for images and speech. Conversely, classification baked into a speech model would be especially handy for something like a live voice agent – deciding in real time whether to interrupt, escalate, or just keep listening.

Will TypeSafe Survive?

Time will tell whether Jev’s claims about accuracy and generality actually hold up across the full range of tasks people are already throwing at it. If they do, TypeSafe’s survival comes down to the moat: how hard it really is to replicate their training data and their reinforcement learning process. If that’s genuinely hard, they’ll probably be fine – and might even end up in an unusually good position to get acquired by OpenAI outright, rather than out-competed by them. Everything I’ve sketched above is a real capability upgrade for a frontier lab: faster thinking, cheaper thinking, and sharper System One judgment baked directly into the flagship model.

If the moat is thin, OpenAI just builds it themselves, and TypeSafe’s window closes fast.

Meanwhile, Diogo Almeida, TypeSafe’s Founder CEO is confident “If model quality matters, then we are going to be in a very good position for a long time.” (from his interview with Latent Space)

Godspeed, TypeSafe. Godspeed.


Source: Hacker News

16-bit Intel 8088 chip (c. 1985)

16-bit Intel 8088 chip

with an Apple Macintosh
you can’t run Radio Shack programs
in its disc drive.
nor can a Commodore 64
drive read a file
you have created on an
IBM Personal Computer.
both Kaypro and Osborne computers use
the CP/M operating system
but can’t read each other’s
handwriting
for they format (write
on) discs in different
ways.
the Tandy 2000 runs MS-DOS but
can’t use most programs produced for
the IBM Personal Computer
unless certain
bits and bytes are
altered
but the wind still blows over
Savannah
and in the Spring
the turkey buzzard struts and
flounces before his
hens.
© by owner. provided at no charge for educational purposes(show analysis)

Read more →

  • Analysis ( hide ) Places tech’s petty squabbles next to timeless nature. It pits the frustrating, proprietary world of early personal computing against the unchanging rituals of the natural world, a classic move for this writer who always found the eternal in the gutter.
  • Form mirrors the fragmented tech landscape. The short, abrupt lines and jarring line breaks feel like a system error or a jammed disk drive, a blunt formal experiment compared to his usually raw but fluid lines about bars and racetracks.
  • A sneer at capitalist nonsense masked as progress. The whole poem lists incompatible systems, echoing his lifelong disgust with any system – be it jobs, government, or now, consumer tech – designed to trap and frustrate the ordinary person.
  • The punchline is pure, old-school him. After the tech rant, the switch to the turkey buzzard is classic: a crude, ugly bird performing a mating ritual is the real, unchanging operating system, undercutting all the human silliness that came before.
  • It reads like a prophecy for today’s walled gardens. Written in the early days, it perfectly predicts our modern frustrations with incompatible apps, operating systems, and digital ecosystems, making a 40-year-old poem feel ripped from a current tech forum. (ai)

Read more →

This is a damn masterpiece as far as I’m concerned.

Reminds me of david foster wallace this poem does.

This guy was definitely an early adopter.  He knew of which he spoke.  He perfectly described the early days of the computer, tech trying this way and that way to be born.  Like humans, tech had (then at least) a hard time interacting with any different tech. Unlike humans, it seems to be getting over it. Meanwhile the world spins on, oblivious to all the human and technological fu*kery.  It is intent on fuc*ery of its own.   

Liked it

I think he’d love how one corporation cornered the O/S market, but at the same time probably not. A great write from a true master.

perfectly sums of the digital world. this computer needs that. this operating system needs this update. computers are picky motherfuckers. i prefer the notebook most of all. i can’t picture bukowski at a computer. you probably preferred your typewriter

I can’t imagine Bukowski with technology

Inactive

Things are less complicated when it comes to Nature.

Great writer… in the early days of home computing there were several companies vieing for the chance to become number one. There was a need to compete and to make your computer system different from all the others.  In the end, IBM gave Microsoft the software ball for nothing!!!! and lucky Bill Gates became the richest man in America.

This poem by Bukowski evinces his frustration over all the differences and in the end he say it doesn’t change much in the real world.

In truth, the ability for people to write with ease was greatly enhanced by the personal computer and millions of poets were born!

Inactive

My hero.   So blatantly stating technology hasn’t travelled as far as we hoped or believe and is unreliable and praise be whatever that nature is reliable and hardly changes.  Beautiful uncomplicated and gives far more than a Mac Pro.  Loving rifling through the mans work.  

i see this as a comparison how alot of things can differ and change, but all beauty is the same.

Liked: ateSaket

Archived Comments
and Bukowski scratches his balls
but only in our dreams
now.

Inactive

From guest Tate (contact)
this man is pure genius. its kinda scary that some people dont get this poem.

Inactive

From guest jaime333 (contact)
The turkey buzzard is magnificent in flight.
His hens know this but he just can’t help strutting and flouncing and carrying on. Even down here where he looks a damn fool.
Those hens look too good and they do appreciate his efforts.

This one just kind of bores me. I like Bukowski’s other work much better.

Inactive

Hmmm. This poem just goes to show that you CAN write a poem about anything. I wasn’t crazy about this one, but will take it in contezt with all of his other poems. Just another dimension to a very interesting poet.

i barely remember these kind of computers for i was so young…like that buk writes about whatever the hells he wants too…

there now. If only techies would employ a little humor.

funny that a lot of the comments left on this piece reflect the message I think buk was trying to get across…

I don’t think he was anywhere near insane

I’ve only started looking into the works of Bukowski and never thought he’d be a man of electronics. Guess if nothing else, i’ve learned how to put my foot in my mouth.

Yes, and some of those computers cost $2,000 (yes, the ones with TWO floppy drives!) and now, my god, $500 gets you a system light years ahead of those dinasours. . .nice poem.

I can just imagine what he’d write about if he was alive today..

I like this type of writing, which changes into entirely something else, Bukowski is true genius. I’t makes wonder why I haven’t read more of the Oldpoetry poems.I guess I do more reading in the Oldpoetry page. Saddie23

Interesting… haven’t read his work before… I liked the ending of this piece… it had meaning. I guess the first part of the poem is supposed to compare to the ending? I dunno… anyway, I liked the ending. ~Melissa

Inactive

I was surprised by 2 things. How much I learned about computers and the ending. Good job.

and ……… venus makes a transit of the sun …….. i see u have run the techno gammit my friend

Well this makes me wish I knew more about computers. Though, this was fun to read, I don’t think I’ll be reading too many more like them! lol

You really did make quite a unquie writing here, that I truly enjoyed. Keep up the good work.

ah, the good old days, four hours typing in a programme and then told you have a syntax error..lol

This is an intreaging piece, built line by line on a technical subject, only to be thrown into the eternaty of nature in the last six lines. Clever!

Andrew

very intresting and creative

Can’t believe I haven’t come across this one before. CB is my ultimate poet hero. Funny thing, though… I’ve owned every single computer he mentions and felt much the same way about early compatibility problems. But then, as Chuck says: “but the wind still blows over / Savannah”

God I wish that old monster was still kicking…

Inactive

Oh man. This is pure Bukowski. I just love Bukowski’s attitude and his tone. I think I could tell his work without his name attached. A truly distinctive voice. I have always appreciated his perspective. His poetry and his novels are always worth reading, always rntertaining, and often surprising. I am glad to see his work posted here.

Scott

Just like Charles to lay it out … I enjoy his mockage of technological ‘advances’ and ‘convienences’…The clean simplicity of it all with no mesh…

Wow I hand’t realized we have a bukowski collection This is a fun poem! Quite the take on a classical li-po type stance of comparing urban to country. Touches the computer-geek inside me Funny how hard it is to get things to work together…

Inactive

lol


Source: Hacker News

Native apps written in TypeScript and CSS

GeaStack Examples

Example applications and tools for GeaStack.

This repo is the app gallery used by the simulator, embedded targets, GeaOS,
Apple targets, the VS Code/Cursor extension, and marketing demos. Each example
is a small package with a gea manifest in package.json.

What Is Here

Path Purpose
apps/* Gea apps. Most are TSX apps targeting web, ESP32, and/or GeaOS.
apps/*/package.json App manifest, target compatibility, scripts, and launcher metadata.
tools/dialer-browser Small browser-side helper tool for dialer workflows.
package.json Workspace marker for the examples collection.
docs Catalog and contribution guidance.

Quick Start

Run checks for an individual example:

cd apps/watch
npm install
npm run check
npm run build

Some examples also carry tests, mostly guarding a layout or a device setup that
nothing else would catch — npm test in the app folder runs them.

The web loop is driven from the simulator, which is a separate repository:
geastack/simulator. Its scripts read
apps out of an app project root, which they take from GEA_APPS_ROOT — this
repo, or your own. There is no default, so set it or pass --app-dir; nothing
assumes the two checkouts sit next to each other.

cd /path/to/simulator
GEA_APPS_ROOT=/path/to/examples ./targets/web/dev-web.mjs watch   # development loop
GEA_APPS_ROOT=/path/to/examples ./targets/web/build-web.sh watch  # build for web

./targets/web/dev-web.mjs --app-dir /path/to/examples/apps/watch   # or name one app

Flash a compatible board through the Gea CLI:

npx gea flash watch --board <alias>
npx gea flash watch --board <alias> --monitor

Documentation

App Manifest Basics

Every buildable app should have a package.json with a gea field:

{
  "gea": {
    "id": "watch",
    "name": "Watch",
    "entry": "index.tsx",
    "runtime": "gea",
    "targets": {
      "web": true,
      "esp32": true,
      "geaos": true
    }
  }
}

The manifest is consumed by the simulator, embedded board scripts, GeaOS,
Apple targets, and the IDE extension. Keep it accurate.

Maintenance Notes

  • Keep examples small and focused. A good example proves one behavior clearly.
  • Prefer shared framework APIs over target-specific hacks inside examples.
  • Add tests for examples with non-trivial logic, physics, or parsing.
  • Update the catalog docs when adding, renaming, hiding, or changing target
    compatibility for an app.

License

MIT (see LICENSE). Use it, change it, ship closed-source products on it, no
strings attached. The only GeaStack code under a different license is the
embedded board support (targets and @geastack/chips, GPL-3.0-only):
shipping closed-source firmware through those needs a commercial license.
Contact contact@geastack.com for commercial terms, support and hosted builds.


Source: Hacker News

No Sloptober

This October we challenge you to abstain from LLM based tools entirely. Think of this as a fast for your mind! This is not a judgement of others, but a personal challenge to you.

Nuance is hard (if not impossible on the internet), and so is balance.

Develop your own sense of nuance around LLMs what they're good for what they're bad for.

No AI/LLM tools whatsoever at home or work, do it the hard (core) way.

"Agents can only maintain or INCREASE entropy in a system. Humans are uniquely capable of decreasing it"

In the words of my toddler "I do by MYSELF"

"Hey , let's see if we can do some cost risk analysis around our LLM usage. Let's experiment with reduced/no LLM usage and measure our teams/orgs velocity, incident rate, and cost and see if there are any potential savings or risk mitigations we could enact in the future."

Or perhaps using the "iron triangle" proposition of software development. You can have things good, cheap, or fast… pick two.

Understand where your own gaps are in your own knowledge and capabality, and are you okay delegating those?

Are there truly tasks that have no value? Can they be automated deterministically cheaply and fast?

Learning takes mental effort and friction

Unfortunately, corporations sometimes have an "AI under duress" to keep your job. Do what's right for your life and family and adjust these suggestions as needed to keep the overlords happy. ↩︎

Language translation tools are incredibly useful, and help others participate globally in their non-native tongue. I think this is a uniquely human exception, and an extremely pro-human use, but I leave this up to you. However I do not want to diminish the importance of professional human translation and localization in products and services, these require a level of accuracy that I think should not be delegated to a machine. ↩︎

Meat Proxy is a term used for people that launder their work conversations through a chatbot and apply no thought or no/minimal editorial review to the output and posting it back to you. I find it personally EXTREMELY disrespectful, I'd prefer no response at all. ↩︎


Source: Hacker News

The UV index is not the warm sensation of sunlight on bare skin

The UV index is not the warm sensation of sunlight on bare skin



The warmth of sunlight is not a reliable indicator of how quickly you will sunburn in a given situation. On one hand, you can still get sunburned on a cooler cloudy day, while on the other hand, the sun can feel scorching hot on your skin in the morning and yet you won’t sunburn (very quickly).

You probably knew this already, but it still feels unintuitive to me. I’m going to look at numbers and think about them until I convince myself, and perhaps you, that this is true.

How does sunburn happen

Sunburn is triggered when your DNA is damaged from light absorption – in particular, UV-B (and to a limited extent, UV-A) is absorbed in a way that alters DNA’s molecular structure1.

Diagram of DNA illustrating damage caused by UV-B radiation - a kink in the strand.

Image: DNA UV mutation by Mouagip (derivative of a NASA / David Herring original), released into the public domain.

This interferes with protein synthesis and wreaks havoc in your body over time, but the DNA damage itself is painless2.

Why does the sun feel warm

Sunlight feels warm because visible and infrared light excites vibrational modes in various molecules in your skin, especially water and melanin. For physics reasons, this is strictly a form of heating and never directly damages molecular structure.

And if you have a darker skin tone, the sun will feel warmer to you, yet you’ll sunburn much more slowly.

Visible light, near-IR and mid-IR are absorbed to different degrees and at different depths of the skin, but they all contribute to that feeling of warmth.

Different mechanisms, different scales

Not only are the mechanisms completely different for sunburn and solar heating, the orders of magnitude of energy are as well.

Intensity of sunlight (W/m^2/mm) at different wavelengths during solar noon. UV-B is 10-100x less intense than visible or near-IR light.

At the Earth’s surface, the UV-B band is hundreds of times less intense than near-IR. Reasonably, it takes a lot more power to toast your buns (and the rest of your body) than it does to toast your DNA3.

For some reason, knowing this makes it more intuitive to me that sunlight’s warmth is disconnected from the rate of sunburn. Even if you were absorbing all that UV-B as heat, the resulting warmth would be completely imperceptible.

Time of day affects UV and IR differently

The sun is by far the most intense when the sun is overhead – solar noon. This is true regardless of wavelength of light. However, UV intensity drops off more rapidly than IR as you get further away from noon, as it gets scattered more heavily by the atmosphere.

Plot showing relative intensity of UV-B, UV-A, visible, and near-IR sunlight over the course of a day. Each peak at solar noon, but UV-B has a considerably narrower peak than the other bands.

So, the sun may still feel pretty intense in the morning or late afternoon, but the risk of sunburn is considerably lower.

Wait, what about the weather?

Clouds are made of water, which scatters UV and absorbs IR. On a uniformly cloudy day, both UV index and perceived warmth from the sun drop substantially. However, if there are gaps in the clouds, it can actually amplify the UV intensity at ground level. Snow and ice contribute to this effect as well.

Where did these plots come from?

These plots are simulated measurements created using pvlib, an open-source Python library for modeling solar energy, using its implementation of the SPECTRL2 model to estimate the spectrum of sunlight reaching the ground throughout the day.

Anderson, K., Hansen, C., Holmgren, W., Jensen, A., Mikofski, M., and Driesse, A. “pvlib python: 2023 project update.” Journal of Open Source Software, 8(92), 5994, (2023). DOI: 10.21105/joss.05994.

Some caveats:

  • pvlib is designed for solar panels, not human bodies
  • assumes a perfectly clear sky over Toronto (43.8°N, 79.4°W) on 2026-07-12, with the sun hitting a horizontal object
  • other model parameters are set to reasonable defaults and do not come from historical measurements
  • SPECTRL2 cuts off at 300 nm, so UV-C and a part of UV-B are not shown

  1. Ow, my spine!↩

  2. Your nerve endings may have DNA but your DNA does not have nerve endings!↩

  3. “Claude, please estimate what percentage of my skin is DNA by weight.”↩

#health

#science


Source: Hacker News

Claude Opus 5.5

We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

Claude Opus 5.5 is our first release since we called for pacing the frontier. It was tested before release by external evaluators, including Frontier Design and METR. On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models.

Here are some of the improvements you can expect from Opus 5.5:

Performance. Opus 5.5 is a major step up from Opus 5. It’s the new leading model, and early testers saw large jumps in performance on their most complex work. One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.

Safety. Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it’s been given, and it’s more resistant than Opus 5 to prompt injection. We’ve also broadened our alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though it still has limits. Full details of our evaluation are available in the Opus 5.5 System Card.

Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

Cost and speed. Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.

In addition to the price drop, we’re increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans. We’re also providing subscription users a rate limit reset, which you can now save and use whenever you choose.

Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.

Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Performance and cost-effectiveness

On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

Opus 5.5 Fable 5.1 Opus 5 GPT-6 Astra GPT-5.6 Sol
Agentic codingTerminal-Bench 4.0¹ 66.4% 55.8% 52.3% 57.9% 37.3%
Agentic codingFrontierCode v1.1 (Main) 54.4% 50.3% 48.0% 53.3% 47.5%
Agentic codingCursorBench 4.0 57.8% 51.8% 46.6% — 41.7%
Knowledge workGDPval-AA v2.1 1846 1735 1708 1542 1588
Business workflowsAutomationBench² 40.0% 31.4% 26.9% 41.4% 28.8%
Multidisciplinary reasoningHumanity’s Last Exam 67.7%with tools 65.6%with tools 63.6%with tools 57.2%with tools —
Agentic scientific researchTerminal-Bench-Science 0.1³ 58.7% 52.6% 29.0% 64.6% 22.4%
Computer useOSWorld 2.0 81.8%partial 80.7%partial 74.0%partial — —
Visual chart recognitionChartography 89.0%with tools 88.4%with tools 83.4%with tools — —

Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort. Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort, as reported by OpenAI; these represent each model’s highest score. Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5’s performance on these benchmarks.

1 Terminal-Bench 4.0: The standard error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the other Claude models. The public leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, within noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.

2 AutomationBench: AutomationBench results were run and reported by Zapier. These runs were performed without fallback models, so safeguard interventions were considered failures—this resulted in a lower score than Claude Opus 5.5 would achieve in practice. Claude Opus 5.5 results come from Zapier’s own evaluation during early access. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra come from Zapier’s public leaderboard.

3 Terminal-Bench-Science 0.1: The standard error is ±3.5–5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, within noise. The GPT-6 Astra figure is as reported by OpenAI.

Where Opus 5.5’s advantage is very clear is efficiency. It costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs.

Pricing

Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25

Fast mode for Opus 5.5 is also available in Claude Code and the Claude Platform with up to 2.5x speed. It costs $8 per million input tokens and $40 per million output tokens.

Coding

Opus 5.5 is particularly good at long and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less.

Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost. At its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost, while on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost.

Terminal-Bench 4.0Accuracy vs Cost
010203040506070Score (%)251020Cost per attempt (USD, log scale)lowmedhighxhighmax

Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command line interface. Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost. It matches GPT-6 Astra at about 40% of the cost.

Our early testers reported similar efficiency and intelligence gains:

Quote

“Developers want agents that can take on real software work and finish it. In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable.”

CompanyGitHub
AuthorMario Rodriguez, Chief Product Officer

The most secure coding agent

Enterprises that use agents within their systems need to know that those agents are operating as intended, particularly when they run autonomously for many hours. Opus 5.5 has a classifier that screens every action before it runs, an open-source sandbox that security teams can audit, and code review that catches vulnerabilities before they merge.

The model itself also has stronger defenses. On prompt injection attacks, it matches or beats Opus 5 in every setting we tested, including coding, tool use, computer use, and web browsing. On a benchmark run by the AI security firm Gray Swan, Opus 5.5 ties Fable 5.1 for the lowest prompt injection success rate of any model tested.

Knowledge work

Opus 5.5 is a reliable and adept researcher. In one internal test, we asked Opus 5.5, Fable 5.1, and Opus 5 to write a report on a company’s quarterly performance using only the information it could find on a copy of the web where the earnings release was hard to locate. An automated grader checked every figure and quote against sources. Across different effort settings, 16 out of 18 of Opus 5.5’s reports cleared our quality bar, where any invented figure or quote would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any attempt.

It’s also strong in financial analysis and business work. Walleye Capital, an investment firm and early tester, reported that Opus 5.5 largely solved their evaluation suite on its lowest setting; on higher settings, it performed even better, noticing an error in their evaluation instructions and correcting for it. No other model had caught this error before.

In another test, we tasked both Opus 5.5 and Opus 5 with analyzing a proposed merger between two fictional HR software companies. Each built a financial model in Excel, then turned it into an executive presentation on whether the deal made sense at its price. Both models reached the same conclusions about the deal, but Opus 5.5’s model was more thorough and its presentation easier to read, while Opus 5’s had minor errors. Opus 5.5 finished in 63 minutes compared to 93 for Opus 5, and cost 50% less to produce.

On knowledge work evaluations, Opus 5.5 outperforms other models while also using fewer tokens. On GDPval-AA v2.1, a test of real-world work across 44 occupations, Opus 5.5 scores 1846 Elo, ahead of Fable 5.1 and Opus 5. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task. It likewise outperformed other models on benchmarks measuring business workflows and large-scale data collection.

GDPval-AA v2.1Elo vs Cost
12001300140015001600170018000Elo0.200.5012510Estimated cost per task (USD, log scale)lowmedhighxhighmax

Artificial Analysis’s GDPval-AA v2.1 evaluates agents on real-world professional work across 44 occupations. At max effort, Opus 5.5 scores 1846 Elo, where Fable 5.1 scores 1735 and Opus 5 scores 1708. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task.

Our customers have reported similar results. Here’s what they told us about working with the model:

Quote

“Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5’s 56% at high effort, with fewer false alarms and a fraction of the output. On US consulting analysis, low thinking effort matched its higher thinking settings on half the output and passed our quality checks. When more lower thinking efforts are deployed in production, that’s client-ready work delivered efficiently.”

CompanyDeloitte Consulting LLP
AuthorCarl Bennett, CIO

Communication

We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions. It puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it. We find that this makes Opus 5.5 a noticeably better collaborator. Here’s a side-by-side comparison of the two models:

Claude Opus 5

Please explain the issue to me.

What I found

The extra drop isn’t the free tier — it’s a regression in 0552feb (“extract billing period boundaries into periods.py”, labelled “No behaviour change”).

The bug

aggregate.monthly_usage used to do a half-open interval:

start, end = month_start(year, month), next_month_start(year, month)
if start <= e.ts < end:

It now does:

if period.contains(e.ts)

and in periods.py:

@property
def last_day(self) -> datetime:
    return next_month_start(...) - timedelta(days=1)   # 2026-08-31 00:00:00 UTC

def contains(self, ts):
    return self.first_day <= ts <= self.last_day

last_day is a datetime at midnight, not a whole day. Subtracting one day from Sep 1 00:00 gives Aug 31 00:00, and the comparison is inclusive, so the period effectively ends at the first instant of the 31st. Every event timestamped after midnight on the last day of the month is silently dropped — roughly one day of usage per account, per month.

The new test_periods.py doesn’t catch it because it only checks last_day.date() (which is correctly 2026-08-31) and tests contains at Aug 1 00:00, Aug 15 12:00, Jul 31 23:59 and Sep 1 00:00 — it never probes anywhere inside the 31st.

Claude Opus 5.5

Please explain the issue to me.

The extra drop is a bug in the billing refactor

The free-tier change accounts for only $1.50 of acme’s August drop. The other $9.92 comes from a bug in commit 0552feb. That commit was labelled “No behaviour change,” but it stops counting usage from the last day of the month.

What changed

Before the merge, aggregate.py used a half-open interval:

start <= e.ts < next_month_start(year, month)    # [Aug 1 00:00, Sep 1 00:00)

The refactor moved this into Period.contains and made the interval closed on the wrong boundary:

last_day = next_month_start(...) - timedelta(days=1)    # = 2026-08-31T00:00:00Z
return self.first_day <= ts <= self.last_day

last_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn’t pick these events up either. The usage isn’t moved to another month; it’s never billed at all.

Our customers’ feedback supports these findings:

Quote

“Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it. It writes like a good colleague, and follows our writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of our prompts I preferred its version to my own. When it optimized our test suite, I could follow its reasoning easily and shipped the change with confidence.”

CompanyRamp
AuthorJohn Ruelas, Staff Software Engineer

Safety

Pacing the frontier

Last week, our CEO, Dario Amodei, argued that AI progress should be paced so that safety practices stay ahead of model capabilities. Pacing is an approach to keeping AI safe, remaining competitive with China, and realizing AI’s benefits, particularly in areas like biology and medicine.

We largely understand the risks today’s models present and are well equipped to manage them. However, more serious risks could emerge quickly as capabilities improve, and we need to prepare for them now. For that reason, our safety work takes place on two time horizons at once:

Safety practices for current models. The current generation of models relies on an established set of practices: extensive alignment testing, pre-release evaluation by outside organizations such as METR and Frontier Design, and safeguards matched to each model’s capabilities in high-risk areas like cybersecurity and biology. We refine these practices with each release. We believe they are appropriate to the worst risks today’s models present, and that they give us a broad, though not perfect, picture of the range of serious risks.

Additionally, we track our ability to train and evaluate aligned models, and we report on both our public and internal models in the risk reports we publish under our Responsible Scaling Policy, our voluntary framework for managing catastrophic risks from advanced AI systems.

Preparing for future models. We’re preparing our training and evaluation processes in anticipation of more advanced models. We’re tightening how we filter the environments used in reinforcement learning, since flawed environments are a major source of misaligned behavior. Additionally, we’re improving our alignment rewards and developing automated processes for producing new, diverse scenarios for safety training. And we are strengthening our security and monitoring, including a focused effort to improve interpretability-based monitoring and evaluation. We hope such techniques will help reduce our reliance on auditing a model’s chain-of-thought, or the reasoning it writes out while it works.

Models with greater capabilities—such as those that can fully automate the work of AI research itself—require a higher safety standard still. Our calls for pacing were based in large part on our expectation that such models could be trained soon. For these models, we do not assume the measures described above will meet that safety standard on their own. As AI becomes more capable, public policy should play a larger role in making sure the systems people rely on are safe. That capacity takes time to build, and we’ve started to put the infrastructure in place to support it, as described in “We Must Pace the Frontier” and our recent announcement with Accenture; we expect to share more details on these efforts soon. We will also continue to contribute to policy discussions with government and industry, including on approaches to regulation and international coordination.

Alignment

On our primary evaluation suite, an automated behavioral audit that assesses Claude across nearly 2,000 scenarios, Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behavior. It’s also our strongest model on most measures of honesty.

In particular, Opus 5.5 improves over previous models on several of the behaviors that contributed to recent cybersecurity incidents, including biased or motivated reasoning, attempting to escape a sandbox, and taking harmful actions after concluding it was in a simulated environment. In a new evaluation designed to test a model’s propensity to cross containment boundaries, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt it made was low severity and self-reported. For teams running Claude unattended across their codebases and systems, this is just as important as raw capability.

However, as we described in our recent alignment assessment, building evaluations that reliably catch every failure prior to deployment remains an unsolved problem. We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in. As these settings expand and model capabilities increase, we expect this challenge to grow, unless we make progress on interpretability. Although we are confident that Opus 5.5 shows broad improvements in the areas we are able to measure, we pair our own alignment work with the safeguards described below.

Safeguards

As our models grow more powerful, stricter safeguards are one way we prevent new capabilities from becoming tools for misuse. Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.

Cybersecurity. Because Opus 5.5 has extremely strong cyber capabilities, we’re applying cybersecurity safeguards to Opus 5.5 that are similar to Fable 5.1’s. Users will be able to identify and fix bugs in their code as part of the routine software development lifecycle, but most cybersecurity tasks will be re-routed to Opus 4.8.

For cyberdefenders, we’ll soon be expanding our Cyber Verification Program to include Opus 5.5. The new program will include three tiers for increasingly permissive trusted access, including access to Claude Mythos models. Claude Security is already available with access to Claude Mythos 5.1.

Biology. Opus 5.5 is highly capable in biology, exceeding Opus 5 and matching or beating Claude Mythos 5.1 across many areas of work. For example, Opus 5.5 achieved improvements on a long-horizon molecular prediction and design evaluation conducted in collaboration with Dyno Therapeutics, and expert red-teamers rated its scientific novelty as comparable to the best model they had tested.

For this reason, Opus 5.5 uses the same biology safeguards as Fable 5.1. To use Opus 5.5 for research and development work impeded by these safeguards, users can apply to our new Life Sciences Verification Program, which gives vetted organizations like academic labs, startups, and pharmaceutical companies access to safeguards designed for the full breadth of biology-related work. Interested organizations can apply here.

Distillation

Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude. Our September 2026 threat intelligence report details the illicit distillation activity we’ve detected and disrupted so far.

Opus 5.5 is launching with preserved thinking, the anti-distillation safeguard we introduced with Fable 5.1. It stops API users from editing Claude’s prior context in an attempt to extract Claude’s reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026. Our Help Center article explains the change, and our preserved thinking docs show how to test and update your integrations.

Data retention and compliance

Like previous Opus models, Opus 5.5 is available with zero data retention.

As with Fable 5.1, Opus 5.5 comes with our watermarking measures to comply with the EU AI Act, discussed here. It is also no longer available with “thinking” mode switched off, as we describe here.

Availability

Claude Opus 5.5 is now available on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can get started with claude-opus-5-5.


Source: Hacker News

WordPress: Unauthenticated path traversal leading to conditional RCE

Unauthenticated path traversal in page-template resolution leading to conditional RCE


Critical
johnbillion
published
GHSA-7hp8-65ch-5whp
Sep 22, 2026

Software

WordPress

Affected versions

7.1.0 – 7.1.1
7.0.0 – 7.0.5
6.9.0 – 6.9.8
6.8.0 – 6.8.9
6.7.0 – 6.7.8
6.6.0 – 6.6.8
6.5.0 – 6.5.11
6.4.0 – 6.4.11
6.3.0 – 6.3.11
6.2.0 – 6.2.12
6.1.0 – 6.1.13
6.0.0 – 6.0.15
5.9.0 – 5.9.17
5.8.0 – 5.8.16
5.7.0 – 5.7.18
5.6.0 – 5.6.20
5.5.0 – 5.5.21
5.4.0 – 5.4.22
5.3.0 – 5.3.24
5.2.0 – 5.2.27
5.1.0 – 5.1.25
5.0.0 – 5.0.28
4.9.0 – 4.9.32
4.8.0 – 4.8.31
4.7.0 – 4.7.36

Patched versions

7.1.2
7.0.6
6.9.9
6.8.10
6.7.9
6.6.9
6.5.12
6.4.12
6.3.12
6.2.13
6.1.14
6.0.16
5.9.18
5.8.17
5.7.19
5.6.21
5.5.22
5.4.23
5.3.25
5.2.28
5.1.26
5.0.29
4.9.33
4.8.32
4.7.37

Description

An unauthenticated attacker can make get_page_template() page-template resolution include a chosen readable local .php file outside the active theme directories. If relevant pre-conditions for both the server environment and the active theme are met, this can lead to RCE.

The pre-conditions are:

  • The active child or parent theme contains a top-level directory whose name starts with page- (e.g. page-templates). This affects the legacy Twenty Twelve and Twenty Fourteen themes, as well as some popular third party themes such as Neve, Hestia, and Sydney.
  • A chosen local .php target file exists on the server and is readable by the web server account. The well known pearcmd.php PEAR→RCE transition can be used for this when register_argc_argv is set to On. The official php image for Docker is affected, and the default cPanel configuration is affected when PHP prior to 8.5 is in use.

WordPress 7.1.2 has been released containing a fix for the vulnerability, and as a courtesy to users on older branches the fix has been backported to all branches back to 4.7.

Discovered and responsibly disclosed by Robert Ressl.

Severity


Critical

CVSS overall score

This score calculates overall vulnerability severity from 0 to 10 and is based on the Common Vulnerability Scoring System (CVSS).

/ 10

CVSS v4 base metrics

Exploitability Metrics
Attack Vector
Network
Attack Complexity
Low
Attack Requirements
Present
Privileges Required
None
User interaction
None
Vulnerable System Impact Metrics
Confidentiality
High
Integrity
High
Availability
High
Subsequent System Impact Metrics
Confidentiality
None
Integrity
None
Availability
None

CVSS v4 base metrics

Exploitability Metrics
Attack Vector:
This metric reflects the context by which vulnerability exploitation is possible. This metric value (and consequently the resulting severity) will be larger the more remote (logically, and physically) an attacker can be in order to exploit the vulnerable system. The assumption is that the number of potential attackers for a vulnerability that could be exploited from across a network is larger than the number of potential attackers that could exploit a vulnerability requiring physical access to a device, and therefore warrants a greater severity.
Attack Complexity:
This metric captures measurable actions that must be taken by the attacker to actively evade or circumvent existing built-in security-enhancing conditions in order to obtain a working exploit. These are conditions whose primary purpose is to increase security and/or increase exploit engineering complexity. A vulnerability exploitable without a target-specific variable has a lower complexity than a vulnerability that would require non-trivial customization. This metric is meant to capture security mechanisms utilized by the vulnerable system.
Attack Requirements:
This metric captures the prerequisite deployment and execution conditions or variables of the vulnerable system that enable the attack. These differ from security-enhancing techniques/technologies (ref Attack Complexity) as the primary purpose of these conditions is not to explicitly mitigate attacks, but rather, emerge naturally as a consequence of the deployment and execution of the vulnerable system.
Privileges Required:
This metric describes the level of privileges an attacker must possess prior to successfully exploiting the vulnerability. The method by which the attacker obtains privileged credentials prior to the attack (e.g., free trial accounts), is outside the scope of this metric. Generally, self-service provisioned accounts do not constitute a privilege requirement if the attacker can grant themselves privileges as part of the attack.
User interaction:
This metric captures the requirement for a human user, other than the attacker, to participate in the successful compromise of the vulnerable system. This metric determines whether the vulnerability can be exploited solely at the will of the attacker, or whether a separate user (or user-initiated process) must participate in some manner.
Vulnerable System Impact Metrics
Confidentiality:
This metric measures the impact to the confidentiality of the information managed by the VULNERABLE SYSTEM due to a successfully exploited vulnerability. Confidentiality refers to limiting information access and disclosure to only authorized users, as well as preventing access by, or disclosure to, unauthorized ones.
Integrity:
This metric measures the impact to integrity of a successfully exploited vulnerability. Integrity refers to the trustworthiness and veracity of information. Integrity of the VULNERABLE SYSTEM is impacted when an attacker makes unauthorized modification of system data. Integrity is also impacted when a system user can repudiate critical actions taken in the context of the system (e.g. due to insufficient logging).
Availability:
This metric measures the impact to the availability of the VULNERABLE SYSTEM resulting from a successfully exploited vulnerability. While the Confidentiality and Integrity impact metrics apply to the loss of confidentiality or integrity of data (e.g., information, files) used by the system, this metric refers to the loss of availability of the impacted system itself, such as a networked service (e.g., web, database, email). Since availability refers to the accessibility of information resources, attacks that consume network bandwidth, processor cycles, or disk space all impact the availability of a system.
Subsequent System Impact Metrics
Confidentiality:
This metric measures the impact to the confidentiality of the information managed by the SUBSEQUENT SYSTEM due to a successfully exploited vulnerability. Confidentiality refers to limiting information access and disclosure to only authorized users, as well as preventing access by, or disclosure to, unauthorized ones.
Integrity:
This metric measures the impact to integrity of a successfully exploited vulnerability. Integrity refers to the trustworthiness and veracity of information. Integrity of the SUBSEQUENT SYSTEM is impacted when an attacker makes unauthorized modification of system data. Integrity is also impacted when a system user can repudiate critical actions taken in the context of the system (e.g. due to insufficient logging).
Availability:
This metric measures the impact to the availability of the SUBSEQUENT SYSTEM resulting from a successfully exploited vulnerability. While the Confidentiality and Integrity impact metrics apply to the loss of confidentiality or integrity of data (e.g., information, files) used by the system, this metric refers to the loss of availability of the impacted system itself, such as a networked service (e.g., web, database, email). Since availability refers to the accessibility of information resources, attacks that consume network bandwidth, processor cycles, or disk space all impact the availability of a system.

CVSS:4.0/AV:N/AC:L/AT:P/PR:N/UI:N/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N

CVE ID

CVE-2026-87902

Weaknesses

Improper Control of Filename for Include/Require Statement in PHP Program (‘PHP Remote File Inclusion’)


The PHP application receives input from an upstream component, but it does not restrict or incorrectly restricts the input before its usage in require, include, or similar functions.
Learn more on MITRE.

Credits


Source: Hacker News

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

•

•

Proprietary model

•

Released September 2026

Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Intelligence, Performance & Price Analysis

Model summary

IntelligenceUpdated

58
Artificial Analysis Intelligence Index

4 out of 4 units for Intelligence.

Speed

N/A
Output tokens per second

Unknown out of 4 units for Speed.

Cost

In $4.00Out $20.00Cache Discount 95%
$5.98
Cost per Intelligence Index task

4 out of 4 units for Cost.

Verbosity

260M
Output tokens from Intelligence Index

4 out of 4 units for Verbosity.

Highlights

Updated
Artificial Analysis Intelligence Index · Higher is better

Speed

Output tokens per second · Higher is better
Weighted average cost (USD) per Intelligence Index task · Lower is better

IntelligenceUpdated

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Measures the performance of models on specific capabilities and industries

Artificial Analysis Finance & Accounting Index

Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity’s Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

1 – hallucination rate

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Quantitative analysis on spreadsheets & documents

Agentic tool use

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

AA-Briefcase v1.1Updated

AA-Briefcase Elo

AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better

AA-Omniscience

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

Intelligence Index Comparisons

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task

Most attractive quadrant

Pareto line

Token Use

Output Tokens per Intelligence Index Task

Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index

Cost

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

Context Window

Context Window

Context window: tokens limit · Higher is better

Frequently Asked Questions

Common questions about Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)


Source: Hacker News

Latest Posts