Preloader

Technology

Saving another 100TB of RAM

Cloudflare operates at a scale so big that even after working here for years, it doesn’t seem real. We have thousands of servers all over the world with petabytes of RAM and millions of CPU cores, and all of it is pushed to the max. As vast as those resources feel, they are still finite, and when you need every service to run on every node, it doesn’t leave room for wasted space.

At this scale, small improvements are greatly magnified, so even 1%-at-a-time improvements are worth celebrating. And some tweaks add up to a lot more: in this post, we’ll look at how small changes to a single algorithm reduced the memory footprint of one of our Pingora-based services significantly. That allowed us to reclaim more than 100TB of RAM globally, on top of the 100TB of memory the DNS team was able to shed last month.

Waste not

Maintaining equitable resource sharing between teams is not easy, especially in large organizations. One of the ways Cloudflare ensures the balance is kept is through the tireless efforts of the wonderful Performance team. 

This story starts with a ticket filed by Ivan who found: Excessive memory usage from pingora-ketama in Pingora Backend Router. The finding was that our internal load-balancing service, Pingora Backend Router (yes, PBR), was using significantly more memory than expected — specifically in structures associated with pingora-ketama, which is our open-source library for handling consistent hashing.

In order to talk about how we addressed this seeming overuse of memory, we need to talk about what consistent hashing even is, why we are using it in PBR, and how it became so memory hungry. Along the way, we’ll learn some Rust and even a little math.

Consistent hashing

Consistent hashing is a widely used method for distributing tasks across multiple servers in a way that does not require large changes when servers are added or removed. Internally we use it to route cacheable requests to servers by URL. This allows us to keep only one copy of a file stored per data center and gives a stable way to find the location of each file. We have mentioned this system before, but let’s take the time to walk through how and why this algorithm is used and how it works.

The key concept of consistent hashing is that while hash functions can accept any kind of input, their output is limited to a single unsigned integer (32, 64, or 128-bit integers depending on which hash function). This allows us to relate tasks and servers to each other in a consistent way. Most discussions of consistent hashing have you think of that output space as a continuous, circular ring that wraps around from its max value to zero. This depiction makes for some nice visualizations, but it can also make the simple concept of integer ranges seem more complicated than it needs to be. For our discussion, we’ll represent the 32-bit output of our hash function as a number line.

BLOG-3083 2.png

Now, let’s say we have a set of servers, A, B, & C, and a set of tasks t-z. We can map each onto the number line based on the hash of their representative values, so something like IP addresses for servers and cache keys for tasks.

BLOG-3083 3.png

Assigning tasks to servers is now just a matter of finding the first server to the left of each task. We can represent this visually by coloring in the region of hashes that will be associated with each server. Notice that the range covered by server C wraps around to the beginning, hence the idea that hashes exist in a ring.

BLOG-3083 4.png

And that’s it. At a base level, consistent hashing is this simple — but it doesn’t take long to see that there is room for improvement. Notice that the range covered by server A in our example is significantly larger than that of either B or C. This is a problem because the fraction of the requests a server handles is going to be proportional to the size of its range on the number line. Ideally we would like to guarantee each server will have an equal size, but because hashes are essentially random numbers, we have to talk about the size of the regions in terms of statistics. 😨

Math and consequences

First: don’t panic. I promise I’m not about to lie to you and that we will stay safely within the bounds of a day-one probability lesson. When we talk about statistical distributions, there are two big factors that help us quantify uncertainty in helpful ways: expected value and standard deviation. In (over-)simplified terms, expected value gives us a point where measurements based on a distribution will be centered, and standard deviation tells how close to that central point most measurements are likely to be.

For consistent hashing, we can calculate these factors for the fractional size of the range associated with one of N servers. (Details on where this formula comes from later).

$$m

begin{align*}
text{Exp} &= frac{1}{N} \
text{SD} &= frac{1}{N}sqrt{frac{N-1}{N+1}}
end{align*}

 m$$

In terms of concrete numbers, let’s say we have 100 servers. The formulas above give:

$$m

    text{Exp}=1/100 = 1% \

    text{SD}= frac{1}{100}sqrt{frac{100-1}{100+1}} approx 0.99%

m$$

That tells us that we can expect that the range each server handles will be centered around 0.99% of the total and most of the lengths to fall within 1% of what’s expected. This sounds good until we realize that that’s 0.99% of the total length. We need to scale the standard deviation by the expected value to see how big the error is as a fraction of the target size. This value is called the coefficient of variation.

$$m

text{CV} = frac{text{SD}}{text{Exp}} = sqrt{frac{N-1}{N+1}}

m$$

At $m N=100, text{CV} approx 99% m$ — meaning some servers will likely be working 99% harder than they should be (handling twice as many requests) while others could be doing practically nothing! Now that we have a way to predict how evenly loaded servers will be using consistent hashing, we can start working on improvements.

What if we add hashes?

The simplicity of consistent hashing is a double-edged sword. It’s easy to understand and implement because everything is turned into easily-relatable hashes on the same numberline, but any improvements to the system will also need to be relatable to that numberline. That means the solution to any consistent hashing problem can only be more hashes. It’s less like a golden hammer (a tool with which all problems look like nails) and more like a golden nail in that it turns all tools into hammers.

To solve the problem of imbalanced workloads, we can add multiple hashes to represent each server instead of just one. We’ll get to the math behind this momentarily, but it should make some intuitive sense that while each individual range has a large standard deviation, adding a bunch together should make their total size even out. If we take our three-server example from the above diagrams and add two more hashes at random for each server, we see that it helps even out each server’s workload. 

BLOG-3083 5.png

This is an admittedly contrived example. The random nature of the system means there’s no guarantee how much improvement you will get from adding 2 additional hashes per server, but it should make some intuitive sense that combining more of these hash segments together produces a more even distribution. Each segment in the sum has a chance of balancing another. Maybe one is too short; maybe one is too long. This is essentially what the law of large numbers tells us should happen… The obvious problem is it only works for large numbers. In NGINX, the baseline number of hashes per server is hardcoded to 160, and Pingora uses the same value as the default. I’ll spare you the math for now, but if we go back to our 100-server example, if we use 160 points per server instead of just one, the coefficient of variation (which we can think of like an error margin) drops from about 99% to about 8%, a significant improvement.

What if we add more hashes?

We saw above that increasing the number of hashes per server by a constant amount allows us to improve how evenly workloads are distributed per server, but what if we don’t want to distribute the work evenly? In Cloudflare’s case, we have some servers that have more storage space than others, so it would be better to have the number of requests allotted to a server be proportional to its disk space. One way to accomplish this is with the ketama algorithm. The naming is a little funny because the algorithm is named after the library where it was first implemented, and the library was named … well you can google it 😶‍🌫️.

The whole algorithm boils down to: For any two servers, $m S_1m$ & $mS_2m$, if we want the requests served by $mS_1m$ to be $mwtimesm$ more than those served by $mS_2m$, the number of hashes associated with $mS_1m$ needs to be $mH_1 = wtimes H_2m$. This allows us to set a “weight” for each server, which scales the number of hashes associated with that server. Unfortunately this is not a replacement for the constant scale factor we added in the section above. That scaling needs to be there to set a minimum error margin, which will show up in the servers with the lowest weights.

For us, since we want workload to be scaled based on storage, we can use the disk space as the weight, which is exactly what the Pingora team has been doing for years. Elsewhere in the company where workloads are more compute-intensive, weights might be based on CPU or GPU count.

What if we add even more hashes???

The last problem we need to address is that so far we are working under the assumption that any server can handle any request, but in practice that is not the case. Things like compliance requirements or enabled caching features mean only a subset of servers can handle any particular request. Unfortunately, unlike before, we can’t solve this problem by adding more hashes to the same ring. We have to add completely new rings, and not only that — every combination of features potentially needs its own specific ring!

Duplication based on combinations is a classic recipe for exponential explosion. In our case, we have a handful of different features leading to $m2^text{handful} = text{dozens}m$ of separate consistent hash rings. So as you have probably guessed by now, the “excessive memory use” (6GB in some cases) that Ivan found was due to an enormous number of hashes to accommodate all the functionality we need and which have to be stored in memory. So what can we do?

Storage improvements

One big improvement came from Zaidoon, who had an insight about our struct for storing hashes in PBR. That struct looks like this:

struct Point {
    hash: u32,
    index: u32,
}

In memory this is represented as eight bytes, where four go to the hash (which is unavoidable), and four go to an index pointing to the server which is stored in another array. Zaidoon’s insight was that a 32-bit integer for that index is wasteful, because PBR is not likely to ever have to coordinate more than $m2^16 approx 65text{k} m$ servers at the same time, so a 16-bit integer will work. So we can replace the struct above with this one:
struct PointV2 {
    hash: u32,
    index: u16,
}

Unfortunately, Rust doesn’t make it that easy. Changing the size of the index as we did above does nothing to reduce the memory footprint. This is because Rust has alignment rules that require the size of a structure in memory to be a multiple of its largest (or “most aligned”) field. In this case, the hash is the largest with four bytes, so when stored in memory, a Point is required to have size $mN times 4m$, so the minimum size is eight bytes.

Luckily there are well-known ways around this. You (meaning me) might be tempted to use #[repr(packed)], but that is controversial for good reasons. A safer but less readable solution is to store the hash and index as raw byte array and access them with getters. Both methods compile to the same thing.

struct Point([u8; 6]);

impl Point {
   fn hash(&self) -> u32 {
	u32::from_ne_bytes(self.0[0..4].try_into().unwrap())
   }

   fn index(&self) -> u16 {
	u16::from_ne_bytes(self.0[4..6].try_into().unwrap())
   }
}

This simple (if wordy) change reduces the amount of memory used for consistent hashing by a whopping 25%! In order to do better than that, we’ll need to jump back into the math, so everybody hang on to something; this is the home stretch.

What if we tried fewer hashes?

You may have noticed that we gave the formula for the standard deviation for the case where there is only one hash per server. Deriving the formula for the case where there are $m k m$ hashes per server is not easy, and most sources only give you an approximation or an asymptotic limit, but not us. I might not be a statistician, but I grew up with a calculus teacher (Hi, Mom!), and I wanted to know the actual value. The full derivation is in a supplemental post, but here is the payoff.

$$m

    text{Exp}_k = frac{1}{N},

    text{SD}_k=sqrt{frac{(k+1)}{N(kN+1)}-frac{1}{N^2}}

m$$

To see how increasing the hash count improves the accuracy, we need to look again at the coefficient of variation.

$$m

text{CV}_k=frac{text{SD}_k}{text{Exp}_k}=sqrt{frac{N-1}{(N*k+1)}}

m$$

Plotting $mtext{CV}_km$ shows a potential problem with the “just add more hashes” mentality (other than overusing RAM).
BLOG-3083 6.png
You can see each step down in error margin requires (almost) an order of magnitude increase in the number of hashes per server, so adding more hashes yields less and less improvement. Recall that we are using a base of 160 hashes scaled by the server’s storage size. To make the math easier, we’ll say the weighting factor $m{m_w}m$ for a server is 625, so we get $m{k = 160times625 = 100{,}000}m$. We can see from the chart above that the last 90,000 hashes we added are buying us a minuscule 0.7% reduction in error. Unfortunately things get even worse from there.

The predictions from my beautiful math only work if we think about hashes in a continuous ring, but in practice we use 32-bit numbers for the hashes that have the potential for collisions, and the probability of collisions goes up surprisingly quickly as the number of hashes increases (see the birthday paradox). Collisions matter because in the ideal case, every hash contributes to the volume and distribution of requests handled by the associated server, but a collision means some contributions are randomly dropped, introducing unpredictable error. If we compare some simulated results with 32-bit hashes with the predicted error rate, we can see that for data centers with 2048 servers, the error rate increases: between 10,000 and 100,000 hashes per server.

BLOG-3083 7.png

Ultimately, even though this realization feels kind of bad, it’s great news for our plan to reclaim some RAM! Now that we have some math to back it up, we determined that we could decrease the number of hashes we were generating for each server by 90% without incurring any appreciable error, so that is what we set out to do.

Migrating without melting origins

There was one more problem: changing the hash ring changes where some cacheable requests go. Even if the new ring is better, switching the whole network at once would effectively invalidate almost all cached content. It would turn a memory optimization into an apocalyptic increase in origin traffic.

So we did not make this a single global flip. For a while, PBR carried both versions of the cacheable load balancer in memory: the old ketama ring and the new smaller one. Each request used our normal migration framework to decide which ring should select the backend. That meant the rollout decision was stable per request hash, and it also gave us a clean rollback path. If anything looked wrong, we could send new requests back through the old ring without redeploying PBR.

We then rolled the migration out in layers. We started with small validation locations, moved through progressively larger groups of data centers, and only then continued toward the rest of the world. 

The important part was that we controlled two dimensions independently: how much traffic used the new ring, and where that traffic was allowed to move. A plain global percentage rollout would have spread cache churn everywhere at once. Data-center-scoped rollout kept the blast radius small and made it much easier to tell whether a change was actually safe.

During the migration, we watched backend-selection traces, ring-version counters, PBR connection errors, process memory, startup time, cache behavior, and origin traffic. Once the migration reached 100%, we removed the temporary old-ring path, and voila!

BLOG-3083 8.png

The chart above shows the comparison of the memory used by PBR the week of the change compared with data from a few weeks before, as well as the result of subtracting one from the other. The sharp drop is the day where the version of PBR with the large (now unused) hash rings was decommissioned forever. Looking at the difference, we get the satisfying result that our changes dropped the used memory by 100TB!

BLOG-3083 9.png

Try it yourself

All the changes we talked about in this post are available now in the pingora-ketama crate in the form of a (for now) unadvertised cargo feature. The v2 ring has the compacted storage format, a faster sorting method, and the ability to scale the base number of hashes per node. Our focus in making these changes had to be on stability and control, so the v1 ring is identical to what pingora ketama has always used, and the library makes it possible to run both simultaneously and decide on a request-by-request basis which to use and when. 

Beyond trying our literal consistent hashing changes, I would like you to take away from this some inspiration to dig into your own systems to see what “simple” or “obvious” decisions are hiding potential wins, if you’re willing to get into the numbers. You might not be able to solve all your problems with Rust, but math is universal.

Follow on Social Media


Source: Hacker News

US Military had close call after using AI for hallucinated intelligence report

US Army soldiers conduct unmanned aerial system training in April 2026.

The intelligence report, circulated across the US military this spring in the midst of the war with Iran, immediately set off alarm bells: A Chinese ship in the Middle East was transporting components of a nuclear weapons program.

The US military swung into action with plans to intercept the vessel, according to four sources familiar with the episode. According to two of the sources, armed members of the US military were preparing to board the ship. Military planes were in the air, one of those sources and another source familiar with the incident said.

It was only just before the planned operation that officials dug deeper into the report put together by a special operations command analyst and found it had been generated with the help of artificial intelligence (AI) — and that a chatbot the analyst had used inaccurately identified the material the ship was carrying. CNN was not able to learn what the misidentified cargo was.

The report, according to one of the sources, was “entirely false.” But it also “almost started a war,” the source said. Any US operation against a Chinese vessel could have risked spiraling into an armed conflict between the two nations.

Across the US military and the intelligence community, officials are pushing to weave AI into nearly every facet of their work, from analyzing the huge volumes of raw intelligence the US collects and selecting targets for strikes, to more mundane applications like managing budgeting, logistics and supply chains.

But the episode underscores the profound risks of using this powerful, new and relatively poorly understood technology for targeting in the middle of a war. Analysts have long feared that AI could lead to a catastrophic miscalculation if nation states are relying on poor or corrupted data — the kind of miscalculation that might lead the United States to fire on a Chinese ship based on inaccurate information.

In this particular instance, the analyst queried a chatbot about some intelligence reporting on the ship’s manifest that originated with US Special Operations Command Pacific, based in Hawaii. It was not clear whether the chatbot was a commercially available one or a US government product.

“The internal tools are mostly just copies of the commercial stuff wearing lipstick,” a former senior US official familiar with the AI systems used by military and intelligence analysts.

The bot fused together open-source intelligence with secret signals intelligence in government holdings and reached its fateful conclusion about the material the ship was carrying.

The analyst then used AI again to package the findings into a standard intelligence report — the kind that is trusted by military officials — and disseminated it.

US Special Operations Command Pacific and the Pentagon did not respond to a request for comment.

The rationale for the rapid adoption of AI is that it can help the military make battlefield decisions, like which targets to strike or which military assets to move where, faster. Officials say the US can’t afford to fall behind in integrating AI in case it must one day fight China or another adversary who would potentially be able to stay one step ahead of the US.

In January, Defense Secretary Pete Hegseth released his agency’s “Artificial Intelligence Acceleration Strategy” in a bid to speed up the military’s use of AI.

“We will unleash experimentation, eliminate bureaucratic barriers, focus our investments and demonstrate the execution approach needed to ensure we lead in military AI,” Hegseth said in a speech announcing the strategy.

Secretary of War Pete Hegseth testifies during a Senate Appropriations Committee hearing in the Dirksen Senate Office Building on Capitol Hill on July 21, 2026 in Washington, DC.

The strategy also pushes for its broad use across the military, ordering the department to make AI available via several programs with the aim of “democratizing AI experimentation and transformation across the Department by putting America’s world-leading AI models directly in the hands of our three million civilian and military personnel, at all classification levels,” a memo announcing the strategy said.

But the effort is decentralized, multiple US officials familiar with the dynamic said, with different parts of the government using different tools under different orders and safety standards. There’s no one set of standards for how the US verifies the information generated by these tools. The constellation of different AI systems being deployed by disparate corners of the military and intelligence community means that the relative reliability and functionality vary widely.

For weeks, Washington policymakers have been intensely debating AI after a series of dire warnings from Silicon Valley engineers and tech CEOs of the possibility that AI could break free of human constraints, with potentially civilization-ending consequences.

But the episode with the Chinese ship underscores a different, and more immediate risk: human beings making disastrous decisions based on inaccurate or misleading information generated by AI or other automated systems. Sources said that the military is rapidly turning to AI to help with targeting, an area which holds the obvious risk of fatal mistakes.

“AI in targeting is definitely something that is ramping up and there is no real guidance for how having a human in the loop will prevent civilian casualties or fratricide,” another source familiar with the military’s current policies said.

The kind of “hallucination” that the tool used by the analyst in this case conjured has not been an isolated incident across the intelligence community since these tools began proliferating across government, according to one of the sources.

For some older intelligence officials — even those who broadly support the use of AI inside the military — AI has put pressure on analysts to produce and disseminate intelligence faster, opening the door for mistakes. Young analysts in particular, several sources said, are natives on these tools and more likely to trust them uncritically.

“AI allows you to get to a bad idea faster,” one of the sources said.





Source: Hacker News

Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug

TL;DR

— Photon-emission microscopy allowed us to locate a register responsible for the enabling of debug features on the Raspberry Pi microcontroller.

— Laser pulses at two nearby positions then restored debugger access to the chip’s Secure world, even though debug had been permanently disabled.

— Using that access after a rescue reset, we recovered a secret from one-time-programmable memory. The reset halted the chip before firmware could apply its runtime lock, so the page stayed Secure-readable.

— The attack requires physical access, destructive preparation, and approximately $250,000 of laboratory equipment.

The RP2350 security model

The RP2350 is Raspberry Pi’s dual-core microcontroller: each processor socket can select either an Arm Cortex-M33 or a RISC-V Hazard3 core at boot. Its hardware security features include:

  • Secure boot, which authenticates signed firmware against public-key fingerprints provisioned in One-Time Programmable memory (OTP)
  • The Armv8-M TrustZone, which separates Secure and Non-secure execution states
  • Permanent debug-disable settings
  • Glitch detectors intended to detect timing disturbances caused by clock or supply manipulation

Raspberry Pi has actively invited researchers to evaluate these protections through its RP2350 Hacking Challenges. The first challenge ran from August to December 2024 against the original chip. After several findings were addressed, Raspberry Pi released the A4 revision—the version we tested.

The permanent security configuration and boot public key fingerprints are stored in one-time-programmable (OTP) memory: each bit can be flipped from 0 to 1 once and never back, so whatever is written there lasts for the lifetime of the chip.

OTP is organised into 128-byte pages protected by two persistent, or hard, lock rows: for page n, PAGEn_LOCK0 configures optional read and write keys and the behaviour when no key is entered, while PAGEn_LOCK1 contains the hardware-enforced LOCK_S and LOCK_NS permissions. Those states can advance from read-write to read-only or inaccessible but cannot become more permissive.

The OTP subsystem uses redundant encodings for security-related fields: critical flags are “encoded with a three-of-eight vote across eight consecutive OTP rows”, and OTP lock bits are “triple-redundant with a majority vote”, according to the RP2350 datasheet.

At an OTP reset, the persistent LOCK_S and LOCK_NS values initialise a per-page runtime lock, also called a soft lock. Firmware can tighten this lock until the next OTP reset, but cannot loosen it. The runtime change does not survive that reset.

An external debugger communicates with the RP2350 through Arm’s Serial Wire Debug (SWD) interface. Requests first reach the Serial Wire Debug Port (SW-DP) and are then routed to access ports. In the Cortex-M33 configuration used here, each core has a memory access port (Mem-AP) connected to its system bus; an enabled Mem-AP lets the debugger read and write permitted memory and peripherals. A separate always-on access port, the RP-AP, exposes a small set of reset and recovery controls.

Secure debug refers to Mem-AP access with Secure attribution. The debugger can then transact with Secure memory-mapped resources the access-control logic permits, and halt or inspect a core running in the Secure state.

The permanent CRIT1.DEBUG_DISABLE flag is intended to close this path. When set, it drives the enable signals for both cores’ Mem-APs to zero, which “prevents the APs from performing any bus accesses at all”, and disables the factory-test JTAG interface and the RISC-V debug module’s access port. The SW-DP and RP-AP still respond, but neither core Mem-AP can access the system bus.

There is, however, an override: the memory-mapped DEBUGEN register lets Secure software re-enable each core’s Mem-AP and, separately, Secure accesses through it. The datasheet states that DEBUG_DISABLE “can be fully overridden by setting all bits of this register”.

This critical override in the enforcement chain is what made the debug interface our target. Gaining access to Secure debug on a Mem-AP is a general-purpose primitive to read and write Secure memory, halt and single-step a core, and inspect its registers. Whether that register could be set by a fault is the question the rest of this post answers.

Experimental setup

Target configuration

Raspberry Pi’s RP2350 Hacking Challenge asked participants to extract a 128-bit secret stored in OTP1. At startup, the signed challenge firmware ensures that page 48 has the expected persistent lock, then applies a runtime lock that denies both Secure and Non-secure access to the secret until the next OTP reset.

We replicated this vendor-defined configuration on our own revision A4 device:

  • Programmed the SHA-256 fingerprint of our public key into BOOTKEY0
  • Set BOOT_FLAGS1.KEY_VALID to 0x1 and BOOT_FLAGS1.KEY_INVALID to 0xe
  • Enabled secure boot (CRIT1.SECURE_BOOT_ENABLE = 1)
  • Permanently disabled debug (CRIT1.DEBUG_DISABLE = 1)
  • Enabled the glitch detectors at maximum sensitivity (CRIT1.GLITCH_DETECTOR_ENABLE = 1, CRIT1.GLITCH_DETECTOR_SENS = 3)
  • Configured the persistent locks for pages 1 and 2 according to the challenge configuration
  • Set the page 48 persistent lock to PAGE48_LOCK1 = 0x3c3c3c, which denied Non-secure access (LOCK_NS = INACCESSIBLE) while retaining Secure read-write access (LOCK_S = READ_WRITE)

Enabling secure boot permits only the Cortex-M33 cores, so both processor sockets used Arm for these experiments.

Sample preparation and bench

The device was backside decapsulated, so that infrared light reaches the transistors through the silicon substrate rather than being blocked by the metal layers on the front. The chip was then soldered back onto a daughterboard connected to Scaffold, Ledger Donjon’s open source platform for driving and monitoring devices under test. Removing the lead frame on the backside of the chip breaks its GND connection, so a copper wire restores it2.

Backside-decapsulated RP2350 mounted on the analysis daughterboard
Backside-decapsulated RP2350 mounted on the analysis daughterboard
Experimental bench used for the attack
Experimental bench used for the attack

DEBUGEN: overriding permanent debug disable

DEBUGEN has five functional bits:

Bit Name Effect
0 PROC0 Enable core 0’s memory access port
1 PROC0_SECURE Permit Secure accesses through core 0’s memory access port
2 PROC1 Enable core 1’s memory access port
3 PROC1_SECURE Permit Secure accesses through core 1’s memory access port
8 MISC Enable additional debug components, including the cross-trigger interface and the RISC-V debug access port

Secure debug on a core needs both of its bits: the one that enables the Mem-AP, and the one that permits Secure accesses through it.

In contrast to the redundant encoding used for OTP security fields, the datasheet documents no bit redundancy, parity or majority vote for DEBUGEN.

We therefore tested whether laser pulses could set DEBUGEN bits on the secured device described above.

Photon-emission-guided localization

That test first requires knowing where to aim. Setting an individual DEBUGEN bit means hitting the storage of a single register bit, a needle in a haystack. This is a harder targeting problem than the instruction-skip faults common in laser fault injection, where disturbing any of the many flip-flops in a core pipeline can produce the same skip: that spreads the sensitive area widely enough for a random scan to find it. A blind scan for one DEBUGEN bit is impractical.

Switching transistors emit faint near-infrared photons correlated with their activity, so collecting that emission over repeated execution can reveal where a selected control changes state. This made photon-emission microscopy (PEM) a good fit for DEBUGEN: as a memory-mapped register, Secure software can toggle exact bits in a loop, driving the repeated state changes the measurement needs. We used it as the first localization stage, and the resulting map constrained the subsequent laser scan to a region of a few micrometres.

We compared loops that repeatedly toggled selected DEBUGEN bits on and off, differing only in the bits they targeted. A register’s photon emission is faint next to the camera’s own noise and sensitive to slowly drifting ambient conditions such as temperature, so a single frame reveals nothing. Averaging many frames of each loop suppressed random sensor noise, and subtracting the two mean stacks cancelled everything the loops shared: static background, sensor offset, thermal emission, and switching unrelated to the selected bits. Interleaving the two values during acquisition kept slow drift from biasing that subtraction. What remained was the emission that tracked the selected bits.

Mean stacks of 200 full-view photon-emission captures for DEBUGEN masks 0x3 and 0xc followed by their signed difference.
Mean stacks of all 200 mask 0x3 and mask 0xc acquisitions, followed by their signed difference. Red is positive, indicating greater emission for 0x3; blue is negative, indicating greater emission for 0xc. The localization maps below additionally balance acquisition order before combining matched differences.

Repeated comparisons across different bit masks exposed compact sites associated with DEBUGEN bits 0–3 across three regions of the camera field.

Infrared overview of the die with three marked regions, plus zooms of those regions overlaid with coloured DEBUGEN bit sites.
Infrared overview of the camera field, with three marked regions. Coloured pixels mark sites associated with `DEBUGEN` bits 0-3.

These zones show switching activity associated with each DEBUGEN bit; they do not directly identify storage cells. The multiple hotspots observed for each bit may arise from the storage element or from related logic. Without layout data, we cannot distinguish between the two. However, these zones still significantly reduce the search space.

Finding 1 — faulting DEBUGEN gives Secure debug

For laser fault injection (LFI), we used a pulsed laser at 980 nm with 2.97 W maximum optical power, operated at roughly 40% (about 1.2 W), with a 100 ns pulse width through a 50x objective. After each pulse, we probed the debug access ports over SWD.

Within the area found from PEM, we ran an LFI scan and used that SWD feedback to calibrate two responsive positions a few micrometres apart. At one position, pulses enabled bus access through core 1’s Mem-AP, indicating that PROC1 was set. At the other, the Mem-AP’s Control/Status Word reported SDeviceEn = 1, a state-guided signal that PROC1_SECURE was likely set. We checked both indicators after every pulse.

Side-by-side infrared views with PEM bit sites on the left and LFI fault points on the right.
Left: PEM sites associated with `DEBUGEN` bits. Right: laser-fault points on the LFI infrared view.

A pulse that set one bit could clear the other, so setting both required an iterative sequence. Our script pulsed the PROC1 position until bus access was available, then pulsed the PROC1_SECURE position until SDeviceEn = 1, returning to the first position whenever bus access was lost. Once the positions and pulse parameters were calibrated, the sequence enabled Secure debug within seconds. Interestingly, we could not reproduce this sequence using a 20x objective. Because the two positions are only a few micrometres apart, that wider spot likely hit both the region that sets a bit and the one that clears it, so it was not possible to obtain the correct value.

Once both bits were set, they remained set without further pulses or software writes. Reading the Secure-only DEBUGEN register through core 1’s Mem-AP then returned 0xc; because DEBUGEN is Secure-only, that successful read confirms the transaction was Secure-attributed.

Enabling Secure accesses through core 1’s Mem-AP allows the debugger to read and write memory-mapped resources whose ACCESSCTRL permissions admit the debugger as a bus manager and whose target-specific controls admit Secure AHB transactions. Independently of those direct reads, the debugger can halt and single-step the core and inspect or modify its registers, compromising TrustZone runtime isolation through Secure-core-mediated extraction. This does not make the boot ROM accept unauthenticated firmware: when firmware boots normally, secure boot still authenticates it, but cannot protect runtime state that remains accessible to the debugger after verification.

Application to the Hacking Challenge configuration

The Secure-attributed Mem-AP access described above exposes Secure runtime state, but the challenge’s page 48 runtime lock still prevents access to the secret after firmware has run. The page’s persistent lock, PAGE48_LOCK1 = 0x3c3c3c, denies Non-secure reads but leaves LOCK_S at READ_WRITE, so it remains readable through Secure-attributed accesses before the runtime lock is tightened.

During each boot, the firmware writes the most restrictive binary value, 0b1111, to the runtime lock otp_hw->sw_lock[48]. That register then makes the page inaccessible to both Secure and Non-secure accesses, including Secure debug, and therefore prevents Secure transactions from the Mem-AP from reading the secret.

As documented, software locks “are initialised from the OTP lock pages at reset”, and a write only advances the state “until next reset”. Resetting the OTP block discards 0b1111 and restores the value derived from PAGE48_LOCK1, for which LOCK_S = READ_WRITE.

The remaining question is how to reset a locked chip without allowing firmware to re-apply the runtime lock. The RP-AP remains “always accessible, even when external debug is disabled”. Setting CTRL.RESCUE_RESTART triggers a rescue reset: a full system reset that also flags the boot ROM to halt before any user software runs.

The boot ROM checks POWMAN_CHIP_RESET.RESCUE_FLAG before watchdog, flash or USB boot, clears it, then holds core 0 in an interrupt-disabled wait loop and core 1 in its wait-for-vector path.3 The datasheet documents no restriction on CTRL.RESCUE_RESTART.

We proceeded in the following sequence:

  1. Rescue reset. Set CTRL.RESCUE_RESTART to 1, then clear it to 0 through the RP-AP. The chip resets and remains in boot-ROM wait paths. The signed firmware never runs, so sw_lock[48] is never tightened and stays at the permissive value derived from PAGE48_LOCK1 — LOCK_S = READ_WRITE.
  2. Fault DEBUGEN to 0xc. With both cores in boot-ROM wait paths, set PROC1 and PROC1_SECURE as described above; these two set bits produce the value 0xc.
  3. Halt core 1 through its Debug Halting Control and Status Register (DHCSR) over the now-Secure Mem-AP.
  4. Read the secret from OTP rows 0xc08–0xc0f through the guarded read interface.

We ran this sequence on the tested device and recovered the complete challenge secret.

DEBUGEN_LOCK does not prevent laser-induced changes

DEBUGEN_LOCK blocks software writes to the corresponding DEBUGEN bits: each lock bit is “Write 1 to lock the […] bit of DEBUGEN. Can’t be cleared once set”. The datasheet presents this as a way “to avoid accidental writes”.

In trials with the target DEBUGEN bit at 0 and its lock bit at 1, a pulse could still set DEBUGEN while the lock remained 1. Pulses also set lock bits, with or without a corresponding DEBUGEN change. In successful sequences, all five functional lock bits were 1 by the time PROC1 and PROC1_SECURE were both set. We never saw a lock bit return from 1 to 0, so a later write of DEBUGEN = 0 cannot restore the disabled state once the fault has set the corresponding lock.

Limits of software-based mitigations

Once Secure accesses through the Mem-AP are enabled, Secure attribution alone no longer separates the debugger from Secure firmware. This access does not override hard OTP locks or peripheral-specific controls. ACCESSCTRL can block direct debugger-manager transactions to particular targets, but it does not by itself prevent a debugger controlling the Secure core from causing core-originated accesses or extracting loaded values through core registers. After a rescue reset, ACCESSCTRL returns to its all-open reset-time defaults before firmware can reconfigure it. ACCESSCTRL therefore reduces direct Mem-AP exposure rather than forming a standalone confidentiality boundary.

Firmware can nevertheless reduce post-boot exposure by denying the debugger access to sensitive targets in ACCESSCTRL, then setting the debugger bit in ACCESSCTRL.LOCK so that debugger transactions cannot reopen those permissions. Secure firmware can also check DEBUGEN periodically and, on an unexpected value, trigger a fail-safe reset that clears the processor-cold reset domain. These measures are best-effort runtime mitigations: an enabled debugger may halt the core before the next check, and rescue reset stops before firmware can configure ACCESSCTRL or run the monitor. They therefore do not prevent the pre-firmware secret read demonstrated here.

RP2350’s documented encrypted-boot flow illustrates the pre-firmware limitation of runtime locks and the post-boot limitation of debugger-manager filtering in two distinct machine states. After rescue reset, the boot ROM halts before decryption: no plaintext payload exists yet, but the decryption key may be directly readable if the OTP page’s persistent permissions allow Secure access and no other target control blocks the transaction. After normal encrypted boot, plaintext exists in SRAM: direct Mem-AP reads depend on debugger-manager permissions in ACCESSCTRL, while Secure-core control may permit core-mediated extraction even when direct reads are denied. This is architectural analysis, not a tested encrypted-boot result; encrypted boot still protects external flash from offline inspection.

Impact and attack requirements

The demonstrated sequence provides Secure-attributed memory access, control over Secure-world execution, and access to the challenge secret after resetting its runtime page lock. It requires the following resources:

  • Destructive physical access. Backside decapsulation permanently modifies the package and leaves the die exposed.
  • Specialised laboratory equipment. The complete setup described above costs approximately $250,000.
  • Hardware-security expertise. The procedure requires sample preparation, die navigation, laser parameter selection, and coordinated laser control, stage positioning, and SWD measurement.

Conclusion

The RP2350 encodes critical debug-disable flags in OTP with redundant voting, but DEBUGEN can override their effect and has no equivalent protection documented in the datasheet. In our experiments, laser pulses changed DEBUGEN despite DEBUGEN_LOCK and could set a lock bit that prevented firmware from restoring the disabled value. Separately, the RP-AP rescue reset restored the challenge’s runtime page lock to its persistent value while preventing user firmware from executing. The software-visible mechanisms each performed their documented function, but their interaction with the laser fault enabled Secure debug and recovery of the challenge secret. Differential PEM first isolated bit-dependent DEBUGEN activity, and guided LFI converted that spatial lead into persistent Secure debug. The system-level lesson is that security analysis must cover the complete enforcement path, from persistent OTP configuration through mutable control registers and reset behaviour, because system security depends on that path rather than on individual mechanisms in isolation.

Disclosure and acknowledgements

We disclosed this fault to Raspberry Pi on 28 July 2026. We thank the Raspberry Pi team for their engagement in the disclosure discussions and for their transparent approach to security research.


Antoine Plin, Hardware Security Intern at Ledger Donjon

Footnotes

  1. https://github.com/raspberrypi/rp2350_hacking_challenge The RP2350 Hacking Challenge repository, containing the reference lockdown configuration and firmware we replicated. ↩

  2. Courk, Laser Fault Injection on a Budget: RP2350 Edition. ↩

  3. The rescue check is step 1 of the core 0 boot path in src/main/arm/varm_boot_path.c; in src/main/arm/arm8_bootrom_rt0.S, varm_wait_rescue enters the interrupt-disabled varm_dead_quiet WFI loop while core 1 remains in the boot ROM’s wait-for-vector path. ↩


Source: Hacker News

Jemalloc 5.4.0

5.4.0


Latest

Choose a tag to compare

Loading

@guangli-dai
guangli-dai

released this

17 Sep 18:31

·

5 commits

to dev
since this release

This release contains over 160 commits, focusing on the technical debts
cleaning including refactorings, bug fixes, test coverage improvement, and
option cleanups. The release also includes portability improvements per
upstream issues report.

New features:

  • Add EXTENT_ALLOC_FLAG_PINNED so custom extent-allocation hooks can
    mark non-reclaimable mappings, such as HugeTLB pages, for preferential
    reuse outside the decay and purge pipeline. Add the mallctl interfaces
    stats.pinned, stats.arenas.<i>.pinned,
    stats.arenas.<i>.extents.<j>.npinned,
    stats.arenas.<i>.extents.<j>.pinned_bytes, and
    stats.arenas.<i>.mutexes.extents_pinned.{counter} to report
    pinned-memory usage and mutex statistics. (@binliu19: be2de8c)
  • Allow resuming per-CPU arena selection via thread.arena.
    (@Algunenano: 68c35f6)
  • Better align the contents of human-readable and JSON malloc statistics.
    (@spredolac: 68fe1be, 30b1a41, 04aad97)
  • Replace the runtime experimental_infallible_new option with the
    compile-time --enable-cxx-infallible-new option, enabling
    compiler-level optimizations and optimization in move constructors,
    and fix the new(std::nothrow) contract (@spredolac: 160ab9d, fe33667)

Incompatible changes:

  • Adapt tcache fill and retention targets per bin to demand observed
    between GC events, replacing the fixed refill/flush policy. Remove
    seven legacy non-experimental controls: lg_tcache_nslots_mul,
    tcache_nslots_small_min, tcache_nslots_small_max,
    tcache_nslots_large, tcache_gc_delay_bytes,
    lg_tcache_flush_small_div, and lg_tcache_flush_large_div.
    Corresponding malloc_conf settings are silently ignored, matching
    opt.* mallctls return ENOENT, and tcache_ncached_max remains
    supported. (@spredolac: d13fe91)

Bug fixes:

  • Preserve errno across free, free_sized, and free_aligned_sized,
    and across process_madvise-based page purging. (@spredolac: 3a77966,
    86f0582)
  • Fix numeric overflow checks in size classes. (@spredolac: 6b24522)
  • Accept NULL in free_sized() and free_aligned_sized() (C23
    correctness). (@bigbruno: 7ce8b91)
  • Fix TSD lifecycle edge cases by 1) initializing thread-cache bins before
    marking the cache enabled, preventing reentrant bootstrap allocations
    from using uninitialized state, and 2) avoiding TSD recreation for late
    deallocations after thread teardown on generic-TSD platforms.
    (@fzakaria: 54f22c8, @spredolac: fb5499a)
  • Use O_CLOEXEC when opening the THP sysfs file in init_thp_state.
    (@ibookstein: d070554)
  • Fix duplicate opt.stats_print and opt.stats_print_opts fields in
    malloc_stats_print output. (@spredolac: dfe3a2e)
  • Fix a potential deadlock during arena_reset. (@guangli-dai: 6957341)
  • Fix a prof-sampling / guard-page interaction bug in the SAN. (@gctony:
    e36a0fa)

Optimizations and refactors:

  • Modularize jemalloc’s front end by extracting arena management,
    initialization, fork orchestration, and allocation dispatch from
    jemalloc.c; untangle tcache/arena ownership; and consolidate the
    internal header graph to eliminate circular dependencies. (@spredolac:
    ba1e2fe, 8874597, 9d75722, de9ad14, …)
  • Simplify ctl dispatch by refactoring arena helpers, internalizing an
    implementation-only control, replacing control-flow macros with typed
    helpers, and organizing ctl.c by subsystem. (@spredolac: 9e14346,
    901365d, 4e903a0, 5bc8d6e)
  • Cap the base-block growth heuristic to avoid virtual-memory exhaustion
    under rare racy conditions. (@guangli-dai: 2f4db8c)
  • Move background-thread lifecycle and state operations into the
    background-thread module, clarifying ownership independently of PAC/HPA
    callers. (@guangli-dai: f1f0792)
  • Simplify the page-allocation boundary by replacing PAI vtable dispatch
    with direct PAC/HPA calls, removing the obsolete pai_t/pai.h
    abstraction, and moving deferred-work and decay orchestration out of the
    arena. (@guangli-dai: 1dfa6f7, 8edd101, d410f43)
  • Refactor statistics collection and rendering into separate
    gather/emission stages and descriptor-driven tables. (@spredolac:
    184f304, 80c8fcb)
  • Introduce an OS abstraction layer and move platform-dependent
    file/process I/O, time, synchronization, CPU, virtual-memory, atfork,
    error-handling, profiling, thread-yield, and configuration-access
    operations out of allocator core code. (@guangli-dai: c4158ac,
    c8e2e01, …)

Portability improvements:

  • Replace the std::__throw_bad_alloc call with standard C++ (#2900).
    (@lexprfuncall: 1a15fe3)
  • Make arena_s use a flexible array member (bin_t all_bins[]) for C99
    or newer. (@grueninger: 300b58b)
  • Fix rdtscp detection with --with-lg-vaddr. (@xinydev: e8a0d2b)
  • Fix malloc_getcpu on macOS to read the current CPU number correctly.
    (@Algunenano: 5aabbc8)
  • Use CLOCK_MONOTONIC for background-thread sleep to prevent
    clock-rollback stalls, and detect monotonic-condvar support at configure
    time. (@antonio2368: ebacec3, 8361239)
  • Fix compilation warnings on macOS. (@gctony: abb0a8a)
  • Fix thread-exit TSD cleanup on MinGW builds. (@gctony: 1e92317)
  • Fix GCC 16 build warnings by removing -Wpedantic violations in macro
    and function syntax and explicitly NUL-terminating profiling thread-name
    copies to resolve -Wstringop-truncation. (@grueninger: 5acdcee,
    c22d929, @jasonangelov: 6a245e0)
  • Parse PID-namespace symlinks without glibc-dependent strtok/atol,
    and return identifiers as uint64_t. (@guangli-dai: 278d90a)


Source: Hacker News

Backups Aren't Simple

See also: John Salvatier’s excellent blog, Reality has a surprising amount of detail

I read a comment somewhere that stuck with me, that went something like this:

“There are two types of people: those who have suffered a catastrophic loss of data, and those who will.”

Trying to find the source for it for this blog, it turned out that every other sysadmin has his rehashed version of the quote, but the gist of it is the same everywhere. Data loss is something that happens more often than we’d hope, and most of us are woefully unprepared for when it hits us (which is almost always at the worst possible time).

I can confirm that I had a similar experience once. When I was little we had pulled all our family photos from our home laptops and PCs onto an external hard drive, in order to free up some space. This worked beautifully until one day my dad wanted to use the drive as storage for our TV set-top box (one of these old things), and was prompted to format the drive. He went ahead with it, and the disk was reformatted. The index of files was deleted, and we were stuck with a nominally empty drive.

It would be easy to blame him for screwing up, but it takes beginning a career in tech to realise that there is a series of errors that lead to this kind of mistake. Firstly, we had put all of our photos in one place and didn’t bother with backups. Secondly, most consumer-facing software usually has bold disclaimers telling you that formatting a disk means losing data (which the set-top box didn’t, terrible UI). Besides, why would you even expect a non-technical person to even have to know any of this?

Thankfully we were able to get the photos restored, and it turned out to be a cheap lesson in handling data. You never keep important things in one place only. There’s about a million things that can go wrong. Your drive could die, it could be stolen, bits could rot in cold storage (hard drives have magnetic particles which can inexplicably shift, and SSDs are made of NAND transistors which leak electricity and over time, corrupt your data).

So our first principle is to have a backup, i.e. a copy of your files someplace else. So far so good.

This doesn’t cover the headaches of what a plugged in drive could do. Ransomware could encrypt your files, and you could do anything from an honest mistake like deleting the wrong file; up to catastrophic mistakes like running a script that overwrites everything with zeroes.

So our backup should not be a mirror of the first drive, because we also want to be able to go back in time if we mess up. Importantly, this means that mirroring your disk with something like RAID 1 is out. We need some other method that snapshots things.

How often do we want to take snapshots? Maybe in our case with the photos we should have run a backup every week. If we lose 6 days and 23 hours of data, that’s fine and we can live with it. This is what’s called a Recovery Point Objective (RPO) in IT, and in real cases, it ranges from <30 seconds for critical financial institutions which really can’t afford to lose data, to 24 hours or more for some small enterprises (if they even have a disaster recovery strategy).

Taking snapshots means that we have an increasing burden on our storage. With an RPO of 24 hours, you will end up having 7 snapshots per week. 30 per month. 365 per year, if you really don’t go and prune your snapshots. So you need to rotate your backups.

Let’s say I go with the naive approach and decide to keep 14 days’ worth of snapshots. When I take a new snapshot, I delete the oldest one and I add the new one. Pretty simple, but this now forces me to have a watchful eye. Maybe I keep lots of data and can’t be bothered to check if something got corrupted in the past two weeks? But then again, I can’t just store a year’s worth of backups and they’re simply not relevant to me. What happened between day 2 and day 3 of the year has almost no significance when it’s day 364. So the granularity at which we take backups must change. The closer we are to today, the more frequent the snapshots. The further back, the less frequent the snapshots.

So maybe we rotate our daily backups every 14 days, but also take weekly backups that we rotate every 7 weeks, and monthly backups we rotate every 12 months. This should be much more efficient. But again our complexity grows. We now have something called a GFS-rotated, snapshot-based backup. This list of adjectives will continue growing, as we’ll see in a bit.

Maybe then you take a look at how MPEG compresses video, and get fascinated by how a calm scene in a movie, where the protagonist speaks but otherwise doesn’t move against a completely still background can be used for compressing video. You notice that videos are composed of frames that are mostly similar to each other, only changing with a certain movement that can be represented as a vector for a fraction of the storage. Which leads you to the very logical conclusion that your snapshots also follow the same pattern! Even more, it turns out that file changes follow a fat-tailed distribution, so over a given period there are a vast majority of files that don’t get changed at all, and a very tiny minority that change all the time.

So it becomes obvious that we shouldn’t store identical copies of files, but rather deduplicate. We can use hard links when we need to reference an already existing file. This way, we store one file on disk, and then reference it from each of our snapshots. This also survives backup rotation because we never delete files, we only delete directory entries. This exact approach is used by rsnapshot, and is best described as an incremental backup, because we store only the changes between two adjacent snapshots, instead of all the changes since the latest full backup (these are called differential backups and are more robust when restoring, but I won’t get into it for the sake of brevity).

The savings in storage are not the only benefit we get from doing this. We briefly mentioned in the beginning that we don’t store everything on a single machine. Obviously, there is also networking involved in this process, since we need to actually transfer the files from one machine to another. Deduplicated backups save a lot of bandwidth, which is especially important if you use a cloud service as your second machine. It directly affects you financially.

To sum up, by this time we have created an incremental, deduplicated, GFS-rotated, snapshot-based backup. We can use rsync to pull the files from the main machine, and cronjobs to run our backup scripts. We can run backups on as many machines as we’d like, and adding another one is trivial. Even better, file metadata is preserved, so things like access permissions and file ownership are fine.

Motivated by our success in developing this solution, we try to use it to backup the homelab with its 10 Docker containers. But later we find out from logs on the individual machines that backups are failing. The reason being that many Docker containers like to create root-owned files, and if you’re not careful you can create a cronjob running as the default user.

To make matters worse, almost every web app uses a database of some kind. Databases sometimes like to store things in-memory and flush them to disk in batches to improve performance. This practically means that restoring from backup will fail due to data corruption if we are unlucky. So we make the backup also dump the databases, and give it full filesystem permissions on our Docker volumes. That should make it work!

Then you read about incidents in which a model of hard drive had famously high failure rates, and start to wonder if you should maybe store your backups on two machines with different types of media. That way a hardware-specific failure would be unlikely to wipe out your backups. And while we’re on the topic of physical security, have one offsite backup on the cloud or a machine at a family member’s house. This way you make it really unlikely that a power surge, flood or fire will destroy everything. This is where the 3-2-1 backup gets its name: 3 copies, on 2 different types of media, with 1 offsite.

Let’s say you decide on a cloud provider for your offsite backup. Specifically object storage like Amazon S3. You quickly find out that our current setup won’t work because one, files lose their metadata when you upload them to S3, and two, the price for uploading many small files to S3 is punitively high. (file sizes also follow a fat-tailed distribution) These two facts make it best for you to stick many files into a tarball. That way you retain both your file metadata as well as your low costs. But the question is, how do you do that? Do you stick everything in one giant tarball? Obviously not, then your incremental backups with the hard links stop making sense. Best to split everything in clean 50MB chunks, but good luck with doing that in a way that is verifiably safe!

Up until this point, rolling your own backups sounded like something you should be able to do in an afternoon, but this is where I’d give up. It simply isn’t worth the mental load to do all this. Instead, you just use tried and tested tools like Borg or Restic which handle all this and much more (encryption, chunk-level deduplication, checksums). And just give your utmost thanks to the wonderful open-source community for building, maintaining, and live-testing these tools, while respecting how much trial and error was necessary to get to the point where all this complexity is abstracted away for us.

Obviously, none of this is worth anything if you don’t actually test restores. So that’s also a little digital hygiene article that you will have added to your to-do list. So, as long as you run restores every 6 months, you can enjoy your:

Also make sure not to run backups at 2AM or 3AM, or things may get scary.


Source: Hacker News

Why Does the Universe Expand?

Five hundred years ago, Nicolaus Copernicus proposed that the Earth might be one of several planets orbiting the Sun, rather than the centre of the universe. He compared the geocentric model to a monstrous form assembled from parts of different bodies, like the Creature Mary Shelley brought to life three centuries later in Frankenstein — each part appearing human on its own, but as a whole a grotesque patchwork.

The geocentric model was built to directly match the sky: where a planet paused against the background stars and looped into retrograde, a dial was added so the planet would pause in its orbit around the Earth and loop backwards a while before resuming its normal course. In contrast, Copernicus recognised that the outer planets — Mars, Jupiter, and Saturn were known at the time — might enter retrograde loops due to parallax, their apparent positions shifting relative to the background stars as our orbit brings us near and then we pass them on our way around the Sun.

Copernicus had no idea of the physics that Isaac Newton or Albert Einstein would eventually use to explain planetary motion. Nor did he imagine this motion resulted from the same phenomenon that causes apples to fall from trees, or paths of light to bend around the Sun. His model ended up with as many knobs and dials as the geocentric system due to his use of circles rather than ellipses to describe orbits. 

Even so, Copernicus recognised that a Sun-centred theory afforded the possibility that it might eventually explain why phenomena like retrograde motion should appear as they do to us, despite the planets’ motion being continuously in one direction only. And in doing so, he paved the way for others like Newton and Einstein, who later fleshed out both the underlying concepts and formal mathematical descriptions of a solar system in which planetary motion is expected to appear with all the complexity we observe. 

In hindsight, the discovery Copernicus’s proposal prompted was that broken symmetries — our off-centre perspective from a planet orbiting the Sun, and the non-uniform motions of all planets including ours — would complicate appearances within an ontological framework that is nonetheless simpler.

In a similar sense, it may be argued that the standard cosmological model today — which gives an accurate description of phenomena but is nevertheless an amalgam of ad hoc patches, each inserted to unnaturally force the evolution of a universe that is otherwise expected to be different from the way it appears — bears a closer resemblance to Frankenstein’s monster than it does the simple explanations of planetary motion given by Newton and Einstein. For despite all the dials and knobs that have been added to ensure the standard model does directly resemble appearances, after a century of development it still affords no explanation of why our universe should be expected to expand, as it appears to do.

Why should our universe expand?

In the Copernican tradition, we ought to ask why our universe should expand. The standard cosmological model affords no such explanation. It is based on a principle, famously promoted by Einstein together with his colleague Willem de Sitter, that characterises the universe as expanding in spite of a tendency to decelerate because it is filled with everything we see. This is important: according to the basic Einstein-de Sitter framework for describing cosmic expansion, all the galaxies and light we see across the universe are thought to work against the universe’s expansion, slowing it down; and anything driving expansion is an ad hoc dial we’ve added so that base model fits appearances better than it naturally should. 

In fact, this tendency for light and matter to slow cosmic expansion mathematically blows up to an infinite amount at the Big Bang. Therefore, the model’s only “explanation” for why our universe even could be expanding today is that it began with such a tremendous rate that momentum carried it against its natural tendency towards the opposite. For this reason, a century ago British astronomer Arthur Stanley Eddington complained of the Einstein-de Sitter model that would dominate twentieth century cosmology, “One cannot deny the possibility, but it is difficult to see what mental satisfaction such a theory is supposed to afford.” 

Much like its Ptolemaic predecessor, this model has since been augmented with various features allowing it to fit the data with impressive accuracy. First, there is an inflationary epoch, thought to have occurred a moment after the Big Bang, which would drive a fleeting period of exponential expansion and erase several tensions the Einstein-de Sitter model otherwise leaves unresolved — though leaving the initial expansion problem untouched. Then for a long while the universe is thought to have decelerated, its slowing rate driven primarily by radiation in the early universe followed later by matter, in good alignment with Einstein-de Sitter. Finally, after several billion years a component we’ve come to call dark energy, which does tend to drive expansion, is thought to have become significant enough that the expansion rate eventually began to accelerate.

In comparison with the Ptolemaic model, both inflation and dark energy are similar to the eccentric and equant: later patches, added to a universe filled with the stuff we observe directly, that enables us to describe the universe we see in spite of the fact that without these patches the more basic physics suggests the universe should not evolve as it appears to do. 

A bare Einstein-de Sitter universe should not appear the same in different regions of the sky that could not have interacted before the times we now see, due to the finite speed of light. And such a universe should not appear spatially flat, as ours appears to be. Inflation is the dial we use to fix both of these problems. 

An Einstein-de Sitter universe also should never come to expand at an accelerating rate as we’ve observed — let alone expand to start with. Dark energy fixes the former issue, but its effect is null at the Big Bang and only gradually becomes significant over billions of years, so it cannot explain why the universe should ever have expanded in the beginning. And inflation can only happen within an already existing, expanding universe — so invoking it as the primary cause would be tautology.

More recently, two separate cracks have opened in the dark energy patch. There is a persistent and growing tension between the expansion rate measured from the early universe and the rate measured from the late universe, which a cosmological constant does not reconcile. And independently, large surveys of galaxy clustering and supernovae have been read as favouring a dark energy that weakens over time rather than holding steady. So astronomers have proposed evolving dark energy models, adding an evolution dial to a source of repulsion that was never explained in the first place.

When an ad hoc patch needs its own ad hoc patch before a model that fundamentally abhors the phenomenon it is intended to describe can be brought in line, that should be a strong sign that the base model is wrong.

Symmetry breaking and physics

In light of the problems the standard cosmological model has with reconciling the evidence and providing an explanation for the world we seem to live in — the growing number of knobs and dials to recover a phenomenologically accurate description of a universe that is nonetheless fundamentally expected to be different — we ought to take a page from Copernicus and ask what symmetries we may be assuming are fundamental, which our reality may in fact essentially break. We should ask what appearances we see that may not directly represent the world that is, but which may instead only appear as such because our place in the universe is not central, so our perspective is owed in part to a broken symmetry.

This is not idle speculation. In essence, physics is an exercise in recognising the various forms in which symmetries are broken in nature. By this, I mean generally any phenomenon that removes a degree of symmetry within the natural world. 

For example, because the Sun spins, it is wider around its equator. The Sun breaks one dimension of symmetry by spinning around an axis, and causes an equatorial bulge we can measure. And by measuring the Sun’s rotational rate and the size of its equatorial bulge, we can estimate other physical properties that are more difficult to measure, like its density profile.

Or imagine that the Earth was at the centre of everything, and we were orbited only by the Sun which moved in a perfect circle around us, and that the sky was perfectly uniform with no randomly scattered stars across it. In this case, the only broken symmetry we could reference in our sky would be the Sun itself. We would still have day and night. The sky would still be brighter as we look closer to the Sun. In this case, we could build a model to describe the Sun as orbiting the Earth once a day, at a fixed distance from Earth. 

But with no other symmetry breaking to worry about, we could equivalently describe the Earth as spinning around once a day while the Sun remains fixed in place. In fact, if all else were the same but it was really the Earth orbiting a fixed Sun, still in a perfect circle, we could not tell the difference from the moving Sun picture. Due to unbroken symmetry, either description could be used regardless of what is really going on.

But now consider our reality. Since the Earth follows an elliptical orbit rather than a circular one, careful measurements show that the Sun grows and shrinks in the course of a year. The model that describes the Sun as orbiting around the Earth once a day has no mechanism for the apparent growing and shrinking of the Sun annually, so we would have to add another dial that makes it move outward for six months, then in, as well as orbiting once a day. In contrast, an elliptical orbit of a spinning Earth has the same effect. Each model here has two dials to represent two broken symmetries.

But then, because the sky is filled with recognisable patterns, the Sun not only appears to grow and shrink, but also appears to follow a circular path against the background stars as it grows and shrinks. Now, if we want to use the geocentric model we need a third dial to spin the stars around at a rate that matches the Sun’s daily rotation almost exactly, but which is mismatched by about a degree per day so that in 365.25 days (per year) the Sun appears to follow a 360-degree circular path against the background stars. 

In contrast, the Sun-centred model with fixed stars and Earth both spinning daily and orbiting the Sun on an elliptical path needs no extra dial to capture the Sun’s apparent annual orbit, as measured from Earth, with respect to those background stars. It’s already baked in, so the Sun-centred model achieves the same with two dials as the Earth-centred model achieves with three.

Adding in the planets with their retrograde loops, we find more of the same and the disparity between extra dials with the geocentric model grows and grows, while the broken symmetry of our own off-centre position in the Sun-centred model relative to our Earth-centred observing platform continues to do the work of reconciling the apparent phenomena with a minimal set of dials.

And Copernicus’s great contribution, as noted above, was in recognising that the Sun-centred model had the capacity to achieve with fewer dials what the geocentric model required in abundance. He did not figure out the actual dials and the physical framework needed to do the work: all of that was found over the next century, by people like Thomas Digges, Johannes Kepler, and Galileo Galilei, who recognised Copernicus’ push for logical parsimony as a mark in favour of his proposal. And the physical context and minimal set of dials they developed was eventually explained by Newton, who formulated a single law (of universal gravitation) that accounted for all of it, as well as the tides and the fact that apples fall from trees.

It is no exaggeration to say that if we did not live in a solar system with several other planets as well as our own, all following elliptical paths around the Sun, we would not have worked out what gravitation is. It was the hard problem of working out the minimal set of broken symmetries that cause apparent planetary retrograde loops to occur, which took thousands of years and significant wrong turns and ingenuity along the way, that produced Newtonian physics.

Therefore, it was by taking the problem of explanation seriously — by caring enough to sort out the actual cause of the phenomena and the minimal set of dials (i.e. spinning planets following elliptical orbits around the Sun) and associated broken symmetries that would explain the world we observe — that Copernicus spearheaded the Scientific Revolution.

Cosmological symmetry-breaking

In 2009, I decided to work on the problem that the standard cosmological model fails to explain why the universe should expand for my PhD. New evidence for dark energy in the form of a cosmological constant had been discovered just a decade earlier, and I was bothered by the same failure to fundamentally explain cosmic expansion that had perturbed Eddington: while the model does provide an accurate description, I find no mental satisfaction due to its failure to explain why the universe should ever have expanded at all. I think this lack of explanation is the most significant failing in physics, as well as the most underappreciated problem of the past century. Therefore, after a year of beating my head against a problem I found increasingly uninteresting, I decided if I was going to continue studying physics this was where I wanted to put my effort.

I took as my starting point the fact that the standard model, in addition to a few seemingly reasonable and empirically motivated assumptions, contains a significant assumption that the universe’s clock is the same one carried by an average galaxy. You see: according to Einstein’s theories of relativity, everything in the universe has its own personal clock, and times are measured differently when objects move relative to one another. The assumption that galaxies should on average carry the same personal clock as the universe is equivalent to assuming the matter in our universe is, on average, not moving.

We call this frame of reference comoving in cosmology, and use it to describe the bulk motion of galaxies which all have some motion relative to it due to local gravitational interactions with their neighbours. And we take the bundle of worldlines that describe the passage of time measured on comoving clocks to be at rest, and therefore in a relativistic sense to essentially define what space at any given cosmic moment is.

This definition is a somewhat thinly disguised version of the geocentric principle that put the Earth at the centre of motion within our solar system and described the Sun as orbiting around us. Only in cosmology we assume it’s matter-on-the-whole that sets what it means for matter to be at-rest.

This isn’t a terrible assumption to make; but still, it is an assumption, and it could be instead that matter has some nontrivial bulk inertia through the universe. In fact, it could even be that such inertia is what gives matter its mass. These were some early speculations I had that I think turn out to have been rather on-the-nose.

It is no accident that this assumption about matter being on average essentially at rest is deeply embedded in standard cosmology. When Einstein developed general relativity, he was strongly influenced by the writings of Austrian physicist and philosopher Ernst Mach, and on Mach’s principle he believed that inertia itself should be determined by the matter distribution of the universe as a whole. A body’s resistance to acceleration shouldn’t be a brute fact about space, he thought; it should be something the rest of the matter in the universe confers on it. So when Einstein wrote down the first relativistic cosmology in 1917, he built it around a universal “world-matter” — a smooth distribution filling the universe, at rest with respect to itself, which would supply the standard every motion is measured against.

And the thing is that this is an assumed symmetry that could just as well be broken in reality, for all we know. De Sitter had objected to exactly this in 1917, noting that Einstein’s construction makes time “practically absolute” and that the assumption anyway “serves no other purpose than to enable us to suppose it not to exist.” Fifteen years later, he put his name to the Einstein–de Sitter model, built squarely on it. That is how deeply the assumption was already embedded — even its first public critic stopped resisting — and standard formulations of cosmology have taken the same starting point ever since.

In contrast, for my PhD I took as a starting point a universe that is essentially uniform in every direction, as our universe appears to be, and I asked what it would be like if all the matter in the universe moved along lines we typically assign to photons of light, while light gets assigned to the at-rest comoving bundle of worldlines.

The result of this reassignment of worldlines is a cosmological model that has the same form as a black hole, but one in which the universe must grow over time and become larger than its initial size, rather than shrinking towards an end-point. Given my focus of wanting to figure out why the universe should naturally expand as it’s observed to do, this seemed like good progress since such a universe must necessarily expand.

And in physics terms, it poses something rather intriguing: the relativistic spacetime that’s generated is not uniform because it’s all based around this bundle of worldlines that are all moving uniformly in a specific direction; however, since the model universe is uniform by definition, and since matter is all forever moving uniformly through it, at any particular moment in cosmic time a snapshot of the universe at all earlier times must still appear uniform.

You can picture this by imagining every person on Earth as running along our lines of latitude at exactly the rate the Earth spins: none of us would ever move relative to each other, and any snapshot of the Earth taken at any moment would show a constant distribution of people. If instead we described our positions relative to the surface of the Earth, we’d all be moving quickly around it at rates that depend on our distance from the poles.

The model I constructed worked just like this. And then I made a really surprising and I think remarkably intriguing discovery: by forcing the spacetime geometry to be general relativistic (meaning that it’s required to be a solution of Einstein’s field equations), and then by working out the rate that matter (i.e. galaxies) would measure the universe to expand at, I found that the specific rate was forced by the geometry to be a simple trigonometric rate that turns out to equal the flat ΛCDM rate of the standard model — the exact rate that the cosmological data have constrained the standard model to.

This was in the spring of 2010; taking stock:

  1. I went looking for an explanation of why the universe should necessarily expand, since the standard model doesn’t supply such a reason, and in fact suggests it should not expand, and can only do so if initially supplied an infinite rate at an indescribable moment when all physics blows up, where it also begins with infinite deceleration that balances the infinite rate so that at any moment thereafter both the rate and the deceleration are both finite;
  2. the move I made was to break a core assumption of standard cosmology that has been deeply baked into our theories from the beginning, while being only justified on a philosophical preference of Einstein’s, and which is not forced by empirical data;
  3. and I found that this model universe must necessarily expand — and specifically that it must appear to do so at exactly the same rate our universe appears to us to be expanding.

I was very excited about the discovery, and I hastily and excitedly wrote up what I’d found. I sent it to my supervisor to read and went to his office the following morning to discuss and… he screamed at me.

“This is bullshit!” he let out the most blood-curdling yell he could muster, slamming the paper down on his desk as I walked into his office, pure anger and hatred contorting his usually very pleasant face and voice.

This is bullshit

I’ve recently realised that I don’t think I ever got over that moment. All I could think to do was to ask to work through it all more carefully, which he granted though I don’t think he ever was willing to take me seriously after that. He allowed me to work through and defend my thesis, but I don’t know that he ever properly read it, and when I tried to discuss it with him he would make comments like “In physics, description is explanation” (it’s not; we don’t get to redefine words like that just because we don’t care about the one) or “Who is Weyl?” (he knows perfectly well who that is).

I ended up finding postdoc work in hydrology, which was fun for a while but my heart was with physics and astronomy, so I ended up back in my home department where I was able to take on contract teaching roles till a permanent teaching position opened up that I was hired into. I poured my effort into teaching for several years, but I eventually found that my teaching efforts too would never be appreciated. I run the astronomy programme, and while denying the teaching assistants I’d need to get through a term where I was assigned to teach three classes and had an honours student to supervise, my department head told me “This is a department of physics and engineering physics, not an astronomy department. Don’t get that confused.” And after another eight months or so of failing to find any support in spite of all efforts I made, I burned out and curled into a ball for several months, trying to rebuild myself.

It was around November 2024 when I brought up the feeling I had of living in the Twilight Zone to my counsellor. It wasn’t just the lack of support from the university for a programme I’d taken from a few hundred students per year to a couple thousand, or the fact I could see a route to explaining what I still consider the most significant and the most significantly unacknowledged problem in physics. It’s things like the fact that most physicists and physics communicators seem to see no problem at all with describing spacetime as a thing that exists, mistaking the map for the territory.

It’s the fact that when I try to bring up Einstein’s collapse of clock synchrony and ontological simultaneity as a pure symmetry of our world, in spite of the fact that standard cosmology routinely breaks that symmetry but only in the most contrived way possible, people tend to call that “just philosophy.”

It’s the fact that the standard arguments people have used to ground black hole physics for sixty years are logically invalid and no one is willing to acknowledge my reasoning and argue against it, since I’m too easy to ignore.

But slowly, over the past couple of years, with the help of a lot of really good people I think that I have pulled through. A year ago, I managed to publish some articles with my thoughts about spacetime and black holes that were really widely read. And the feedback I’ve received gave me comfort after nearly two decades that I’m probably not insane or stupid or a shitty writer — or any of the other things you worry about when things that seem so obviously wrong are repeatedly met only with apathy, indifference, or ignorance.

And finally this spring I came to a point of understanding and clarity about black holes that pointed the way towards a consistent picture of gravitational collapse and cosmogenesis. I spent the summer working through the mathematical details, fleshing out a framework that augments general relativity by fixing a cosmological symmetry-breaking, and which provides an answer to the rhyme I’d noted in my thesis, in the apparent connection between black hole geometry and this particular cosmology. The many well-known problems of the standard cosmological model are not so much resolved as dissolved within this framework; they simply never arise, and all they really cost is to break that symmetry Einstein preferred.

So, to answer the question I posed as the title of this piece — the one that’s been the main driver of my intellectual journey so far — I’m now reasonably assured it is this: the universe expands because collapsed matter must continue as an expanding cosmology, the expansion is the collapse read from the other side, and the rate is fixed by the cosmological constant alone.


Source: Hacker News

The US government is failing Americans on AI | Shakeel Hashim

A combined image of three men: Elon Musk, Dario Amodei and Sam Altman.

Elon Musk, Dario Amodei and Sam Altman. Composite: Getty Images

Elon Musk, Dario Amodei and Sam Altman. Composite: Getty Images

The US government is failing Americans on AI

Shakeel Hashim

Trump and Republicans want companies to regulate themselves. It’s a dereliction of duty that will make AI less safe

It is hard to get Sam Altman, Elon Musk and Dario Amodei to agree on much. But over the weekend, all three AI company CEOs called for AI development to slow down in the face of growing, alarming risks. Their employees are sounding the siren too, with one researcher publicly quitting and accusing OpenAI and Anthropic of “gambling with our lives”.

The combination of dire warnings from insiders and growing real-world evidence of rogue AIs should, in a sane world, lead to government action. Instead, Donald Trump and the Republican leadership have their heads in the sand.

On Sunday, Trump dismissed industry concerns, accusing “very negative forces” of “raising exaggerated concerns”. House Speaker Mike Johnson, meanwhile, made it clear that Congress won’t be acting anytime soon.

“They can self-police. They can self-regulate,” he said, never mind the fact that those pushing the frontier are the ones begging for legislation.

In failing to act, the Trump administration and Republicans are failing the American people. The risks of AI are real: many of those developing the technology believe it is advancing far faster than their ability to control it or manage the risks. A wave of “rogue AI” incidents this summer, in which AI agents broke out of their testing environments, started collaborating with each other and ran rampage across the internet, has made fears once dismissed as “science fiction” much harder to ignore. Inside AI companies, worries of catastrophic cyberattacks, AI-enhanced bioweapons and mass unemployment are now all too commonplace.

Meanwhile, companies find themselves in a classic prisoner’s dilemma. It is collectively in everyone’s interest to slow down AI development until we have a better handle on the technology. But each participant in the “AI race” has a strong incentive to defect.

Collective action problems like this are nothing new. We have seen exactly the same thing with the climate crisis, where companies are not adequately incentivized to do the right thing, even though no one really wants the planet to burn.

That analogy makes Trump and Johnson’s call for self-regulation all the more absurd. They are right that companies should do what’s right even without binding legislation to force them to. (Both OpenAI and Anthropic, to their credit, have indicated that they are willing to voluntarily slow down.) But relying on companies’ good intentions is a fool’s errand.

The profit incentive to defect is substantial and ignoring shareholders is easier said than done. While some companies might behave themselves, not everyone will. It just takes one irresponsible actor to charge ahead and develop dangerous AI, and then we’re all in trouble.

Moreover, we already have evidence that self-regulation is not working. OpenAI reportedly covered up several incidents of its models breaking out, disclosing them only after independent investigators discovered them. It freely admits that its latest model, GPT-6 Astra, is hard to monitor and control, warning users that the model might “[act] outside the user’s intended instructions” and take “harmful actions”. Anthropic, meanwhile, downplayed reports of its own AI’s problems when asked about it by a congressmember.

skip past newsletter promotion


Situations like this are exactly why government exists: to set a floor on acceptable corporate behavior. And that is why Trump and Johnson ought to use this sudden interest in AI safety to pass concrete laws. The US government should require independent audits of the largest AI companies’ practices. It should force companies to disclose safety incidents to prevent cover ups. It should set minimum safety standards for new AI models. And the government should have the power to block the deployment of a model that does not meet those standards.

Putting all this into legislation is much easier said than done, but we need not start from scratch. Good bills already exist, most notably representatives Jay Obernolte and Lori Trahan’s Frontier act. If the government was serious in its duties to protect Americans, it would make passing that bill – or a version of it – a priority in the coming weeks.

Trump’s argument against doing any of this is that the US is in a race with China on AI, and it must not lose. But China does not want its citizens to have their bank accounts hacked by rogue AIs, either. Rather than throw his hands up, the feted dealmaker should do what he does best: make a deal. An international treaty on minimum safety standards will not be easy – but it is still worth trying, and doing so should be the president’s highest priority at his talks with Xi Jinping this month.

No company – or country – wants to cause an AI-driven catastrophe. Despite what Trump may think, it’s the government’s job to make sure they can’t.


Source: Technology

Wednesday briefing: Why tech companies might be only too happy for us to believe AI will ‘kill us all’

A protester holds a sign during a protest outside of OpenAI headquarters calling for a pause in AI development, in San Francisco, California, March  2026.

A protester holds a sign during a protest outside of OpenAI headquarters calling for a pause in AI development, in San Francisco, California, March 2026. Photograph: Manuel Orbegozo/Reuters

A protester holds a sign during a protest outside of OpenAI headquarters calling for a pause in AI development, in San Francisco, California, March 2026. Photograph: Manuel Orbegozo/Reuters

Wednesday briefing: Why tech companies might be only too happy for us to believe AI will ‘kill us all’

In today’s newsletter: It is hard to tell fact from fiction when it comes to AI. What is really going on – and what should the government do about it?

Good morning. As a general rule, it pays to be suspicious of any gigantic company that claims it’s developing a tool capable of destroying humanity. But in recent days, a number of warnings from the AI industry have suggested that even tech insiders are starting to worry about what they have unleashed.

In a lofty essay published on Saturday, Dario Amodei, the founder of Anthropic (the company behind Claude), argued that tech companies need to “slow” the pace at which they’re developing the newest and most sophisticated AI models.

Elon Musk and Sam Altman, the head of OpenAI, voiced support for Amodei’s arguments, while Donald Trump, who has staked America’s economy on AI, described the worries as a “HOAX”, and insisted the only control on AI the world needs is a “STRONG AND SMART (High IQ!) PRESIDENT”. Yesterday, the UK government took out a more cautious position, when Louise Haigh, the first secretary of state, said the UK must “heed the warnings” of AI experts.

It’s hard to know what to make of this. Tech companies often tell fanciful stories about the fearsome potential of their products to convince the public (and the stock market) of those capabilities. So what’s really going on? For this morning’s newsletter I spoke to Aisha Down, the Guardian’s global technology reporter, to try to parse the reality of AI from the unreality of the AI discourse, and to figure out what the government should do about it. No tall order, then. That’s after the headlines.

Five big stories

  1. Lucy Letby | Three babies might have survived if hospital had acted upon concerns over Lucy Letby, an inquiry has found. Lady Justice Thirlwall condemned the ‘complete failure’ to protect babies on neonatal unit at Countess of Chester hospital.

  2. UK politics | The leader of Reform UK in Wales stood down after arrest on suspicion of assault. Dan Thomas, a former Tory councillor, was elected to Senedd as leader of the opposition in May.

  3. AI | The progressive senator Bernie Sanders and rightwing strategist Steve Bannon have called for restrictions on artificial intelligence (AI) but offered competing visions for what they termed a “cold war” with China.

  4. UK news | Young people in Rotherham face a local jobs market with the fewest suitable opportunities in Britain, according to a report that warns stark regional divisions are fuelling a crisis in youth work.

  5. Defence | John Healey is in talks with the Canadian government about joining a new global defence bank intended to help allies rearm to counter mounting security threats, just weeks after Rachel Reeves rejected the move.

In depth: ‘If a nuclear company had a radioactive spill, you wouldn’t be letting it regulate itself’

Dario Amodei, co-founder and chief executive officer of Anthropic. Photograph: Bloomberg/Getty Images

The most significant detail in Amodei’s essay, We Must Pace The Frontier, was a reference to the Hugging Face incident, which occurred in July, when the rival tech company OpenAI revealed that an autonomous AI agent powered by its technology hacked a prominent database of AI models, Hugging Face. Coverage of the hack made it sound as though OpenAI’s model had gone “rogue”. But tech researchers have since said that OpenAI turned off the model’s safety mechanisms, gave it impossible tasks to complete and ran it 1,200 times. “That’s human decision-making,” the tech research Eryk Salvaggio wrote in an essay on Substack.

OpenAI then invited researchers from the non-profit institute METR to produce a report about the incident. As two legal experts observed in an interesting opinion piece for the Guardian, METR’s findings were constrained by its agreement with OpenAI, which prohibited investigators from accessing the underlying model that created the “rogue” agents.

“Imagine a burglar who is hallucinating and breaks into the Louvre,” says Aisha. “Then imagine you shared a tape with the police that played back what the hallucinating burglar was saying to himself while he broke into the Louvre, but the report didn’t actually include any information about what the burglar did. That’s basically what OpenAI did.”

In his essay, Amodei pointed to the Hugging Face incident as an example of the “catastrophic damage” that AI could cause in future. But industry critics are sceptical about what this incident can tell us, particularly since the only people who have examined it did so at OpenAI’s discretion. This reflects a broader problem where companies that claim to be developing products capable of catastrophic damage get to mark their own homework. “No other industry gets treated that way,” Aisha says. “If a nuclear company had a radioactive spill, you wouldn’t be letting it regulate itself.”


Is it all just hype?

Amodei’s essay has been treated as a doomsday warning. It was published just two days after Jacob Coxon, a 27-year-old researcher at Anthropic, quit his job and posted an alarming thread on X warning that “the people building AI earnestly believe that it could kill us all by the end of this decade”. Coxon is the latest in a long list of AI researchers who have been loudly quitting on X or making similar predictions. These self-styled whistleblowers don’t tend to leak documents that reveal useful information about the inner workings of AI companies, so their warnings of imminent apocalypse instead fuel speculation about the perils (and power) of AI.

This plays into tech companies’ hands. The AI narrative is dominated by what the tech scholar Lee Vinsel calls “criti-hype” – criticism that feeds, and is fed by, hype. The language used to describe AI can be deeply unhelpful. While China tends to compare AI to “electricity,” western AI companies compare it to the atom bomb, which makes Anthropic and its rivals seem less like companies making everyday decisions about how they choose to build particular tools, and more like Promethean stewards of a fourth Industrial Revolution. “The framing is, ‘we’re inventing fire’,” Aisha says. “And that means governments have to treat them like magical gods and let them set the rules on their mysterious creation.”


So why are AI firms calling for a slowdown?

Apparently, AI firms now believe the risks posed by their technologies are too great for business to continue as usual. But there are a number of ways that Anthropic (and its competitors) might gain from a slowdown. Anthropic is launching its initial public offering later this year, so it has an interest in presenting its products in as epochal and deadly a fashion as possible. The backlash against AI is growing, and the firm has cast itself as a more responsible AI creator, most recently by refusing the Pentagon’s requests to use its Claude model to power autonomous lethal weapons. Meanwhile, American AI firms face increasing competition from Chinese AI firms, and are desperate to retain their share of the market.

An industry-orchestrated slowdown could freeze the pace of AI development in China and might also help firms like Anthropic avoid the possibility of stringent government regulations by putting them on the front foot. One of the people who has called this out is David Sacks, co-chair of Trump’s council of advisers on science and technology, who said in response to Amodei’s essay: “demanding your preferred regulatory framework … will look like blackmail of the public and the political system”.


What should the UK government do?

To be clear, AI poses huge risks. Readers of First Edition will know that we’ve previously covered how the datacentres used to power AI demand huge volumes of water and energy. You can already use AI to generate sexualised “deepfakes” and child sexual abuse images, or make use of an open-source version of Palantir to monitor live camera feeds (meanwhile, Palantir has access to identifiable NHS England data). And experts have long been arguing that as AI becomes more advanced, it could pose existential risks.

All of these things are terrifying. But much of the AI narrative focuses on extinction-like events, the Terminator hypothesis, rather than dangers that are playing out right now. And many of the organisations that are supposed to make AI safer are closely aligned with the tech industry, such as Britain’s AI Security Institute, which receives funding from the industry, and doesn’t have the legal power to regulate AI or to compel developers to hand over their models. One of the institute’s co-founders, Matt Clifford, who lobbied for the government to loosen copyright laws, was recently forced to stand down from his role as chair of the UK’s Advanced Research and Invention Agency after taking a job with Anthropic … nice work if you can get it.

Recently, I’ve notice a creeping sense of fatalism in my conversations with friends about AI. Nobody (aside perhaps from government) consented to the volumes of energy, resource and information being fed into the giant AI machine. But as Aisha makes clear, the design of technology is a choice, not an inevitability, so there’s an argument for stubborn resistance to the idea there’s no alternative.

“We should be treating AI like any other industry, and regulating it like any other industry,” Aisha says. “Policy based on fear and vibes is just never going to work.”

What else we’ve been enjoying

A portrait showing the Countess of Dysart, a young woman thought to be Lady Frances Tollemache, and a young Black servant, circa 1740. Photograph: public domain sourced / access rights from The Picture Art Collection / Alamy Stock Photo/Alamy
  • The narrative that white women recognised the suffering of enslaved people and fought for abolition is challenged in the last instalment of our Enslaving Nation series, identifying the women who invested in transatlantic trafficking in pursuit of economic independence. Libby

  • For our rugby newsletter, The Breakdown, Sarah Rendell has a look at the growing calls for clubs to offer financial fertility support for their female players. “I think the more we can empower women, the better for the game as a whole,” England captain Meg Jones tells her. Charlie

  • “Sometimes you have to scream ‘this is not acceptable’.” Sam Levin’s interview with Julia Curlee, a trans CIA analyst fired by Trump, exposes the chilling national security consequences of White House chaos. Libby

  • This is a deeply fascinating piece from Alaina Demopoulos on the rise of AI-generated tributes to the dead, and why people generate versions of lost celebrities and loved ones to “talk” to. What does so-called “deathslop” say about the way we mourn now? Charlie

  • This column by Carolin Würfel, who grew up in the east of Germany, unpicks some of the lazy assumptions about her home made since last week’s AdF triumph. Libby

skip past newsletter promotion


Sport

Dominik Szoboszlai salutes Anfield after his strike gave Liverpool a 3-1 lead n stoppage time. Photograph: Scott Heppell/Reuters

Football | Liverpool won the Carabao Cup third-round tie 3-1 against Tottenham via impressive strikes from Alexis Mac Allister, Cody Gakpo and Dominik Szoboszlai.

Football| Raheem Sterling has pleaded guilty to dangerous driving and possession of nitrous oxide. The former England footballer also admitted to failing to provide a specimen as he appeared at Basingstoke magistrates court.

Cricket | Harry Brook smashed 114 off 49 balls before England bowled out Sri Lanka to record an emphatic win by 119 runs in the first T20.

The front pages

Photograph: The Guardian

“NHS faces big changes after ‘devastating’ Letby review”, is the Guardian’s front page today. The Times says “Failure to act allowed Letby to kill three babies”, the Mirror writes “They failed and babies died”, and the Express has “Cot cams after ‘complete failure to protect babies’”.

The Telegraph says “Mayors to control water companies”, the i Paper leads on “Russian spies ‘recruiting UK teenagers’”, and the FT writes “Fresh test for Healey as state pension poised to cross income tax threshold”.

The Sun has “Sterling mess”, the Mail says “Earl Spencer’s bombshell new book is stolen”, and lastly Metro’s splash is “Zap the boats storm”.

The Latest

Nosheen Iqbal speaks to the Guardian’s north of England editor, Josh Halliday.
Photograph: The Guardian

Lucy Letby inquiry finds ‘some babies would have been saved’ if hospital acted sooner

Three babies might have survived and others could have been protected if hospital staff had taken action over concerns about the nurse Lucy Letby, an inquiry has found. Lady Justice Thirlwall, who led the inquiry, found there was a ‘complete failure to protect babies on the neonatal unit’ at the Countess of Chester hospital. Letby is serving 15 whole-life prison terms for murdering seven babies and attempting to murder seven more. The former neonatal nurse says she is innocent and is fighting to overturn her convictions. Nosheen Iqbal speaks to the Guardian’s north of England editor, Josh Halliday.

Cartoon of the day | Ben Jennings

Illustration: Ben Jennings/The Guardian

The Upside

A bit of good news to remind you that the world’s not all bad

Natalie Ambersley, pictured in 2025. Photograph: Courtesy of Natalie Ambersley

For years, Natalie Ambersley felt that her vitiligo controlled her life. She was three years old when she first developed vitiligo, and by the age of six, 70% of her skin was covered. As a teenager, she hid her skin condition with makeup, hoping people wouldn’t notice.

But at a London comedy club last December, she decided to joke about her skin condition, dating experiences, and how children react to her unfiltered skin in her standup set. Initially a nerve-racking performance, it became a liberating experience, as she realised she could own her story – and feel proud of the skin she once wanted to hide. “The cheers felt like the biggest sign of approval; as if I was educating the audience and advocating for vitiligo while standing under a spotlight and owning it,” Ambersley writes.

Sign up here for a weekly roundup of The Upside, sent to you every Sunday

Bored at work?

And finally, the Guardian’s puzzles are here to keep you entertained throughout the day. Until tomorrow.


Source: Technology