Preloader

Technology

Pirate Face Rescues LLM Models from Deletion

The permanence layer for sovereign AI.

Pirate Face is decentralized infrastructure for sovereign AI. Every open model is mirrored from Hugging Face as a torrent, held peer-to-peer instead of by a single company.

Why models should never die →

Censorship-resistant

Every model is a torrent that also downloads straight from Hugging Face. The day it’s gone, the swarm keeps it alive – there’s no single host to shut down.

huggingface.co/modelremoved
fallback ↓
pirateface.co/model1,240 seeding
Checksum-verified

Every file carries its official Hugging Face SHA-256 – so you verify every byte. Download it anywhere and the hash still has to match: the real weights, never a tampered copy.

model-00001.safetensors6.7 GB
SHA-2569f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08
✓ Matches Hugging Face – bit-for-bit, untampered
Drop-in API soon

Point your existing pipeline at Pirate Face – same paths, same API, just a different endpoint. Zero code changes.

pirateface – zsh

# same pipeline – one env var
$export HF_ENDPOINT=https://pirateface.co
$python train.py
✓pulling meta-llama/Llama-4 from the swarm

Claim your handle

Reserve pirateface.co/yourname for your public profile. Verify your matching Hugging Face identity to earn the Verified creator badge and help prevent impersonation.

Join the community for upcoming account benefits, including free compute credits and exclusive model releases.

Create account

zai-org ✓ verified
pirateface.co/zai-org

Z.ai · Zhipu AI – GLM foundation models, Beijing
LLMzai-org/GLM-5.2🌱 1,800
LLMzai-org/GLM-4-9B🌱 1,240
VISzai-org/CogVideoX-5B🌱 760

FAQ

What is Pirate Face?

A decentralized, peer-to-peer layer for sovereign AI. Open models from Hugging Face become checksum-verified torrents, held by a global swarm so they survive any single host taking them down. The goal is permanence for open-source AI.

Do I need an account to use Pirate Face?

No. You can browse, download, and seed without an account.

  • Claiming is optional. A public claim reserves your Pirate Face handle and helps the network spread. Only your first eligible handle earns welcome points.
  • Hugging Face creator claims require verification. To receive the Verified creator badge for a current Hugging Face username or organization, you must verify the matching Hugging Face account. Until then, the handle is only reserved and does not prove identity.
  • Direct publishing is planned. Accounts will eventually support publishing and managing models directly on Pirate Face, without first uploading them to Hugging Face. That is not live yet.

Today, submitting a model requires a Hugging Face account and the model must already be on Hugging Face.

Accounts also support upcoming community benefits. See account benefits.

Why create a Pirate Face account?

An account gives you a place to manage your claimed handles and track your contributions. Claiming reserves a public handle; verifying the matching Hugging Face identity adds the creator badge and helps prevent impersonation.

Pirate Face is preparing community benefits for account holders, including free compute credits and exclusive model releases.

You can browse, download and seed without an account.

Why use Hugging Face verification if the goal is independence from centralized hosting?

Hugging Face verification confirms that someone claiming a creator’s handle controls the matching HF account or organization. It helps prevent impersonation and preserves attribution as models spread beyond their original host.

You don’t need verification, or an account, to browse, download or seed.

Pirate Face’s short-term goal is for model availability to outlive any single hosting platform. Today, parts of Pirate Face depend on Hugging Face and its community. Pirate Face is working to allow direct model publishing soon, without first uploading to HF.

Does my Pirate Face handle have to match my Hugging Face username?

No. You can reserve an available name, but the verified badge requires a matching Hugging Face identity.

For example, a Hugging Face account named john123 can verify john123, not john.

You can still reserve john, but it stays unverified. The matching Hugging Face owner could reclaim that handle without deleting your Pirate Face account.

Manage your connected accounts and reserved handles in Account.

Could someone impersonate a creator here?

Both are protected:

  • The weights can’t be faked – every file is checksum-verified against Hugging Face’s official SHA-256.
  • The handle – anyone can reserve a name (it shows reserved, no Verified creator), but only proving the matching Hugging Face identity earns the Verified creator. A squatter can’t lock you out: the real owner reclaims a reserved name by verifying.
What does “checksum-verified” mean?

Every weight file carries its official Hugging Face SHA-256 – a unique fingerprint of the exact bytes. Download it from anywhere and the hash still has to match, so you always know it’s the real weights, never a tampered copy. It’s the #1 fear with mirrored models, so it’s a first-class feature here.

How do downloads work, and what if Hugging Face removes a model?

Every model is a magnet link (a torrent) with a Hugging Face web-seed built in – so it works even with zero peers.

  • While it’s on Hugging Face: you pull the bytes straight from HF – same weights, checksum-verified, same speed.
  • The day HF removes it: the web-seed dies and the download falls back to the peer-to-peer swarm. We mark it Rescued – still reachable, kept alive by whoever’s seeding.

That fallback is the whole point: the same bytes as downloading from Hugging Face directly, but with no single point of failure (and the swarm can serve far-from-HF regions faster).

Can I add my own model?

Yes – if you have a surviving copy, submit a magnet with source evidence. Opening a model page already records Hugging Face checksums when they exist. A listing is not a download, and listing does not start seeding.

  • Sign in and submit a peer-only magnet, pinned revision, file checksums, and license evidence. MIT and Apache-2.0 only, plus the approved Kimi-K3 exception.
  • Matching Hugging Face LFS hashes at the pinned revision lists immediately. Mismatches or a gone source stay in the review queue.

Pirate Face mirrors from Hugging Face, so the model has to live there first. If yours isn’t on HF yet, upload it there (free) and then submit it here.

What’s a “web-seed”?

The one bit of jargon worth knowing. A web-seed (BitTorrent spec BEP-19) is a plain HTTPS URL built into a torrent – here, the model’s file on Hugging Face. It’s really just the download link, carried inside the torrent as a guaranteed source – which is why models download even with zero peers, and why the swarm takes over the moment that link dies.

What’s the drop-in API?

Set HF_ENDPOINT=https://pirateface.co and your existing pipeline resolves models through us – straight from Hugging Face while it’s up, from the swarm the moment it isn’t. Same paths, same API, zero code changes. soon

What are points?

Points recognise participation on the leaderboard. They are not money or a compute-credit balance.

  • Welcome points: one qualifying handle per account. Extra reservations and linking HF do not earn another award. Connect your claim to an account to qualify.
  • Referrals: 50 points for a new account’s first handle, up to 10 credited referrals per referring account. Self-referrals and extra handles do not count.
  • Rescued models: 25 points per distinct rescued model, credited to its first verified handle.
  • Seeding rewards: planned, not active. We need attributable seeding evidence before awarding them.

Pirate Face is building an ecosystem of contributors, with planned community benefits including free compute credits and exclusive model releases. Eligibility, limits and launch dates will be announced before benefits become available. Points do not automatically convert to credits, and reserving more handles does not increase your benefits.

There is no token. Watch our official X for announcements.


Source: Hacker News

Amiga Unix, Again

Amiga Unix, again

Amiga Unix — Amix — was Commodore’s System V Release 4 for the Amiga: shipped in 1990–92 for the
A2500UX and A3000UX, then left where it stood. amigaux.org is an unofficial community project that
picks it up again: Amix on 68040 and 68060 machines and on today’s accelerator hardware, with a
modern toolchain, a package manager, and drivers for cards that never had any. The work is done in
the open, and written down as it happens in the
grimoire.

Soon: install it yourself

The install medium boots from floppy and CD-ROM on real hardware — floppy image and ISO under
emulation — and takes care of the whole installation, interactively or unattended. With supported
hardware, getting Amix onto a machine comes down to following the prompts. The classic tape
install follows later.

From there the network takes over: apkg installs packages straight from
pkg.amigaux.org — grep, gzip, less, patch and zlib to start with, and more GNU and BSD
tools as they are built and tested.

# apkg update
catalog updated: 39 packages
# apkg install less
downloading extra/less-704.pkg …
installed less-704

Where we are

Working today

  • 68040/68060 support — Amix 2.1 on a real 68060, with or without FPU; an MMU is mandatory.
    The kernel: asokero/amix-040-060-port
  • Z3660 accelerator — native SCSI and ethernet drivers, proven on real hardware
  • A4091 / A4092 — a Zorro III SCSI driver and an auto-detecting kernel
  • Installation — from floppy + CD-ROM on a real machine, or floppy image + ISO image under
    emulation, interactive or unattended; the classic tape install follows later
  • apkg — a remote package client with a hosted repository (catalog, dependencies, upgrades),
    and the first modern GNU tools packaged: grep, gzip, less, patch, zlib
  • A cross toolchain — Linux-hosted m68k-cbm-sysv4, building all of the above
  • An OpenLook desktop that comes up ready on a fresh install
  • Quake — runs; a benchmark more than a game, for now

In progress

  • RTG graphics for the Z3660 (ZZ9000-compatible) — running under emulation, real hardware next
  • Ethernet throughput in the accelerator firmware
  • X11R6.3 and Mesa as packages, from the community ports; the wider userland after that
  • A read-only CD filesystem (ODFileSystem port)
  • Install-media polish: a repair path, guard rails; the tape version

Next

  • A native, modern gcc on the box
  • RTG on the real accelerator
  • The packaged X11R6.3 / Mesa stack
  • The firmware ethernet fixes
  • Then the full launch here: packages, source, instructions and manuals

How it’s built

Part of this work is done with generative AI in the loop — models reading the old kernels in binary, as no
source is available, and writing drivers and notes — with people setting direction, reviewing every change and testing on real
hardware, where it either boots or it doesn’t. Other parts are done the classic way. The grimoire
records both, confidence-tagged, so you can see what is verified and what is still a guess.

The code

The kernel

Drivers and hardware

Graphics and X11

Built on

Tools and docs


Source: Hacker News

Frontier Labs Are Selling Garbage to Fools in Washington


Source: Hacker News

Sherline Tools Is Going Out of Business

If you buy something through our links, ToolGuyd might earn an affiliate commission.


Sherline Precision Tools for Precision Work with Lathe

Sherline Tools, a USA manufacturer known for precision lathes, mills, micro-machining accessories, and more recently small CNC machines, has announced that they are “winding down manufacturing operations.”

Following are some highlights from Sherline’s recent message to their customers.

Sherline Products, Inc. has begun the process of winding down manufacturing operations

Advertisement

When we took over the company in 2017… we invested in new products and designs, modernized our computer systems and manufacturing processes

Unfortunately, the manufacturing environment has changed dramatically. The effects of COVID, increasing manufacturing and operating costs, and significant changes in consumer purchasing habits have made it increasingly difficult for a small American manufacturer such as Sherline to maintain the workforce and production levels necessary to remain competitive.

we have reached the very difficult conclusion that continuing manufacturing operations is no longer sustainable

We plan to continue building and selling machines, tooling, accessories, and replacement parts to the extent our remaining equipment, materials, staffing, and inventory allow through the end of October [2026].

Some manufacturing operations, particularly those requiring our larger production equipment, will necessarily end sooner

some products may become unavailable before others

Advertisement

We plan to maintain an online presence and make replacement parts available as inventory permits. We will also continue to address warranty issues in accordance with our warranty obligations.

Sherline emphases that they won’t simply disappear or leave customers without support. They also say they will maintain their archive of technical and educational information and resources.

They also mention saying goodbye to employees.

Production is coming to a halt. Equipment is being “removed from service,” which sounds a lot like a liquidation sale. It seems Sherline is effectively shutting everything down and is likely closing the doors.

Update: one of the owners posted to a Facebook group page, saying:

We are not sure of the actual final day, but it would suffice to say that by the end of the year at the latest, Sherline will no longer be in business.

Such sad news for a storied USA micro-machining brand.

<!–

–>


Source: Hacker News

What happened to the Snowden archive

The last document from the Snowden archive was published on 29 May 2019. The Guardian stopped publishing documents in February 2014, Der Spiegel in January 2015, and The New York Times and ProPublica in August 2015. After that, only The Intercept was still publishing documents, with only a few exceptions, until it closed its archive in March 2019. Eleven weeks later, on 29 May 2019, it released what would become the final batch of documents from the archive. Since then, no news outlet, journalist, or institution anywhere has published a single document from the Snowden archive.

Snowden started reaching out to journalists months before having any material to share. In December 2012, he contacted Guardian columnist and former attorney Glenn Greenwald, requesting that he set up a secure communication channel, though Greenwald didn’t know how to do so.[1] After a second attempt in January failed to make progress, Snowden shifted his focus to Laura Poitras, an American documentary filmmaker. Poitras then teamed up with an investigative journalist Barton Gellman.[1]

Snowden shared the archive gradually. On 31 March 2013 he sent Poitras a link to an encrypted file called astro_noise,[2] which she downloaded and filmed herself doing so, but had no key to open it.[3] On 10 May Snowden posted a box from Hawaii to a Brooklyn address he had asked her to supply. It went to Jessica Bruder, a journalist who had agreed to receive a package without being told what it was, and who passed it unopened to Dale Maharidge, and then Maharidge handed it to Poitras on 15 May.

On 21 May 2013 Snowden sent Poitras and Gellman the keys to an encrypted archive called Pandora containing about 50,000 documents.[4][5][1]

At some point in late May, Snowden warned Poitras about what he called a “single point of failure,” and urged her to get copies of the material into other hands. She distributed three. One went to Trevor Timm of the Freedom of the Press Foundation, with a note asking him to hold the material and not to give it to anyone but Greenwald, and only if Greenwald asked in person. One went to a person who has asked not to be identified. The third went to someone who remains unknown. Maharidge kept a copy himself, saying in 2017 that he still had it.[6] Aside from Poitras, Gellman, and Greenwald, it’s unclear whether the others holding copies can actually read the documents, since the copies are likely encrypted and they may not have the key.

Greenwald received the Pandora copy from Poitras on 1 June 2013, on the way to the airport,[1] and read it for the first time on the flight to Hong Kong.[7] The Guardian journalist Ewen MacAskill was handed the GCHQ material in Hong Kong,[1] and brought it to the Guardian‘s London office.[8] Der Spiegel obtained their archive from Poitras in that summer.[9] The New York Times and ProPublica got theirs from the Guardian.[10][11]

Greenwald said in 13 March 2019 that he and Poitras “individually and independently, continue to possess full copies of the archive, as do other individuals and institutions.”[12]

Poitras said in 2022 that the archive “still exists, and there is still more to report,” describing “a vast amount of information that hasn’t been reported of enormous contemporary and historical significance.” Greenwald said in 2023 that the archive runs to “hundreds of thousands of documents, if not more.” But neither has published anything from it in seven years, and neither has explained why.

The Guardian

Snowden gave a large portion of the archive to the Scottish journalist Ewen MacAskill in Hong Kong, containing “tens of thousands of documents.” In Citizenfour, Snowden is filmed handing GCHQ material to MacAskill in the Hong Kong hotel room. Snowden describes it as coming from the agency’s internal wiki — “at the top secret, you know, super classified level, where anybody working intelligence can work on anything they want” — and adds: “That’s what this is. I’m giving it to you. You can make the decision on that, what’s appropriate, what’s not.” MacAskill brought his copy to the Guardian‘s London office.[8] The Guardian later stated that their archive contained around 58,000 documents.

The Guardian published its first Snowden story with a document from the archive on 6 June 2013. The following day, UK defence officials issued a confidential D notice to every major British media outlet “in an attempt to censor coverage of surveillance tactics employed by intelligence agencies in the UK and US.”

The DPBAC committee (the Defence, Press and Broadcasting Advisory Committee) issuing D notices runs what was then called the DA-Notice system, a voluntary arrangement dating to 1912, under which news organisations consult the Ministry of Defence before publishing material touching national security. In 2013, the serving secretary was Air Vice-Marshal Andrew Vallance.

The D Notice was issued to a press already inclined to cooperate. Reviewing the academic literature in 2018, Vian Bakir and Andrew McStay from Bangor University noted that:

Many longitudinal studies of British press and broadcasting coverage across the two-year period following Snowden’s initial revelations find most outlets privileging political sources seeking to justify and defend the security services and mass surveillance.

According to Bakir and McStay the alignment was close to uniform. The Daily Mail, Mirror, Star, Telegraph, Sun and Times were coded pro-surveillance. Only the Independent, the i and the People leaned the other way. The Guardian and the Express were balanced. What that meant in practice, they write, was that “citizens’ privacy rights and surveillance regulation were minimally discussed, while mass surveillance was normalised by suggesting that it is necessary for national security.” Prominent press themes “were that social media companies should do more to fight terrorism, and that while surveillance of politicians is problematic, surveillance of the public should be increased.”

The Daily Mail – which published more articles on Snowden than any other British outlet besides the Guardian – is a prime example of this trend. In the six months following the initial leak, the outlet concentrated on Snowden’s “lavish lifestyle on the run and his sexual appeal.” As Bakir and McStay note, these tactics appeared designed to discredit or draw attention away from the GCHQ revelations he exposed.

Following pressure from Downing Street and Cabinet Secretary Jeremy Heywood, three Guardian executives destroyed the computers holding their London-based archive in a basement on 20 July 2013, six weeks after their first Snowden story. Two GCHQ technicians supervised the destruction. The paper kept the action secret[11][13], only making it public a month later on August 19.

Paul Johnson, then deputy editor of the Guardian and one of the three who destroyed the computers, later called the destruction “purely a symbolic act”, because the government knew “that the material had been taken to the US to be shared with the New York Times. The reporting would go on. The episode hadn’t changed anything.”

The DPBAC committee’s own minutes record something happening in the same weeks. Vallance told that “at the outset the Guardian had avoided engaging with the DA-Notice system before publishing the first tranche”, that as a member of the Newspaper Publishers Association it “was obliged to seek (but not necessarily to accept) DA Notice advice under the terms of the DA Notice code,” and that this failure “was a key source of concern and considerable efforts had been made to address it.” Towards the end of July 2013, he said, the Guardian “had begun to seek and accept DA Notice advice not to publish certain highly sensitive details.”

The Guardian‘s then editor-in-chief Alan Rusbridger was publicly questioned on 3 December 2013, in the Home Affairs Committee, about the Guardian‘s Snowden reporting. “In terms of publishing documents, I think we have published 26,” Rusbridger said, “I would not be expecting us to be publishing a huge amount more. With 26 over six months, I would say it has been a trickle.” Rusbridger said that during the last six months “there have been more than 100 contacts” with UK and US government agencies.[14] “We were in continuous contact with the intelligence agencies and the White House,” Rusbridger writes in 2018.[15] When asked about how many of the 58,000 documents were read, Rusbridger replied “I could not tell you. I don’t know.”[16]

Rusbridger told the Home Affairs Committee that of roughly 35 stories published, the Guardian had consulted the authorities on all but one. Of Vallance he said: “We have, in fact, collaborated with him since and he has been at The Guardian to talk to all our reporters.” Rusbridger added in 2018 that Vallance was in time invited into the Guardian‘s morning conference.[17] Rusbridger’s account of the arrangement is that Vallance was reviewing stories as they came and rarely objecting.[18] Gellman wrote that the Guardian “killed some of its stories for legal reasons and delayed others for months.”[19]

Three months later, on 27 February 2014, the Guardian published its last Snowden documents. Over nine months it had published around 30 of the 58,000 documents it held — 0.05%. Rusbridger would later describe the episode as the story with the biggest global impact in the Guardian‘s 190-year history.[20]

What stopped in February 2014 was the publication of documents, not the reporting. For example on January 2015 the Guardian reported from the Snowden documents that GCHQ had captured the emails of journalists at major international outlets. And on June 2015 it published a joint investigation with The New York Times into GCHQ’s role in American drone strikes in Yemen and Pakistan, drawn from the documents provided by Snowden to the Guardian “and shared with The New York Times.”

So even sixteen months after its last document publication, the Guardian was still analyzing the archive and sharing documents with its American partner, but it chose not to release any documents (nor did The New York Times).

In May in 2014, the DPBAC committee’s chair reported that engagement with the Guardian had continued to strengthen and that the process “had culminated by the appointment of Paul Johnson (Guardian deputy editor) as a DPBAC member.”

Jacob Appelbaum – who worked on the archive, published extensive reporting in Der Spiegel and elsewhere, and helped Poitras vet Snowden — sharply criticized the Guardian in March 2016:

[…] And while writing technical stories [about the Snowden documents], [the Guardian] directly consulted with the White House and with GCHQ and other government officials in order to do essentially line by line redaction of things that were for the most part not actually even worth redacting. They weren’t even worth compromising yourself. It reminds me of the Winston Churchill story about whether or not someone would sleep with him for a million dollars. And of course, when the person says no to one dollar but yes to a million, we know what kind of person that person is. And the same is true for the Guardian. They were willing to compromise and to give editorial control to the state. What are they then? They’re stenographers.

In his 2018 book Breaking News, Rusbridger covers the Snowden era in extensive detail, discussing the Heywood meetings, the destruction of the hard drives at the Guardian‘s London office, Miranda’s detention at Heathrow, critical editorials from rival newspapers, and politicians demanding prosecutions. However, the reporting ends in a single clause: “Once the Snowden coverage had died down.”[21] This stands out as the only major event in the chapter lacking a clear actor or motivation. While all the other actions have clear motives, media coverage just fades away, with no explanation who decided to stop, when the choice was made, or why.

When asked in 2023 about the lack of document publication, MacAskill blamed waning interest, noting that “each story attracted smaller and smaller readerships, as interest dwindled.”

That is hard to square with what the Guardian‘s own editor says about the same period. The paper published its last Snowden document accompynied with a story which was significant front-page news, as the most Snowden stories before it had been. Two months later it won a Pulitzer for the reporting. Eight months after the last document, Rusbridger records, the paper overtook The New York Times to become the leading serious English-language newspaper website in the world, having been the ninth-biggest paper in Britain.[22] He attributes the rise to doing “the stuff that journalism was — at least in our reckoning — supposed to do”: a list of subjects he expected nobody to read, in which security and civil liberties come second and third. “No one — surely — would read you if you banged on about climate change, security, civil liberties […] But they did.” He concluded that “there was clearly a huge, and global, appetite for significant news presented seriously. There was a need for it.”[22]

Nor had the appetite gone elsewhere. Der Spiegel went on publishing documents into 2015, with stories that made international headlines. The Intercept continued until 2019, across almost a hundred pieces. Whatever caused the Guardian to stop publishing documents in February 2014, it was not that there was nothing left or that nobody was reading.

According to Rusbridger[10] and MacAskill, Snowden had asked from the start that the reporting stay on surveillance and privacy rather than the wartime use of intelligence, and that Rusbridger had issued that internally as an edict. MacAskill describes being asked by Rusbridger, at some point after the London copy was destroyed, to go back to The New York Times — “which still had the material” — and review it for stories that might become reportable if that restriction was lifted. He returned to London with a list of about a dozen. Rusbridger rejected them, MacAskill writes, “not only because he had not intended to renege on the agreement with Snowden, but also because none of them was as explosive as the original stories.”

There is a gap in the account of where the Guardian‘s copies went. Rusbridger writes in Breaking News that after the London destruction the paper was “left with the dossier intact — in New York,” noting that the UK authorities showed little interest in “the material we held at 536 Broadway,”[23] the Guardian‘s own New York office. But when Rusbridger later asked MacAskill to review what remained, MacAskill didn’t go to his own paper’s New York office, but instead he was sent to The New York Times, “which still had the material.” And in 2023, describing where the archive sits now, MacAskill mentions only the copy locked in an office at The New York Times, with the Guardian retaining responsibility for it. The Guardian‘s own New York copy is not mentioned.

The Guardian shared their archive with The New York Times and partly with ProPublica,[11] under an agreement Rusbridger describes in Breaking News.[10]

[…] I typed out our conditions for collaboration on a side of A4. The main stipulation was that the NYT (and also ProPublica) wouldn’t use the archive as a bran tub to go fishing for stories unrelated to Snowden’s primary focus. I knew there were documents about Afghanistan, Iraq, even Kenya, in the archive. We weren’t going to look at that. This was about surveillance and civil liberties, not the wartime use of intelligence. The NYT agreed. Sometimes I think they (and ProPublica) wished they hadn’t, but they stuck by the agreement.

In March 2016, Appelbaum said that the Guardian holds ProPublica in a gag over the Snowden archive.

[…] For example, the Guardian holds ProPublica in a gag. You may not know this. But ProPublica has access to the Snowden archive. But they are not allowed to publish things, unless the Guardian will allow them. And the Guardian has decided that they will not allow things from the Snowden archive to be published. Things about Afghanistan and Iraq. Crimes, serious war crimes are documented in there. Crimes where civilians are killed. Things that are absolutely political. And we will never see them, because of the collaborationists at the Guardian, who absolutely kowtow to the British political class and the hereditary power structures in the UK. […]

On 18 August 2026 we contacted the Guardian‘s press office with questions. The press office replied saying they had nothing new to share, that they don’t routinely comment on editorial decision making, and that we were asking about issues from over a decade ago. They didn’t address whether they still hold a copy or if any material remains at The New York Times and whether they retain any responsibility for it, or whether they hold any approval rights over publication by their partners – none of which are editorial decisions.

The Washginton Post

Barton Gellman holds a significant portion of the full archive. He received a copy with Poitras from Snowden in 21 May 2013,[1] containing over 50,000[5] documents.[4] Gellman took a copy of his archive to The Washington Post‘s New York office in May 2013, where it was secured[24] until he left The Washington Post in 2014 and it stopped publishing Snowden documents.

Gellman said in 2020 that his “material is now in cold storage, as secure and inaccessible as I could devise.” In 2022 he agreed with Poitras that the archive would be “really valuable for constructive research,” but added that “the opsec needed to share it with anyone else is too hard to deal with. I just put the whole thing in cold storage. I feel bad about that.”

Gellman’s account of what The Washington Post built is the fullest description anyone has given of how the archive was held in any outlet.

He went to the editor, Marty Baron, with a list. Dedicated computers with freshly wiped, encrypted drives. Networking hardware physically removed, cutting the machines off from the internet and the newsroom’s own systems. A windowless room with a high-security lock, a reinforced door, and a heavy safe bolted to the floor. Decryption keys stored on memory cards and never kept in the same room except when in use. Access would require four credentials — door key, safe combination, digital key card, passphrases — divided among the team, with nobody but Gellman holding all four.[25]

The Washington Post‘s first attempt at the room had a wall full of windows, with a view of the Russian ambassador’s residence half a block away. They found another. It got a high-security lock, a camera in the hall outside and a safe weighing four hundred pounds. Gellman kept a second safe at his own office in New York.[25]

Working alongside technologist Ashkan Soltani, they prepared their laptops by removing the internal Wi-Fi, Bluetooth, and batteries, ensuring the machines would instantly shutdown and encrypt themselves if unplugged. They physically sealed the USB ports and maintained strict security by taking their keys out of the room every time they left, even for a quick bathroom break. To catch any potential physical tampering, Gellman applied epoxy and glitter to the laptop screws and tested ultraviolet powder on the safe’s dial. He stored all his notes on encrypted volumes that required five distinct passphrases just to open each morning, a system that backfired when he forgot one passphrase and lost access to some files forever.[25]

The Washington Post published its last Snowden documents in July 2014. At the end, it published around 30 documents in total,[26] or about 0.06% from the archive it had.

Gellman writes that by late autumn 2015, he and Soltani were no longer writing articles for the newspaper. Soltani stopped using his old laptop, gave back an encryption key fob, and cut his ties to the archive.[27]

The New York Times and ProPublica

Both The New York Times and ProPublica received material from the Guardian in 2013, under the conditions Rusbrdiger typed onto the sheet of A4.[10] Neither held the full archive, and neither has ever said how much it received.

The New York Times and ProPublica published a handful of documents between 2014 and 2015. The last story was a joint investigation published on 15 August 2015. Neither has published a Snowden document since, and neither has explained why. Both stopped immediately after a substantial joint investigation rather than after a decline, and both stopped ten weeks after Rusbridger left the editorship of the paper whose conditions they were working under.

There is one account of those conditions operating on a specific story, though it is second-hand and unconfirmed. Appelbaum has said that Jeff Larson, who contributed reporting to ProPublica‘s 2014 piece on the Mumbai attacks, told him the partners who had supplied the files under conditions wanted the account steered so that important parts would not appear, and that Larson was able to report freely only after obtaining the original material from Poitras and Appelbaum directly.

We asked ProPublica whether it still holds any Snowden material, whether that material is subject to an approval requirement, and why the publication stopped. We received no response.

Der Spiegel

German journalists Marcel Rosenbach and Holger Stark, working for Der Spiegel, published a book about the Snowden affair in 2014, called Der NSA-Komplex. In the book they write how Der Spiegel got its material.

At first Rosenbach and Stark tried getting access to the Snowden archive via Greenwald. On a Saturday in mid-June 2013 they took a call from an acquaintance who said Greenwald was willing to sit down with them. The next morning they flew to Rio with hard drives, laptops, a cryptophone and encryption software in their hand luggage.[9]

They met at a hotel on Copacabana. At ten in the evening, two hours late, Greenwald shuffled into the lobby in bermuda shorts, flip-flops and a washed-out T-shirt, a black rucksack over his shoulder holding the drives full of American state secrets. “They broke into my house today and stole a computer,” he said, visibly shaken. “What can I help you with?” He pointed them at various documents concerning Germany, but said that they would have to obtain them by some other route.[9]

Rosenbach and Stark flew home the next morning empty-handed. Back in Berlin they went to Poitras, who had already heard about the Rio trip and knew them slightly from their earlier WikiLeaks reporting. They proposed working together on stories for Der Spiegel, with Poitras as co-author and freelance contributor. She was a byline on the paper’s first Snowden cover story on 1 July, and on much of what followed.[9]

Der Spiegel fought with the NSA every week. “Over months we dealt weekly with headquarters at Fort Meade, at times also with the White House,” they write. “Every Friday, at the Spiegel’s deadline, we wrestled with the intelligence service over the details of the reporting, sometimes by email, sometimes by conference call.”[9]

They describe withholding material. “It cannot be the aim of critical journalism to play into the hands of the opponents and declared enemies of democracies, or to do the business of other intelligence services,” they write. “For that reason we tried, in an extensive process of discussion, to weigh where the boundaries of public interest lie — and at a number of points refrained from publishing sensitive information.”[9]

Der Spiegel published some 140 documents across around eighteen stories between July 2013 and January 2015. Its two largest releases were also its last two: 44 documents in December 2014, and 36 on January 2015.

Der Spiegel has never said why it stopped publishing documents.

The Intercept

In Hong Kong 2013, Snowden told the journalists to spread the material, that they should make sure it was never held by one person alone. “Make sure that never only one single person possesses a copy,” he said. “The United States or its allies will quite certainly kill you if they believe you are the weakest link in the chain, through which the publication of the information could be stopped.” Greenwald, Poitras and MacAskill were to ensure that “people you can trust” held copies.[28]

Within weeks, only the two of them held complete copies, by agreement. Greenwald and Poitras decided that nobody else would ever have access to the full archive, in order to protect the material and avoid legal entaglements, and to keep the outlets they worked with, as Greenwald put it, “on a leash.” Everyone else got a portion. “Only Laura and I have access to the full set of documents which Snowden provided to journalists,” Greenwald told BuzzFeed at the end of August 2013. The New York Times and ProPublica had only the GCHQ material, and Gellman likewise had “only a small subset of the documents, though the number is substantial and relate to NSA.”

In October 2013, tech billionaire Pierre Omidyar founded First Look Media (FLM), a New York nonprofit, as a collaboration with Greenwald, Poitras and Jeremy Scahill, with a promised $250 million in funding. FLM’s first publication, The Intercept, launched on 10 February 2014 with purpose to report the Snowden documents.

The arrangement drew criticism at the time. Writing at Pando in November 2013, Mark Ames argued that the most significant secrets of the era had ended up dependent on the goodwill of a single billionaire.

In April 2014 FLM hired Lynn Dombek from the Associated Press as research director, to build a team providing primary sources and evidence for its publications. Morgan Marquis-Boire joined as director of security in June 2014, and Erinn Clark, previously a Tor Project developer, joined as lead security architect that September.

They built a SCIF (a secure compartmented information facility) on the 19th floor of a building at Fifth Avenue and 16th Street in Manhattan. A former FLM employee described it having double doors that could only be opened by two people with separate keys, computers requiring two operators to start, phones left outside, and permission needed to burn anything from the archive to an air-gapped computer.

The 2014 and 2015 returns list The Intercept alongside Racket, Reported.ly and Field of Vision, with no research or security activity declared, even though during 2014-15 The Intercept published around 256 documents from the archive.[29]

On 16 May 2016, The Intercept announced it would begin publishing the NSA’s internal newsletter in batches and open the archive to outside journalists and researchers. Greenwald wrote that there were “still many documents of legitimate interest to the public that can and should be disclosed,” and encouraged others to comb through the material because “others may well find stories, or clues that lead to stories, that we did not.”

In the same year, FLM declared the Research and Security Group to the IRS as a new significant program service: “a group of award-winning research, security, engineering and editorial experts who make documents available for inquiry and analysis in secure environments.” It was reported at $863,189, one of the organisation’s three largest program services that year.

In 2017 the Research and Security Group was one of FLM’s three largest program services at $1,578,937 in a year the organisation lost $12.2 million. It was the worst year in its history, and it still funded the archive programme at nearly double the previous year’s figure.

By the 2018 return, the Research and Security Group had ceased to be a programme, and it was now a single clause inside the description of The Intercept: “Through the Research and Security group, known as Neptune, which includes award-winning research, security, engineering and editorial experts, The Intercept makes documents available for inquiry and analysis in secure environments.”

The closure

On 4 March 2019, Betsy Reed, The Intercept‘s editor-in-chief, emailed Laura Poitras to ask for a meeting, in confidence, about “how we’ve assessed our priorities in the course of the budget process, and made some restructuring decisions.” Two days later Reed and Jeremy Scahill went to Poitras’s studio in Tribeca and told her that four positions were being cut, among them the staff who maintained the Snowden archive. During the meeting it became “clear they have decided to eliminate the research department,” Poitras wrote.

Poitras objected repeatedly over the following days, in meetings and in writing. “Rather than causing the closure of the archive,” she wrote, “I objected to it in meetings and in writing multiple times. The Intercept‘s editors said the reason for their decision was a budget cut.” Eliminating the research department, she argued, would jeopardise the archive’s security and was therefore negligent, and would make “unauthorized changes to the agreement” under which FLM held the archive.

On 10 March she argued that the department cost 1.5% of FLM’s total budget. She offered several ways round it, including reducing her own salary to pay for the remaining staff. At this point the budget was the only reason anyone had given.

The Intercept paid its senior journalists well. FLM reported $1.6 million to Glenn Greenwald’s company between 2014 and 2017; $349,826 in total compensation to Jeremy Scahill in 2015; $368,249 to Betsy Reed in 2017. Charles R. Davis, writing in the Columbia Journalism Review, noted that such figures “are large for digital media and noteworthy in the world of progressive, nonprofit journalism.”

On that same day, 10 March, the problem appeared to be solved. Drew Wilson, FLM’s chief financial officer, emailed Poitras to say he had spoken with the chief executive Michael Bloom, and that “for now, we will hold off on any actions on the remaining two staff in the research team until clarity on the go-forward strategy is reached. I informed Betsy that for now she will not have to cover this portion of the cost reduction in the upcoming restructuring this week.”

Poitras thanked him the next day for “putting a pause on eliminating positions that directly impact the archive security.” The problem, as she put it, seemed solved.

On 12 March Poitras was told on a telephone call with Greenwald and Wilson that Greenwald and Reed had decided to shut the archive, because it was no longer of value to The Intercept. On the same call, according to her account, Greenwald said the decision should not be made public because it would look bad for him and for The Intercept. Two days earlier the problem had been money, but now it was “value”.

On 13 March, Poitras sent a memo to FLM’s board urging it to intervene, noting that she had not been consulted and the board had not been told either. Hours later, Bloom emailed the company:[30]

As numerous media outlets have discovered over the last several years, the current business environment sometimes requires painful decisions, and it is always agonizing when reorganization results in the loss of colleagues’ jobs. […] It is crucial to recall that major news outlets that possessed large portions of the Snowden archive — newsrooms much larger than The Intercept‘s — ceased reporting on it years ago. Many decided that the resources required to continue to work on the archive were not justified by the journalistic value the remaining documents provide, as those documents have aged. For five years, the company expended substantial resources to continue to report on the Snowden archive, but The Intercept has now decided to focus on other editorial priorities.

Read closely, the only reason Bloom gives for The Intercept‘s own decision is the last clause, that it had decided to focus on other editorial priorities. The aging documents and the unjustified resources are attributed to other outlets, describing what they concluded years earlier. They are given as context and not as a stated rationale.

Poitras replied to Bloom the same day.

Michael, As I have communicated to you and Betsy, I am sickened by your decision to eliminate the research team, which has been the beating heart of the newsroom since First Look Media was founded, and has overseen the protection of the Snowden archive.

I am also sickened by your joint decision to shut down the Snowden archive, which I was informed of only yesterday — a decision made without consulting me or the board of directors. Your email’s attempt to paper over these firings is not appropriate when the company is presented with such devastating news.

Her email leaked to the Daily Beast which published it that evening. On the following day 14 March Poitras telephoned Snowden to apologize, who had not been told.

Greenwald issued a statement on 14 March.[12] Of the other outlets that had held and reported on Snowden documents, he wrote:

But all of them stopped working completely on the Snowden archive and stopped publishing Snowden documents many years ago, presumably because they decided that the massive financial and human resources required to work with the archive – including elaborate security measures to protect it and large teams of journalists and editors to curate and responsibly report on it – were no longer justified by the journalistic value the archive provided as the remaining documents aged.

Then, of The Intercept‘s own decision:

The Intercept‘s decision to stop working with the archive was the by-product of those financial constraints and, critically, of my intent to seek other partners – particularly academic institutions and research facilities – to ensure continued publication of the remaining Snowden documents that are in the public interest.

The last sentence repays reading twice. The construction is that the decision was the by-product of (a) financial constraints and, critically, of (b) my intent to seek other partners. There are two causes, with the second highlighted as “critically”, as the more significant one. The partner search is treated as the main cause. The decision doesn’t come from The Intercept, the board, or from Greenwald and Poitras together, but from Greenwald himself.

Based on his own words, Greenwald closed the archive because he planned to move it somewhere else, which was more important than the finances.

Four reasons had now been given for the closure during eleven days:

  1. Budget cuts — the initial meetings and emails, 4 to 10 March.

  2. The archive was no longer of value to The Intercept — the telephone call of 12 March.

  3. The Intercept had decided to focus on other editorial priorities — Bloom’s staff email, 13 March.

  4. Financial constraints “and, critically, of my intent to seek other partners” — Greenwald’s statement, 14 March.

The laid-off staff were required to sign non-disclosure agreements prohibiting them from discussing their work.

By the 2019 return, filed in November 2020, the Research and Security Group has gone, which is expected. Part III of the form asks whether the organization stopped running or significantly changed any of its program services. FLM answered “No.” However, it had answered “Yes” twice in the past: in 2015, when three abandoned projects ended, and in 2017, for Reported.ly, an experimental social-media news network. Notably, a publication that never actually launched was reported to the federal government as a discontinued program service, while the Snowden archive was not.

The same return shows revenue of $28.4 million against expenses of $28.2 million. Net assets up again, to $21.5 million. Total compensation up from $9.8 million to $16.6 million. Program spending up by $1.26 million. 87 employees, against 50 in 2016. The single largest outside contractor was Enzuli Management LLC, at $458,337, described in the return as “journalism services, including but not limited to services provided by Glenn Greenwald.”

The Intercept kept publishing for another eleven weeks. On 29 May 2019 it released the eighth and final batch of SIDtoday documents alongside four stories drawn from it. That day is the last time anyone, anywhere, published a Snowden document.

Two years later, Greenwald gave a fifth reason for the closure, and it contradicted his own.

In February 2021, New York Magazine revisited the closure. Greenwald said that “nobody wanted to close the archive,” that he “never heard anyone at The Intercept talking about wanting to close it,” and that “the only reason it was closed was because a dispute arose between Betsy and Laura — after Betsy was required by First Look to lay off four employees — about whether the archive would still be securely maintained if Betsy proceeded to lay off the people she chose to lay off.”

In March 2019 he had written that the closure was the by-product of financial constraints and, critically, of his own intent to seek other partners. But now in 2021 nobody wanted it, and it happened because two colleagues disagreed.

Scahill, quoted in the same article, said: “It is my understanding that Laura attempted to place conditions on the continued use of the archive that no independent outlet or editor-in-chief could accept and then sought to blame The Intercept for a scenario she herself made inevitable.”

Poitras says the conditions weren’t hers, that they were the terms of an existing confidential agreement with FLM, covering how the archive had to be secured. Cutting the staff, she argued, would breach them. David Bralow, FLM’s general counsel, said the opposite: “any suggestion that The Intercept violated any contract by making budgetary decisions about its staff is false.”

Poitras responded to the article by updating her open letter, republishing some of the emails Barrett Brown had released in March 2019 and adding others. Her former colleagues, she wrote, “made several unsupported claims regarding the closure of the NSA Snowden Archive.” She concluded: “The Intercept‘s effort to rewrite history is not only full of falsehoods, but is quite literally hard to stomach.”

She had put the underlying objection more plainly two years earlier, in her reply to Bloom:

How a news organization would take such care to secure this archive, and then walk away from that knowledge and its investment without a proper review involving the board and all stakeholders, defies my understanding.

The closure wasn’t only justified by shifting and conflicting reasons, but it was also done very quickly, and without the people who might’ve prevented it. Nine days passed between the first meeting and the public announcement.

Nor was the alternative plans considered. Micah Lee, The Intercept‘s director of information security, said in the 2021 New York Magazine article that his team could have come up with a plan to secure the archive. “I feel like a compromise wasn’t seriously considered.”

And on the call of 12 March 2019, according to Poitras, Greenwald said the decision should not be made public because it would look bad for him and for The Intercept.

The arguments

Bloom’s email and Greenwald’s statement rest on two arguments about the closure: (1) that the remaining documents had lost their journalistic value with age, and (2) that other outlets had reached the same conclusion years earlier. Both arguments warrant scrutiny, because together they have come to form the accepted explanation for why the Snowden archive eventually fell silent.

“The remaining documents have aged”

The argument is that documents lose journalistic value as they get older, and that by 2019 what was left no longer justified the cost of holding it.

The argument is difficult to reconcile. The most recent material in the archive dates from April 2013,[31] so at the time of the closure many of the documents were just six years old. And in journalism decades-old declassified documents regularly become the basis for significant journalism. Also as time passes the case against publishing the documents weakens while the case for disclosure strengthens.

Both Greenwald and Poitras agree that the archive had stopped being news. Greenwald wrote in his statement[12] that it was “far less of a news resource and far more of a historical asset.” Poitras wrote in the same year that “while it is true the archive can no longer be reported on as ‘news,’ it remains the most significant historical archive documenting the rise of the surveillance state in the twenty first century.”

Greenwald and Bloom didn’t say the archive had stopped being news, but that it had lost journalistic value, as if it was no longer worth the effort of journalistic reporting at all.

The Intercept was still publishing documents until the very end. The final batch of documents was published on 29 May 2019, eleven weeks after the closure was announced, and produced four stories. Similarily the Guardian went from a front-page GCHQ story in February 2014 to nothing. The New York Times and ProPublica went from a joint investigation in August 2015 to nothing.

And Greenwald had argued the opposite three years earlier. Announcing the batch releases and the outside-access programme in May 2016, he wrote that there were “still many documents of legitimate interest to the public that can and should be disclosed,” and urged other journalists and researchers to go through the material because “others may well find stories, or clues that lead to stories, that we did not.” In June 2023 he described the archive as running to “hundreds of thousands of documents, if not more.”

By most estimates around 1% of the archive has been published.[32] It’s a quite confident claim that the other 99% had aged out of journalistic value.

Appelbaum published previously unreported findings from the archive in his 2022 doctoral thesis, including the NSA listing the chipmaker Cavium as a “SIGINT enabled” CPU vendor, NSA compromising Russia’s SORM lawful-intercept system, and NSA participation in Internet Engineering Task Force (IETF) standards meetings with the explicit aim of weakening protocol security.

He also offered an explanation in his thesis for why so much remains unpublished:

As part of our research, we uncovered evidence that the telecommunications infrastructure in many countries has been compromised by intelligence services. The Snowden archive includes largely unpublished internal NSA documents and presentations that discuss targeting and exploiting not only deployed, live interception infrastructure, but also the vendors of the hardware and software used to build the infrastructure. Primarily these documents remain unpublished because the journalists who hold them fear they will be considered disloyal or even that they will be legally punished. Only a few are available to read in public today.

And elsewhere in the same work:

Many journalists who have worked on the Snowden archive know significantly more than they have revealed in public. It is in this sense that the Snowden archive has almost completely failed to create change: many of the backdoors and sabotage unknown to us before 2013 is still unknown to us today. The entire Snowden archive should be open for academic researchers to better understand more of the history of such behavior.

“Other major media outlets stopped reporting as well”

The second claim is that The Intercept was simply the last to do what everyone else had already done. Bloom wrote that the larger newsrooms holding portions of the archive “ceased reporting on it years ago,” and many had decided the resources involved were no longer justified “by the journalistic value the remaining documents provide, as those documents have aged.” Greenwald used almost the same construction the following day, with one addition — they had stopped, he wrote, “presumably because” they reached that conclusion. Few hours later Greenwald dropped[33] the hedge, saying the other major outlets stopped “for cost reasons.”

As a description of the outcome, the claim is accurate. The Guardian published its last documents in February 2014. Der Spiegel‘s last release was in January 2015. The New York Times and ProPublica both stopped in August 2015. After that, The Intercept was the only outlet publishing new documents, with a few exceptions.[34]

But as an explanation, it doesn’t hold.

The only outlet that has explained itself is the Guardian, and its account isn’t the one Bloom attributes to it. MacAskill has given two reasons, neither of which is aging documents. The first is readership: “it reached a point where each story attracted smaller and smaller readerships, as interest dwindled. The feeling at The Guardian — and, I assume, at The New York Times and ProPublica — was they had reported on the biggest stories in the documents and there was diminishing interest in publishing more.” Note that MacAskill is also presuming about his partners.

MacAskill’s second reason is scope. Rusbridger had confined the Guardian‘s reporting to surveillance and privacy at Snowden’s request[10], and when MacAskill was asked to review the material at The New York Times for stories that would become reportable if that restriction were lifted, he returned with about a dozen. Rusbridger declined them.

Unlike the Guardian, The Intercept wasn’t bound by the same restrictions in its Snowden reporting. Since its very first Snowden story, The Intercept routinely published stories outside the Guardian‘s scope.

As for The New York Times and ProPublica, neither of them has explained why they decided to stop. In 2016 Appelbaum said publicly that ProPublica held Snowden material but couldn’t publish without the Guardian‘s consent, which he said was withheld.

Bloom treats the other outlets as having made an editorial judgement, but the Guardian had its London copy destroyed under government supervision and was under police investigation, and The New York Times and ProPublica were both working under written conditions imposed by another outlet.[35]

And The Intercept was founded to report the archive. It was the only outlet that existed for that purpose and the only one that had opened the material to outsiders on the grounds that more remained to be found. The fact that the others stopped publishing documents wasn’t an obvious reason for The Intercept to do so.

The budget

Poitras objected to the cuts repeatedly in meetings and in writing. On 10 March 2019 she argued that the research department was only 1.5% of FLM’s total budget, and the chief financial officer agreed to keep the two remaining staff. Two days later she was told the archive was being closed anyway.[36]

From the filings it doesn’t look as though FLM was under pressure. In 2017 FLM lost $12.2 million and still funded the Research and Security Group at $1,578,937, reporting it as one of its three largest charitable programmes. In 2018 it ran a surplus of just over $6 million and held $21.4 million in net assets. In 2019 it increased program spending, raised total compensation from $9.8 million to $16.6 million, and employed 87 people, up from 50 three years earlier. It paid $458,337 to Glenn Greenwald’s company that year.[36]

An organisation running a surplus could afford the archive through the worst year in its history, but couldn’t afford it in the best.

“Editorial priorities”

Bloom wrote in his staff email that The Intercept “has now decided to focus on other editorial priorities.”[30]

It’s the only reason in his email that describes a decision The Intercept itself made, and it’s a reason that nobody has ever explained. For example, no new priorities were stated, and there were no changes in publishing formats, audience targeting, or anything else that would indicate a shift in editorial priorities.

By 2016 the organisation had told the federal government that providing access to the archive was a distinct charitable activity, and was funding it at over a million and a half dollars a year.[36] A shift away from that represents a substantial change of direction for a news organisation, yet it was announced in a paragraph of a staff email concerning staff reductions.

“My intent to seek other partners”

Greenwald said in his statement that he had spent “the last several months” seeking “other partners – particularly academic institutions and research facilities – to ensure continued publication of the remaining Snowden documents that are in the public interest.”[1]

Seven years later, and no such partner has been found, and no institutions has been named. We asked Greenwald which institutions he approached and what came of it. He did not reply.

Snowden gave his own explanation on 22 April 2019, in an interview with Motherboard‘s CYBER podcast.[37] The other outlets had stopped, he said, “because it became more expensive and longer form work, instead of like the short punchy news pieces which is what the large appetite is in journalism today.” What remained in the archive “is going to require much more substantial effort, basically book length work. You need real researchers to sit down, and to connect all the dots that you can’t get across in 750 words. And The Intercept really was not designed for that, unfortunately. None of these news organizations were.”

Snowden’s explanation that the work became longer-form, and that the outlets weren’t built for it, doesn’t fit what they had been publishing. For example, The Intercept‘s January 2019 piece on supply-chain attacks runs to around 5,200 words. Der Spiegel‘s January 2015 story on the NSA’s struggle for control of the internet runs to ~3,700. They were already doing the long-form work he describes as beyond them. Nor is it obvious why an archive of hundreds of thousands of documents, from which 1% had been published, should have produced only the short pieces and none of the long.

Snowden continued saying that the outlets “were supposed to hand this off to academic institutions, but that just hasn’t happened because the academic institutions get cold feet. They get nervous. They go, look, we’re dependent on grants from the federal government in the U.S. And they don’t want to give it to a foreign university because the politics of that, they’re worried the government’s going to complain.”[37] This is the only explanation anyone has given of why the search for a new home failed. But all we know is that the search was given as a reason for closing the archive, and nothing came of it.

While Snowden may well be right about universities, the hand-off doesn’t necessarily require an academic institution at all, because dozens of organisations maintain SecureDrop and employ lawyers, and many of them don’t depend on federal grants. This could be tested just by submitting a single document to them.[38]

“Nobody wanted to close it”

Two years later, in February 2021, Greenwald said in the New York Magazine that nobody at The Intercept wanted the archive closed, and that it closed only because Reed and Poitras disagreed about whether it could be securely maintained after the layoffs.

In March 2019 he had written that the closure flowed in part from his own intent, and that he had spent the preceding months preparing for it. Poitras was told on 12 March that he and Reed had decided it. Two years later, nobody wanted it and it happened because two other people argued.

Two more arguments could be made for not publishing any documents.

The first is legal risk. The documents carries real exposure under the Espionage Act, but Poitras, Gellman, and Greenwald published from the archive for years when the material was newer and the exposure greater, and none of them has ever cited it as the reason. The second is infrastructure. Publishing from the archive does require some opsec, because the material has to come out of (cold-)storage, be worked with secure air-gapped computers and be read closely enough to know what is worth publishing. Gellman has said the opsec of sharing the archive with anyone else is too hard to deal with, and that explains why he won’t hand it to someone, but it doesn’t explain why he hasn’t shared any documents from it or published himself.

What happened to The Intercept’s copy

The Intercept has never said what became of its copy of the archive, beyond that it was closed. In his 2022 doctoral thesis, Appelbaum wrote that it had been destroyed,[39] and in a 2023 interview he said an insider had told him so.

The Intercept destroyed its copy of the Snowden archive. That’s what an insider told me. Those responsible for the archive have failed to live up to their responsibilities. They have withheld many things that are in the public interest.

The claim is based on an anonymous source and remains unconfirmed. The Intercept has been asked twice by two journalists and has declined to answer both times. Ralf Hutter put it to The Intercept in June 2023, and The Intercept declined to comment on the archive’s existence, saying it was a confidential matter. Stefania Maurizi put the same question in the same year and was told The Intercept doesn’t discuss confidential news-gathering materials.

We asked again of The Intercept, of First Look Institute, of Michael Bloom, and of eight people who worked on the archive at The Intercept. Nobody answered.

The holders

Greenwald and Poitras hold complete copies of the archive. Gellman holds the +50,000 files Snowden sent him directly in May 2013. None of them is constrained by anything that constrained The Intercept, like a budget, a board, a funder, or editors.

Greenwald resigned from The Intercept in October 2020, over the spiking of an article, and has published independently ever since. Poitras was terminated in November 2020. Gellman left The Washinton Post in 2014 and has written independently since.

All three have said the material matters. In 2022 Poitras told that “the Snowden Archive still exists, and there is still more to report,” and that “there is a vast amount of information that hasn’t been reported of enormous contemporary and historical significance.” Gellman agreed saying that the archive would be “really valuable for constructive research,” but said “the opsec needed to share it with anyone else is too hard to deal with. I just put the whole thing in cold storage. I feel bad about that.” In June 2023 Greenwald described the archive as running to “hundreds of thousands of documents, if not more.”

None of the three has published a document from it since May 2019.

On 6 June 2023, the tenth anniversary of the first Snowden story, Greenwald hosted Snowden and Poitras for a two-hour discussion on his programme System Update. They talked about the reporting, its consequences, and the state of surveillance since. The closure of the archive was not discussed, nor was the fact that nothing had been published from it in four years. Four years earlier, Poitras had released documents in which Greenwald was recorded as saying that he and Reed had decided to close the archive and that it should be kept from the public because it would look bad for him. Two years earlier, he had told New York Magazine that nobody had wanted the archive closed. Neither came up.

Rusbridger, writing in 2018:

Edward Snowden could easily have published his concerns and/or revelations himself. He chose not to. He deliberately went to journalists he thought would understand the significance of what he was disclosing and asked them to make their own judgements. It was, to some extent, an act of faith in journalism. Whether whistleblowers will behave in such a way in future is unknown. There would, of course, be no incentive in taking such disclosures to newspapers if they declined to print them.

{{ fin. }}

Between August and September 2026 we tried contacting over twenty people and organisations. First Look Institute, including questions for Michael Bloom, and The Intercept. Betsy Reed, Laura Poitras, and Glenn Greenwald. The Guardian‘s press office and its editor-in-chief Katharine Viner. Alan Rusbridger, Ewen MacAskill, Janine Gibson, Julian Borger, Gill Phillips and Zoe Norden. ProPublica‘s press office. Laura Poitras, Jeremy Scahill, Murtaza Hussain, Micah Lee, Erinn Clark, Lynn Dombek, and others.

Two replied. Gill Phillips, the Guardian‘s lawyer throughout the Snowden period, wrote to say she had retired and could not assist. The Guardian‘s press office said it had nothing new to share and did not routinely comment on editorial decision making.

Nobody answered a single question of substance.

If you know something about any of this, we’d like to hear from you.

References

Books cited

  • Gellman, Barton. Dark Mirror. Penguin Books. 2021.
  • Poitras, Laura. Astro Noise. Whitney Museum of American Art. 2016.
  • Rosenbach, Marcel and Holger Stark. Der NSA-Komplex. DVA. 2014.
  • Rusbridger, Alan. Breaking News. Farrar, Straus and Giroux. 2018.
[1]:
Dark Mirror, pp. 388-389.
[2]:
Astro Noise, p. 92.
[3]:
Astro Noise, p. 93.
[4]:
Snowden sent five encrypted containers to Gellman and Poitras, but provided the encryption key for only one of them. The one and only encrypted container Gellman could open was named “Pandora.” One of the containers was bigger than Pandora.[Dark Mirror, Chapter Eight “Exploitation”, p. 284] Inside Pandora, there was an another encrypted archive called “Verax,” and inside Verax, there was yet another encrypted archive called “Journodrop.”[Dark Mirror, Notes, note 21, p. 368] Snowden told Gellman that the other encrypted archives he could not access “would probably be deadman linked, time locked, or the like.” He deliberately kept the details vague, explaining that “discussing those mechanisms weakens them.”[Dark Mirror, Chapter Eight “Exploitation”, p. 284] Also worth noting that it’s not clear whether astro_noise archive Poitras downloaded in 31 March is the same archive that Gellman calls Pandora, or a separate archive.

He had sent what appeared to be five encrypted containers all at once, but he provided the encryption key for only one of them. One of the containers was bigger than “Pandora,” the one I could unlock. “The only thing you have is what you have plaintext access to,” Snowden told me. “Anything beyond that would probably be deadman linked, time locked, or the like.” He declined to explain further, saying, “discussing those mechanisms weakens them.”

[5]:
Dark Mirror, p. 24.
[6]:
Snowden has declined three requests from Bruder and Maharidge to discuss this part of the timeline saying “I’m not sure I’m ready to tell my side of that part of the timeline yet.”
[7]:
Der NSA-Komplex, Chapter 2.

[…] dass Greenwald an Bord der Cathay-Pacific-Maschine mit der Flugnummer CX831 lange Passwortketten eingibt, bis sich hoch verschlüsselte Datencontainer öffnen. Über den Wolken auf dem Flug gen Westen klickt sich Greenwald von einer als »streng geheim« eingestuften Dokumentation zur nächsten. »Ich konnte gar nicht aufhören weiterzulesen«, sagt er, »jede einzelne Minute des Flugs über habe ich gelesen.«

Translated:

[…] Greenwald, aboard the Cathay Pacific flight CX831, enters long strings of passwords until highly encrypted data containers open. High above the clouds, on the westbound flight, Greenwald clicks from one document classified as “strictly confidential” to the next. “I simply couldn’t stop reading,” he says. “I read for every single minute of the flight.”

[8]:
Breaking News, p. 224.

[…] MacAskill was back in London with a thumb drive of documents. […]

[9]:
Der NSA-Komplex, Chapter 8, Eit Überwachten.
[10]:
Breaking News, p. 225.
[11]:
Der NSA-Komplex, Chronik: Die Snowden-Enthüllungen, 20. Juli 2013.
[12]:
Greenwald published the statement (archived1, archived2) in his Twitter account as a two images: one (archived1, archived2), two (archived1, archived2). The statement is attached here as text:

Contrary to the perception of many, the Intercept is not the only media outlet to possess Snowden documents. Many large media outlets with budgets and newsrooms far larger than ours - including the Washington Post, the New York Times, the Guardian, and der Spiegel - have possessed large parts of the Snowden archive since 2013.

But all of them stopped working completely on the Snowden archive and stopped publishing Snowden documents many years ago, presumably because they decided that the massive financial and human resources required to work with the archive - including elaborate security measures to protect it and large teams of journalists and editors to curate and responsibly report on it - were no longer justified by the journalistic value the archive provided as the remaining documents aged.

I'm proud of the fact that long after all those other large media outlets completely stopped their reporting on the Snowden documents, the Intercept, with the full support of First Look Media, continued for five years to devote enormous resources to publishing and reporting on them, with teams of highly devoted and skilled journalists, researchers, tech experts and security specialists ensuring that the reporting continued in the responsible and incremental manner demanded by our source. During those years, the Intercept also expended substantial resources to create a secure environment in which outside experts, journalists, and researchers could have full access to the Snowden archive to ensure that all newsworthy material was identified and reported.

Like all digital media outlets, the Intercept has been confronted with financial constraints. The budget given to the Intercept by First Look Media for 2019 forced its editor-in-chief Betsy Reed, in consultation with the Intercept's senior editors, to make extremely difficult decisions about how best to allocate these limited budgetary resources to maximize the impact and value of the Intercept's journalism.

The Intercept's decision to stop working with the archive was the by-product of those financial constraints and, critically, of my intent to seek other partners - particularly academic institutions and research facilities - to ensure continued publication of the remaining Snowden documents that are in the public interest. Six years after we first began aggressively reporting on that archive - publishing thousands of top secret and classified documents all over the world despite serious government threats - the archive is now far less of a news resource and far more of a historical asset, which is why I believe academics and researchers, not reporters, are now best equipped to oversee its publication.

Critically, the Intercept is not the only entity that possessed the full Snowden archive. Both Laura Poitras and myself, individually and independently, continue to possess full copies of the archive, as do other individuals and institutions. They are all free to do what they wish with their copy of the archive. Speaking only for myself, I have spent the last severals months seeking to ensure that publication of these materials continues under the auspices of experts most competent to do this work, and who work with institutions that have the ample funds required to do so robustly, quickly and responsibly.

Finally, it is worth remembering that Edward Snowden never wanted the full archive of documents to be published; to the contrary, he adamantly insisted - not just privately but publicly - that there never be a full dumping of the archive. Instead, he insisted from the start that journalists work with teams of editors to carefully curate the archive and only release documents in the public interest and to protect people's reputations, privacy and security. Had he wanted the full archive indiscriminately published, he could have easily done that himself by uploading it to the internet back in 2013, or by providing it to some other organization with instruction that it all be released. He didn't do that. Instead, he has repeatedly stated, as recently as 2018, that he is extremely proud of the work done by the journalists with whom he chose to work on how these materials have been reported, and that it was done in accordance with the principles he insisted on from the outset.

I, too, am proud of the work the Intercept has done over five years on this archive. It took substantial courage, risk and resources to do this reporting, and the Intercept never flinched from doing it. I'm am grateful to my incredibly devoted and skilled colleagues who ensured that this complex, challenging reporting was so effectively carried out in the public interest.

[13]:
Menschen Machen Medien, “Snowden und die große Datenmisshandlung“, Ralf Hutter, June 2023. Translated:

Appelbaum, who was living in Berlin at the time, also criticizes the fact that the Guardian did not inform him and Poitras of the operation, to which he had previously agreed: “What if there had been coordinated raids across borders on everyone working on the material?”

LoganCIJ16. Reports from the Front, 2016, timestamp 1:23:42.

When the Guardian was raided, they did not call myself or Laura Poitras here in Germany to tell us that the GCHQ and other political powers and police powers in the UK had in fact come to destroy source material. They did not tell us. We had to find out in public. They left us to hang in public. They did not treat us as equals. They did not protect us. They did not care. And they continued with this. Every step of the way.

In America that has been with the White House, with the Director of National Intelligence, with the FBI, with the NSA, with the National Security Council and with the Pentagon. In this country it has included Downing Street, the Cabinet Office, the National Security Advisor, GCHQ themselves and the DA-Notice Committee.

[15]:
Breaking News, p. 221.
[16]:
Note that the decision to stop publishing documents was made by an editor-in-chief who couldn’t say how much of the archive had been read, indicating that nobody at the paper had a clear picture what was in the rest of the archive.
[17]:
Breaking News, p. 418.

In time, we even invited him [Vallance] into our morning conference.

Q256 Dr Huppert: Did he give you any feedback as to whether what you are publishing posed a risk to life or not?

Alan Rusbridger: He was quite explicit that nothing we had seen contravened national security in terms of risking life. He was explicit about that. That is not to say he would give us a complete bill of health on things that appeared downstream, but nothing he saw had risk to life and most of the time when we have rung him and put stories to him his response is, “There is nothing that concerns me there. This stuff might be politically embarrassing, but there is nothing here that is risking national security”.

[19]:
Dark Mirror. Notes, “Chapter Four: PRISM”, note 58 “the Post and I assembled lawyers.”
[20]:
Breaking News, p. 221.
[21]:
Breaking News, p. 232.
[22]:
Breaking News, p. 241.
[23]:
Breaking News, p. 227.
[24]:
Dark Mirror, p. 142.

The Snowden files, as it happened, were at that time locked in a Washington Post vault room and kept separate from their keys

[25]:
Dark Mirror, p. 68:

The Post could not handle a story this sensitive in anything like a normal newsroom environment. […] When working with the source material, the Post team would need dedicated computers with freshly wiped, encrypted hard drives. Networking hardware should be physically removed from those machines, cutting them off from the internet and newsroom production systems. Baron would have to find us a windowless room with a high-security lock, reinforced door, and heavy safe bolted to the floor. Decryption key files, stored on memory cards, would never be in the same room except when in use. You don’t have to write this stuff down, I said. I brought a list. Once these precautions were in place, access to the classified material would require four credentials: door key, safe combination, digital key card, and passphrases. We would divide the credentials among team members. No one but me would have all of them.

Dark Mirror, p. 143:

I had firmly requested a separate, locked room at the Post for use by the reporters who worked with the Snowden documents. On a subsequent visit, a facilities staff member proudly showed me the new space in a place of honor beside the company president’s office. The room had one feature I had specifically asked to avoid: a wall full of windows. If you craned your neck you could catch a glimpseof the Beaux-Arts mansion half a block to the west. The Russian ambassador’s residence in Washington. “You have to be kidding me,” Ashkan said. Crestfallen, I asked for a change of venue to a windowless space. The Post dutifully found one, installed a high-security lock, put a video camera in the hall outside, and brought in a huge safe that must have weighed four hundred pounds. I acquired a big, heavy safe in New York as well. I will not enumerate every step I took to keep my work secure, but they were many and varied and sometimes self-befuddling. The computers we used for the NSA archive were specially locked down. Ashkan and I cracked open a pair of laptops, removed the wi-fi and Bluetooth hardware, and disconnected the batteries. If a stranger appeared at the door, we merely had to tug on the quick-release power cables to switch off and reencrypt the machines instantly. We stored the laptops in the vault and kept encryption keys on hardware, itself encrypted, that we took away with us each time we left the room, even for bathroom breaks. We sealed the USB ports. I disconnected and locked up the internet router switch in my New York office every night. I dabbed epoxy and glitter on the case-bottom screws of all my machines to help detect tampering in my absence. (The glitter dries in random, unique patterns.) Detection of compromise was as important as prevention, security expert Nicholas Weaver told me, so I experimented with ultraviolet powder on the dial of the New York safe. Photographing dust patterns under a UV flashlight beam turned out to be messy. I kept my notes on multiple encrypted volumes, arranging the files in such a way that I had to type five long passphrases just to start work every day. I hardly ever typed all the passphrases right the first time. I forgot the passphrase to one seldom-used PGP key and lost access to a few of my files forever.

[26]:
See here, curated by us.
[27]:
Dark Mirror, p. 142.

in the late fall of 2015, Soltani and I had stopped writing stories for the Post. I was reporting for this book. Soltani had moved on. He had retired his old laptop, returned an encryption key fob to me, and shed his last connection to classified materials.

[28]:
Der NSA-Komplex, Chapter 2 Die Flucht, “Kontaktaufnahme, zweiter Versuch.”
[29]:
See here, curated by us.
[30]:
The full email is not published in the Daily Beast story. The author of the story, Max Tani, published the email in his Twitter account (archived1, archived2). Poitras also published the email herself two years later. The email is attached here as text:

On 13.03.19 22:04, Michael Bloom wrote:

Team:

This was a difficult day. As numerous media outlets have discovered over the last several years, the current business environment sometimes requires painful decisions, and it is always agonizing when reorganization results in the loss of colleagues' jobs.

The Intercept is proud of its reporting on the Snowden archive, and we are thankful to Laura Poitras and Glenn Greenwald for making it available to us. It is crucial to recall that major news outlets that possessed large portions of the Snowden archive – newsrooms much larger than The Intercept's – ceased reporting on it years ago. Many decided that the resources required to continue to work on the archive were not justified by the journalistic value the remaining documents provide, as those documents have aged. For five years, the company expended substantial resources to continue to report on the Snowden archive, but The Intercept has now decided to focus on other editorial priorities.

It is our hope that Glenn and Laura are able to find a new partner – such as an academic institution or research facility – that will continue to report on and publish the documents in the archive consistent with the public interest.

Best,

Michael

[31]:
Der NSA-Komplex, Chapter 8 Wir Überwachten.

Es ist allerdings der bislang aktuellste (die Materialien stammen teils aus dem April 2013) […]

[32]:
The 1% originates from Rusbridger’s evidence to the Home Affairs Committee on 3 December 2013, where it describes the Guardian’s own output, but in the same session he put the count at 26 documents out of 58,000-plus, or 0.045%. MacAskill used 1% again in 2023 for the partner group’s combined output. U.S. counterintelligence official Bill Evanina said in 2018 that “journalists have released only about 1 percent taken by the 34-year-old American.”
[33]:
On Twitter, 14 March 2019 (archived).

This has all been publicly discussed many times. Many large news orgs – WPost, NYT, Guardian, Der Spiegel – have huge portions of the archive, but stopped reporting on it years ago for cost reasons. Only TI, with a much smaller budget, continued to spend the resources to publish.

[34]:
As far as can be established, almost everything published after 2015 came from The Intercept or from collaborations it had with other outlets. Only independent exceptions seem to be Boing Boing in February 2016, and GCHQ images from the ANARCHIST programme in Poitras’s book Astro Noise, also published in February 2016.
[35]:
All are described above; see The Guardian and The New York Times and ProPublica sections.
[36]:
Described above; see section The Intercept.
[37]:
Motherboard/VICE, 22 April 2019. CYBER podcast, episode “Edward Snowden on Julian Assange, the Mueller Report, and Press Freedom“. Around timestamp 44:45:

[Ben Makuch]: One thing I wanted to ask you speaking of journalistic institutions. Were you disappointed to see when The Intercept closed down the Snowden archive?
[Edward Snowden]: I think the most disappointing thing about this was the fact that I learned about it from the news. Like look. I am a source. I am not a journalist, I don't work at The Intercept, I'm not on the board of First Look. They don't owe me a vote on any of this. And I understand that, right? And they're also not the sole custodians of this. Barton Gellman of the Washington Post, he's got a copy, the New York Times has copies, the Guardian has copies, Der Spiegel has copies, right? This is distributed. All these other news organizations have stopped publishing for years because it became more expensive and longer form work, instead of like the short punchy news pieces which is what, the large appetite is in journalism today. What remains in the archive, I believe, is stuff that is going to require much more substantial effort, basically book length work. You need real researchers to sit down, and to connect all the dots that you can't get across in 750 words. And The Intercept really was not designed for that, unfortunately. None of these news organizations were, and they were supposed to hand this off to academic institutions, but that just hasn't happened because the academic institutions get cold feet. They get nervous. They go, look, we're dependent on grants from the federal government in the U.S., right? And they don't want to give it to a foreign university because the politics of that, they're worried the government's going to complain and go, oh, you know, you're giving this to whoever. So I am sympathetic and I understand the ways their hands are bound. At the same time, I don't think it would have been a lot to be to ask for, hey, you know, could you guys give me a call? Maybe ask my feelings on it...
[Ben Makuch]: Just a quick Signal text.
[Edward Snowden]: ... I've since talked to them. And I understand where it's at. But yeah, you know, I think anybody would agree it wasn't well handled.
[Ben Makuch]: Yeah, you know, just a quick Signal text, or you can Skype.
[Edward Snowden]: Right, right, right.

[38]:
Ironically a fifth of that list have closed their Snowden archive and/or stopped publishing documents they’ve had access to.

Extra notes

  • Electrospaces, otherwise very careful public account of the archive’s distribution, states that Snowden posted copies to four individuals, but the Harper‘s article describes one package, sent to Bruder, with copies distributed afterwards by Poitras. Also worth noting that the counts don’t quite reconcile in the article. It says “there were five of us” and that “three other volunteers had received duplicates” – which would place Maharidge outside the three since those three are said in the article to be Timm, the person who asked not to be identified, and the person the authors couldn’t identify. Yet Maharidge ends up holding a copy too. Whether four copies were made rather than three, and who made his, the article doesn’t say. The ambiguousness might be deliberate.

  • Snowden supplied a blurb for Breaking News, calling Rusbridger “a fearless defender of the public interest.” By then the Guardian had published about 30 of the 58,000 documents it held.


Source: Hacker News

Software Sandboxing: The Basics (2025)

Diving into the territory of software sandboxing is diving into mostly uncharted
territory. The necessary pieces to implement good sandboxing in your software
are scattered all-around and the pioneers haven’t yet gathered enough knowledge
into an unified mappa mundi that can guide new sailors through some well
understood safe routes. In this blog post I’ll offer my own share of experiences
that I have acquired while working on sandboxing support for Emilua. Writing
style will suffer a little because I’ll err on the side of repeating myself too
much to avoid any misunderstandings.

First, let’s get some informal (but useful) definition for sandboxing just to
make sure we’re on the same page. Here’s
the
definition that was used by Julien Tinnes and Chris Evans at Hack In The Box
Malaysia 2009:

The ability to restrict a process' privileges:

Without administrative authority on the machine;

Discretionary privilege dropping.

That’s a very good definition to keep the ball rolling. Let’s quickly iterate
over each point individually to make them crystal clear. However keep in mind
that the opinions I possess today are a little different from the
opinions J. Tinnes and C. Evans had during the 2009 talk (especially around “is
it okay to use superuser APIs?”), so my explanations will differ a little and
guide you towards what I consider better practices for 2025.

OSes present different interfaces to users and software developers. System
administrators traditionally rely on filesystem permissions to isolate services
(UNIX daemons). If we allowed third-party programs to freely change such
permissions then it’d nullify the policies the sysadmin was trying to enforce to
begin with.

Furthermore third-party programs abstract their own virtual worlds and most of
the time UNIX filesystem permissions aren’t a good fit to model the security
policies such other virtual worlds require. Do you use UNIX permission modes to
define who can see your Twitter feed or message you on Identi.ca? Filesystem
permissions aren’t the only knobs sysadmins possess to restrict access rights,
but the reasoning developed here also apply to these other knobs.

Nonetheless a process inevitably runs on top of an OS and there are
kernel-exposed resources the process interacts with (e.g. files). It’s this
interface that matters to the software developer. Web browsers such as Firefox
run DRM plugins and it’s desirable to run such third-party plugins without
allowing them to have full access to every file that Firefox has access to
(usually every file in the user’s HOME directory). Traditional tools such as
setuidgid can’t help here and their usefulness is limited as interfaces
sysadmins turn to. setuidgid and similar tools aren’t interfaces intended for
the software developer to use.

For programmatic privilege dropping, traditional UNIX interfaces are a poor
match, and OSes where this gap actually matters will provide extended interfaces
that go beyond traditional UNIX (e.g. FreeBSD’s Capsicum and Linux’s Seccomp).

When good interfaces for sandboxing weren’t available, programmers found their
way to create sandboxes anyway by abusing mechanisms available only to the
superuser. The most emblematic technique in this class is a helper suid binary
that’ll configure a chroot jail.

The obvious problem with these approaches is that they aren’t available to all
programs. Allowing any program to install suid binaries defeat any security
measures. Suid binaries equal to temporally raising privileges to full
administrative authority over the system. Privileges should only ever decrease,
never increase (principle of least privilege).

Another related concern here is to not design APIs that backfire by
exponentially increasing the kernel attack surface. The Docker boom popularized
Linux namespaces as a mechanism to cheaply isolate services. However within a
nested user namespace, the process runs as superuser (within that namespace),
and code paths within the kernel that would normally only be available to the
superuser are now available to every user. We have over a decade of kernel code
that was never written with this premise in mind. This decision caused security
problems in the past, and it’s bound to happen again. To quote Andy Lutomirski:

I consider the ability to use CLONE_NEWUSER to acquire CAP_NET_ADMIN over
any network namespace and to thus access the network configuration API to be a
huge risk. For example, unprivileged users can program iptables. I’ll eat my hat
if there are no privilege escalations in there.

It’s fine to allow user namespaces as long as you restrict this interface to
trusted containerization tools (e.g. Docker). However Linux namespaces is a
terrible interface for software sandboxing. Newer sandboxing interfaces in Linux
such as Landlock were carefully designed to not exponentially increase the
kernel attack surface as to avoid the disasters we’ve seen with Linux’s user
namespaces. Moreover new ways to restrict
namespaces within Linux are still being developed and long-term it’s a bad bet
to rely on them as a general sandboxing mechanism.

The first few years of software sandboxing research I’ve put into Emilua were
solely focused on Linux namespaces. After a lot of frustration the focus shifted
towards different solutions. Nowadays Emilua still offers support for Linux
namespaces, but the intended use-case now is the creation of containerization
tools. For proper sandboxing within Emilua, you’ll use mechanisms other than
Linux namespaces.

Actually sandboxes might also be defined as:

A restricted, controlled execution environment that prevents potentially
malicious software […​] from accessing any system resources except those for
which the software is authorized.

There’s no actual consensus over what traits are required for some code to be
considered sandboxed and definitions are usually very loose. These definitions
don’t require the properties we’ve been discussing so far. Therefore a different
term altogether might come in handy. J. Tinnes suggested “discretionary
privilege dropping”. That’s the type of sandboxing we’ll be looking into for
this article.

Discretionary privilege dropping doesn’t replace system administration
policies. Rather they complement each other and should be adopted in tandem.

Now we’re hopefully on the same page. Sandbox for us mean the same thing:
discretionary privilege dropping. How do we go from an unsandboxed program to a
sandboxed one on existing real-world OSes? In every mainstream OS today, the
privilege boundary lies at the process level. Credentials are associated with
each process and that’s what the kernel checks to decide whether the process can
acquire new resources using ambient authority.

Linux is actually different and associates credentials at the thread level, but
a design rooted at the thread level cannot work, and that’s why
glibc will do extra work to synchronize credentials
across threads even if the kernel is sloppy about
it. GNOME
developers thought they could work at the thread level just to be proved wrong
with CVE-2023-43641.

Adam Langley actually described a mechanism that in theory can work at the
thread level, but in practice is economically too costly and I don’t think it’ll
ever work:

So that’s what we do: each untrusted thread has a trusted helper thread running
in the same process. This certainly presents a fairly hostile environment for
the trusted code to run in. For one, it can only trust its CPU registers – all
memory must be assumed to be hostile. Since C code will spill to the stack when
needed and may pass arguments on the stack, all the code for the trusted thread
has to carefully written in assembly.

The trusted thread can receive requests to make system calls from the untrusted
thread over a socket pair, validate the system call number and perform them on
its behalf. We can stop the untrusted thread from breaking out by only using CPU
registers and by refusing to let the untrusted code manipulate the VM in unsafe
ways with mmap, mprotect etc.

Let’s not theorize over what alternative designs could work. For today,
processes is what we got. Once we compartmentalise our program as separate
processes, we can proceed to the next steps:

Assigning different privileges to each compartment (the processes).

Handling communication among the compartments.

Researchers from FreeBSD’s Capsicum already had the right mental model to
develop sandboxes for well over a decade:

Compartmentalised application development is, of necessity, distributed
application development, with software components running in different processes
and communicating via message passing.

The means to drop privileges are different on every platform, so we’ll skip this
for now and get back to it later. First let’s focus on the problem of
distributed application development.

The actor model is one of the most well known patterns for the development of
distributed systems. Erlang is perhaps its most iconic user. However Erlang’s
interest in the actor model lies in high availability and fault
tolerance. Nonetheless it’s still useful to look into widely used models even if
we’re not interested in high availability nor fault tolerance.

It’s common for many explanations of the actor model to quickly step into the
world of mathematics (which is fine). However many of them quickly become lost
into the world of abstraction and forget about computers entirely (which is not
fine). So let’s just use a summary of the points we care about in the actor
model:

Actors can manage their own internal state.

Actors can spawn other actors.

Actors can send messages to other actors.

Actors can include the addresses of other actors in messages.

If we summarize the actor model into concrete design choices within our
programming language or framework, here’s what we care about:

There is a function to create actors. This function returns the address of the
new actor.

The address of an actor can be used to send messages.

The address of an actor can also be a message or part of a larger message.

There is a function to receive messages. This function read messages that are
enqueued for the calling actor.

It’s possible to retrieve the address of the current actor.

Actors share no memory with each other.

An actor doesn’t run in parallel to itself. If an actor is currently running
in thread A, it can’t also be running in thread B. However it’s fine for
actors to jump from one thread to another (as in work-stealing threaded task
schedulers). That’s
the same property that Boost.Asio describe as strands.

For Emilua, this design translates into 3 functions:

If you can learn just 3 functions, you can code for the actor model. Let’s go
over some examples now:

If we decide to use the actor model for sandboxing, then each process will be an
actor. UNIX domain sockets can be used for actor messaging. Upon spawning a new
actor, we setup socket inheritance so we can communicate with it. The socket
will be the actor address. We also need to be able to include the addresses of
other actors in messages, but this is also covered because
it’s possible to send file
descriptors over UNIX domain sockets. The inbox file descriptor is never sent
to other actors (i.e. we have a MPSC channel).

Emilua has many implementations for the actor model, so we must explicitly
instruct it to use subprocesses upon spawning a new actor:

This design also solves another problem for our sandboxing concerns: handing
resources over to restricted processes. “Everything is a file (descriptor)” is
one of the most well known phrases within the UNIX culture. If we can send file
descriptors then we have a really broad range of resources that we can work with
from sandboxed processes. To mention just a few:

Device nodes (e.g. /dev/random, GPU communication, …​).

Process handles — pidfds, procdescs.

These are the resources we care about when we sandbox programs. These are the
resources we’ll take into account when we develop our security models. If we can
prove that we aren’t leaking file descriptors to the wrong actors then we can
use the actor model. Fortunately there’s a well researched model that solves
this problem for us: capability-based security. There’s even a programming
language based on the actor model and capability-based security:
the Pony programming language.

There’s only one small gap that we need to fill to combine both models:
capability based security assumes unforgeable tokens, but the actor model uses
addresses (which are forgeable). In our case, this problem was already solved by
the use of channels instead of addresses. The API stays the same and nobody will
notice a thing. Now we can use capabilities to reason about questions such as:

Is it possible for actor A to have effective access to resource X?

How can we design a layout that makes it impossible for any sandboxed actor to
simultaneously have access to files and sockets?

As for the usage of file descriptors as capabilities, the rule of thumb would be
to avoid ioctls, but we’ll be back to this topic later.

The actor model is simple to use, but very powerful. The ability to include the
addresses of other actors in messages means arbitrarily variable topologies. On
most of my own projects, I restrict myself to tree topologies, but the moment a
tree becomes unfit for my project, it’ll be easily replaced by a different
topology. So far I haven’t stumbled on a single sandboxed application that can’t
be modeled using actors.

If you need guidelines on how to develop distributed applications using the
actor model, you’ll enjoy several decades of R&D that’ve gone into it. Whether
you prefer books, small tutorials, face-to-face classes, study groups, or many
other learning approaches, you’ll likely find something of use.

Now that we have messaging solved with the actor model, let’s jump into
sandboxing (security models) again. There are other properties an object must
have so it can be modeled as a capability. A capability isn’t only a reference
to a resource, but the associated access rights as well. Owning a capability is
the same as also having access rights to perform actions. With this in mind, we
need to ponder:

Can file descriptors be modeled as capabilities?

What precautions must we take to use file descriptors as capabilities?

Generally UNIX systems run permission checks to grant or deny access only when a
new file descriptor is created, not when existing file descriptors are
used. This behavior is compatible with capabilities. Here’s the code for a
sample program:

And the output when I run the program as root:

And the output when I run the program as any other user:

This is just the UNIX behavior I was mentioning. Now let’s run some shell command as root:

And the same command as a different user:

Nothing surprising here. It’s just the same behavior. Now let’s run grep as an
unprivileged user, but making sure it inherited a file descriptor opened by
root:

As stated earlier, UNIX systems generally don’t run permission checks when
performing actions on existing file descriptors. That’s why grep succeeded to
read the file contents in the example. This was true for the previous example
(the action read on a regular file), but will it always be true? We could be
afraid of new kernel versions. They could always introduce a new syscall that
breaks this convention. However early in UNIX history the concept of suid
binaries was introduced, and that’s a legacy that’ll keep haunting kernel
developers to make sure they don’t break this convention. Let’s explore suid
binaries now.

In the last example we had the superuser using the syscall setresuid to change
the process credentials. Now we’ll walk into the opposite direction, creating a
child privileged process from an unprivileged one. This is allowed only for suid
binaries, so the privileged process will only ever run programs trusted by the
sysadmin. One of such programs is su:

This example shows that we can easily trick suid binaries to read or write into
any file descriptors that we have by using simple fd inheritance. For this
example, it wrote the string “Password: su: Authentication token manipulation
error” using the credentials of a privileged process. If the credentials of the
writer process had any importance for the security of the system, every UNIX
system would be broken already. Thefefore new interfaces are always designed in
a way that the credentials of the writer process don’t matter at all.

As a recent example to further stress the point,
some of the last syscalls that Linux introduced
were related to filesystem mounting. The initial versions of the contributed
patchset were rejected due to the use of the syscall write in operations
that’d use the credentials of the calling process for permission
checks. Eventually the contributor changed the design by using the new syscall
fsconfig and the patchset was accepted.

It’s important to notice that kernel developers will respect the convention
whether suid binaries are allowed on our Linux distro or not. Even if we block
suid binaries from our OS entirely, we can still assume that no attacker will be
able to gain new privileges by using our process as a proxy to perform some
dangerous write operation (the attacker could just write into the file
descriptor directly instead and the effects would’ve been the same).

The exception to this rule are ioctls. Performing ioctls on fds received from
untrusted processes is always dangerous. Emilua relies on Boost.Asio for async
IO, and Boost.Asio used to rely on FIONBIO when it shouldn’t. After a few
email exchanges, I managed to persuade Christopher Kohlhoff to change this
behavior and
now
Boost.Asio will do the right thing as long as you’re at least on Boost 1.86.
By the way, even isatty() is — at least on Linux — implemented as an ioctl,
so you really need to be careful about non-standard operations.

Awesome. We can indeed model file descriptors as capabilities, but we weren’t
the first to reach this conclusion.

Capsicum is an interface to better support the use of file descriptors as
capabilities that is part of FreeBSD since its 9.0 release. One of the
facilities offered by Capsicum is the function cap_enter. cap_enter drops
process privileges by disabling ambient authority entirely.

That’s it. One function call and we dropped privileges. All system accesses will
have to be performed through open file descriptors. If we don’t already have
access to some resource, the only way to get it now is through inbox. If we
attempt to open files, open will fail because ambient authority is
disabled. If we attempt to connect a socket to some endpoint, the operation will
fail because ambient authority is disabled. That’s the beauty of Capsicum: we
deny access to external resources because the names themselves that could be
used to refer to resources become unavailable.

When
the Capsicum research was published, the following table was also presented:

setuid root helper sandboxes renderer

Restricted sandbox type enforcement domain

seccomp and userspace syscall wrapper

Capsicum sandboxing using cap_enter

At that time, its researchers have modified Chromium to make use of Capsicum and
compared how much effort was required to make use of each sandboxing mechanism
within Chromium. Capsicum required only 100 lines of code. Compare that to
seccomp’s 11301 lines or Windows' 22350 lines. The other mechanisms compared
didn’t actually restrict the sandboxes significantly and can be disregarded. If
you’re only going to study one sandboxing mechanism in your life, it should be
Capsicum. To this date, I have yet to see a better sandboxing mechanism than
Capsicum.

Capsicum also provides finer grained access control to file descriptors. As an
example, one may use Capsicum to allow one process to wait on a semaphore, but
not to post on it. A file descriptor is usually created with all rights
assigned. Then these rights can be reduced through the use of the function
cap_rights_limit.

Although
open() won’t work in Capsicum mode, openat() will, and Capsicum will make
sure the relative paths are only resolved to a hierarchy beneath the given
directory-fd. If only we could force Linux syscalls to always include
RESOLVE_BENEATH…​ but maybe we can? Keep reading until we’re back at this
topic.

I haven’t actually
faced many problems using Capsicum so the only complaint I’ve had was fixed long
ago. There’s not much to talk about Capsicum. The system is incredibly simple
to use yet powerful. This model will be the inspiration for all sandboxes
developed in the rest of this article no matter the OS.

Once we do receive a file descriptor from a sandbox, it’s time to operate
on it. However if we’re sloppy about it, our thread will block. We need to avoid
blocking operations to dodge some DoS
attempts. The mess around non-blocking IO on
UNIX has been long known. Believe me when I say I have my own share of comments
to make here, but this article isn’t about async IO and the text is already
getting too long, so I’ll just present a boring summary of what you need to
know:

close() may block according to POSIX.

Supposedly close() always succeed on
Linux. Don’t bother checking for errors here.

You
may need to create a thread to run close() if “slow-to-close” files are a
problem.

Use fstat on the received file descriptor to check whether it’s a socket.

On sockets, use MSG_DONTWAIT on recv(). Boost.Asio still gets this wrong,
but I’m short on time to send bugfixes anytime soon.

On non-sockets, use a proactor (completion events opposed to readiness events)
to perform IO operations (e.g. io_uring on Linux, POSIX AIO on FreeBSD, …​).

On FreeBSD, you can
now use aio_read2() with AIO_OP2_FOFFSET to read without an offset. Other
approaches will likely fail with ENOTCAPABLE. Boost.Asio also gets this
wrong, and I’m short on time to send bugfixes here too.

On
Linux, io_uring is widely distrusted and disabled. Therefore you might just
as well reject non-socket IO on Linux if the file descriptor was received from
a sandboxed process.

By now you should understand that there are actually two use cases:

Trusted process creates a resource (file descriptor) and sends it to a
distrusted sandboxed process.

The distrusted sandboxed process creates a resource and sends it elsewhere.

If the file descriptor was created by a trusted process then the range of
options we can work with widens. However state such as O_NONBLOCK is shared
among all copies of a file descriptor and become a problem the moment the file
descriptor reaches the first sandboxed process. On Capsicum we can at least
forbid F_SETFL and alleviate the problem slightly, but this approach only
works for FreeBSD.

By now you should have enough tools in your toolbox to sandbox all your future
code. However there’s a reason why we sandbox code. Projects grow high and large
until it becomes impossible to ensure their code is…​ bug-free. A little bug in
just one of the dozens of modules within a project shouldn’t equal to a fully
compromised system when the bug is exploited by a hacker. Damage should be
contained. Security policies implemented by OS tools external to your code can
only work at a program (or user) level. The program will still be fully
compromised and the hacker will have access to all data and credentials the
program has access to.

When your program deals with data that can be processed independently, there’s
an opportunity to implement a safer approach. If you can run multiple instances
of your program as different users on your OS, you can use existing security
solutions in your project. However if the concept of allocating defined portions
of the data to fixed users doesn’t work for your project, you may need something
more complex or custom-tailored. When the relationships between the users are
blurry and your project demands policies that are more dynamic, you may need
sandboxes.

As a rule of thumb, every shell should have sandboxes. Shell are programs that
act as the membrane that sits between the human operator and some virtual
world. Tablets, smartphones, and laptops display graphical shells to interact
with programs, windows, and files. Servers employ textual shells. Likewise web
browsers act as the shells to the www world.

I wouldn’t be surprised if Firefox and Chrome were the only software employing
discretionary privilege dropping that you know. They are shells after all so it
matters to them. More than that, they’re very well funded projects. Sandboxing
used to be very expensive (especially outside FreeBSD). However these shouldn’t
be the only software out there with builtin sandboxing support. Take Telegram,
for instance. The right media parsing bug could mean a hacker having access to
all my chat history. What time does my son leave school? What people do I trust
my credit card info with? When will I go in a trip and leave my house
unattended? These are just a few examples of the damage that might be done due
to the lack of sandboxes in Telegram. Not only Telegram, but every instant
messenger should be employing sandboxes. Media parsing should always be
performed in dedicated sandboxes.

The first step into this direction is a realistic approach to real-world
engineering: let’s not rewrite all code from scratch. Deal? The tricks you
learned earlier will still be useful, but from now on I’ll share tricks to work
on existing real-world code. Capsicum users refer to the ability to run
unmodified code within sandboxes as oblivious sandboxing. Techniques for
oblivious sandboxing most often than not have nothing to do with discretionary
privilege dropping and can’t solve the problems we were mentioning just a
second ago. However it’s possible to combine approaches from both worlds in the
same project so it’s important to study the techniques for oblivious
sandboxing too.

The one place we need to look at to implement oblivious sandboxing is actually
pretty obvious: the ambient authority functions. In fact, that’s what projects
such as Super Capsicumizer
9000 do. They inject a dynamic library into a process using LD_PRELOAD to
interpose ambient authority calls. This technique is actually yesterday news and
projects such as fakeroot have been using it for decades.

Super Capsicumizer 9000 is actually a small experiment hacked together by a very
very small team. The experiment succeeded into opening old software built on top
of complex libraries with a long history of changes. This is very
promising. It’s a sign that maybe a single programmer working alone to interpose
just a few functions for ambient authority access will have success in running
legacy code.

Programmers almost never do syscalls directly, and instead rely on libc to do
the syscalls on their behalf. That’s why this approach works so well. All you
have to do is to write a definition for the function from libc you want to
interpose. If you’re linking against the dynamic libc, your function will be
loaded first and used instead. If you’re linking against the static libc,
chances are that the libc symbol is actually a weak symbol so it’ll be dropped
once the static linker see your definition. Emilua has been using this approach
to support dynamic and static executables on Linux and FreeBSD and so far
getaddrinfo was the only ambient authority function whose symbol lacked the
attribute for weak symbols (please comment on the linked bug reports if you plan
to build your own sandboxes using the same techniques or using Emilua):

https://sourceware.org/bugzilla/show_bug.cgi?id=32509.

https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=283528.

The next step is choosing which functions to interpose. Functions from FreeBSD’s
libcasper are good first candidates. However for some reason libcasper doesn’t
interpose the functions it intends to replace so you’ll need to change names and
parameters accordingly. libcasper functions (e.g. cap_getaddrinfo) always take
an extra parameter. Another good source of inspiration to decide which functions
to interpose is the library used by Super Capsicumizer 9000: libpreopen. Most of
the time, you’ll only need to interpose a few functions even for complex
projects.

Chromium renderers need very little authority. They need access to fontconfig to
find fonts on the system and to open those font files.

Emilua 0.11 abstracts all these details into the module libc_service. The
example below shows how we use this module to override the behavior of open to
return a rogue file descriptor when the subprocess try to open
/dev/null. Actual sandboxing setup (i.e. privilege dropping within the new
subprocess) is omitted for brevity. The example also shows how to prefill the
code cache for the new subprocess so it won’t query the filesystem to fetch the
Lua code to execute.

Emilua uses UNIX sockets behind the scenes for communication between both
processes. This approach allows one to implement fully dynamic security
policies. For instance, if you’re trying use Telegram’s tdlib to implement your
own Telegram client, you could have the following rules for your secutiry
policy:

Only resolve name queries to pluto.web.telegram.org.

Only allow connect requests to the IP addresses we resolved in previous steps.

The little Lua script we send to be executed in the sandboxed side means we can
apply some simple call fixups at the call site to further broaden the use cases
we can tackle. For instance, when sandboxed code try to open a GUI connecting to
/tmp/.X11-unix/X0, we can send a new file descriptor to an unrelated display
server and replace the socket from the original request with the new one using
dup2 from the Lua script at the call site. In fact, we can do that:

This example also shows that Emilua can make use of LD_PRELOAD to perform libc
interposition on existing programs such as xterm.

Another interesting approach that might prove useful to your projects is to use
kcmp in your security policies. This technique would allow you to increase
your policies granularities even further by implementing different subpolicies
for each file descriptor.

By the way, now we can interpose openat() on Linux to make it work as in
FreeBSD’s capability mode (but — as usual — don’t forget to forbid the actual
syscall):

These techniques are already probably more than what you need, but I have few
more tricks up in my sleeve to share, so let’s move on.

A common theme in the threat model of sandboxes is to assume that initially
trusted code becomes malicious once compromised. For instance, we might trust
that ffmpeg developers are well-intentioned and didn’t backdoored their
project. However ffmpeg is a complex project and legitimate bugs lurk around
just waiting to be found. Some of these bugs might be exploitable by hackers. In
this case we can import ffmpeg as a library in our executable and only setup the
sandbox right before we call ffmpeg functions on external data.

However how’d we approach the sandboxing steps if we assumed the code to be
compromised from the start? Let’s take Telegram’s tdlib as an example. Suppose
you don’t trust tdlib at all. Under this threat model, even just loading the
library would be a dangerous operation. For such scenario, first we need to
build tdlib under a secure environment (e.g. jails on FreeBSD, namespaces on
Linux). Once we do have the built plugin, we can move to the next challenge.

If we disable ambient authority to sandbox the code, dlopen() will fail to
access the filesystem. To work around this issue, we can dlopen by file
descriptors instead. On FreeBSD, we can use fdlopen. On Linux, we can pass a
path to /proc/self/fd/ as long as we never close the file descriptor to avoid
path reuse by a different plugin (glibc will deduplicate plugins by
path). That’s not a perfect solution, but it’s a start.

On some previous example, we saw code cache prefilling as an Emilua way to
instruct a policy to avoid filesystem queries. Emilua just follows the same
trend for plugins and exposes native_modules_cache to prefill the native
plugin cache:

The biggest problem with plugins is that they might depend on not yet loaded
dynamic libraries and dlopen would fail to load them. We can try to use the
workaround for libc-service that we saw in the previous section, but there are
cleaner solutions for this problem. On Linux, we can simply use
Landlock. Landlock would already be required for the /proc/self/fd/ trick
anyway. On FreeBSD, we can use rtld_set_var
and LIBRARY_PATH_FDS.

Fortunately we won’t need to worry about this for Tdlib so the code becomes
slightly simpler. However it begs the question: if we don’t trust tdlib, why are
we trusting it with our data anyway? The question here is to evaluate if the
threat model even makes sense. In the case of tdlib, we aren’t giving tdlib’s
developers (Telegram developers) anything they don’t already have (data on
Telegram servers). Running tdlib within a plugin prevents the Telegram company
to have unrestricted access to data on our systems. The answer for tdlib might
be simple, but that’s not what matters in this story. The lesson that you should
take here is to evaluate whether your threat models even make sense.

Let’s explore another threat model story. The Linux kernel can load compressed
initramfs images. Even if the Linux kernel uses a buggy unmaintained library
full of well known exploits to decompress the initramfs image, it doesn’t matter
at all! We only use this library on data generated by a trusted user. If some
adversarial actor had control over the initramfs images we load, the actor could
already do any damage he wished for even if the decompression library had zero
exploitable bugs. However the story would be completely different if the library
were backdoored. Sometimes it’s more important to have auditable code written by
trusted individuals than supposedly better code written by individuals we aren’t
sure we can trust.

I’ve postponed this section for as long as I could because it’s awful. Seccomp
is not a good mechanism for discretionary privilege dropping. Seccomp is a good
mechanism for OS hardening. Seccomp is a simple programmable syscall filtering
mechanism based on BPF programs. The BPF program must choose an action for each
syscall attempted by the process:

This mechanism can be used to deny access (e.g. SECCOMP_RET_ERRNO) to ambient
authority by disallowing syscalls that work on names (e.g. open, bind). All
you have to do is disable ambient authority, then your process will be properly
sandboxed and you can use the lessons learned in previous sections for
compartmentalised application development. Using this mechanism you may either
implement whitelists or blacklists. These projects implement syscall blacklists:

https://github.com/lxc/lxc/blob/v6.0.3/config/templates/common.seccomp

https://github.com/flatpak/flatpak/issues/4187#issuecomment-1075512546

https://github.com/flatpak/flatpak/blob/1.16.0/common/flatpak-run.c#L1839

The problem with blacklists is well known. New kernel versions may add new
syscalls. You don’t know today what syscalls will be added tomorrow. You don’t
know today if tomorrow’s syscall breaks your today’s policy. However the problem
is even worse for
seccomp. Linux
allows multiarch systems and syscall numbering varies wildly among arches. Even
if you block syscalls such as acct on your x86-64 program, a compromised
sandbox could bypass the syscall filter by running a x86 executable. To our
delight, Linux somehow manages to make the problem even worse:

The arch field is not unique for all calling conventions. The x86-64 ABI and
the x32 ABI both use AUDIT_ARCH_X86_64 as arch, and they run on the same
processors. Instead, the mask __X32_SYSCALL_BIT is used on the system call
number to tell the two ABIs apart.

This means that a policy must either deny all syscalls with X32_SYSCALL_BIT
or it must recognize syscalls with and without X32_SYSCALL_BIT set. A list
of system calls to be denied based on nr that does not also contain nr
values with X32_SYSCALL_BIT set can be bypassed by a malicious program that
sets X32_SYSCALL_BIT.

Even when we do migrate to whitelists, we’ll still be haunted by these
implementation details. Here are a few notable projects based on whitelists:

https://github.com/moby/moby/blob/v27.4.1/profiles/seccomp/default.json

https://github.com/containers/podman/blob/v5.3.1/vendor/github.com/containers/common/pkg/seccomp/seccomp.json

https://gitlab.gnome.org/GNOME/localsearch/-/blob/3.8.2/src/libtracker-miners-common/tracker-seccomp.c#L141

https://android.googlesource.com/platform/bionic/+/704772bda034448165d071f68b6aeca716f4220e/libc/seccomp/seccomp_policy.cpp

Once you do go through these lists to implement your own seccomp policies,
you’ll face another problem: Linux syscall parameter ordering changes among
arches. That’s why Docker uses multiple rules for the syscall clone. Why do we
have to become experts in Linux syscall conventions to do what FreeBSD users get
done in 10 seconds by just calling cap_enter()? Welcome to Linux.

The good news is that in theory we could implement some library with a single
function to disable ambient authority and everyone would then just use this
library. I have a few ideas on how to implement this library, but so far no
customer of mine was interested in this problem (at the same time I’m busy with
different projects).

However there’s just another caveat you must keep in mind. Linux userspace
relies on the filesystem too much. Even if you interpose open for
compatibility with legacy code, legacy code will fail trying to access
/proc/self when nested sandboxes are at play. It’s better to just allow open
and filter filesystem access with Landlock instead. The problem now is that file
descriptors can’t be modeled as capabilities if a process can just reopen
/proc/self/fd/ with a different mode. Landlock developers shared a long-term
goal to expose capabilities compatible with Capsicum, so maybe we’ll have a
solution to this problem in the future.

Seccomp’s complexity might be disheartening for the compartmentalised
application developer, but the situation is even worse if you look into Linux
namespaces. For Linux, Seccomp + Landlock is what we have.

Kafel is a language and library for specifying syscall filtering policies. The
policies are compiled into BPF code that can be used with seccomp-filter.

Kafel is the most promising project to enable Seccomp policy reusing that I’ve
seen during my quests. Using Kafel I’ve managed to define a few policy groups
that might be of interest to you:

These are the policies I’ve been using for almost a year. I’d make a few changes
nowadays, but I haven’t gotten the time to it yet. Also do keep in mind that
Kafel’s syscall database is inexcusably poor so you’ll need a few changes (some
sed-like preprocessing to replace the syscall names by their numbers) to make
the above work. Even when it does work, it’ll be omitting many syscalls that
projects using better syscall databases such as Docker handle. Did you notice
that we don’t mention clock_gettime64? That’s because Kafel’s syscall database
is just too poor. Kafel won’t ever be adopted by projects such as Docker for as
long as it retains such poor syscall databases and poor multiarch support.

However a lack of better syscall databases isn’t the only thing we could improve
in Kafel. I’d like to see policy versioning and better policy composition
operators, for instance. I’d be willing to develop these features, but again, no
current customer of mine was interested in this project, so my time will be
spent in different projects.

In this article, I’ve gone just over the basics for sandboxing in Linux and
FreeBSD. However there are more lessons that I’d like to share. I’ll save these
for a later article in the future. Topics that I’ll likely cover if I ever get
into the mood to write another blog post again:

More demos (for years I’ve been running graphical apps in some of my machines
solely in containers so I got quite a few demos to show).

More about the actor model, capability-based security, and access control
policies.

Maybe a comment or two on Windows and macOS.

PR_SET_DUMPABLE & /proc/sys/kernel/yama/ptrace_scope.

Containers (Emilua also works as a container runtime).

Attack surfaces & safe parsers.

UNIX tricks for the C programmer that Emilua makes use of.


Source: Hacker News

AX – Google’s Open Agentic Orchestrator

Declare an agentic task. AX runs it at scale.

AX sandboxes your task, wires up its workspace, fences its network, and helps you run billions of
them per cluster. Either use a single task per agent, or compose as many as your agent needs.


$ cat task.yaml
apiVersion: ax.io/v1alpha1
kind: Workspace
metadata:
  name: golang
spec:
  git:
    - repo: https://github.com/golang/go.git
      branch: "my-fix"
---
apiVersion: ax.io/v1alpha1
kind: Task
metadata:
  name: test
spec:
  workspaces:
    - name: golang
      goal: "Ensure that Go tool chain is available and is built from source"
  debug: true
$ ax apply -f task.yaml
workspace.ax.io/golang created
task.ax.io/test created
$ ax watch task test
Watching task default/test...
[10:42:01] Phase: Pending    Actor: test               WorkerIP:
[10:42:05] Phase: Running    Actor: test               WorkerIP: 10.20.3.67
Task reached terminal phase "Running".
$ ax get tasks
NAME   ATESPACE   PHASE     ACTOR   WORKER-IP    AGE
test   default    Running   test    10.20.3.67   5s
$ ax ssh test -- ls /workspace
go
$ ax ssh test -- cd /workspace/go && go build ./...
$ ax ssh test -- ps -o pid,cmd
  PID CMD
    1 /usr/local/bin/ax-task-runner
   12 go build ./...
$ ax ssh test -- touch notes.txt
$ ax suspend task test
task.ax.io/test suspended
$ ax resume task test
task.ax.io/test resumed
$ ax ssh test -- ls notes.txt
notes.txt
$ ax suspend task test
task.ax.io/test suspended
$ ax delete task test
task.ax.io/test deleted

How it works

Scales up to billions of tasks.

AX runs on top of Agent Substrate,
a compute runtime designed from the ground up for massive density and fast stateful actor lifecycles.

Billions of tasks

Every task runs as a lightweight actor, allowing you to scale to
billions of concurrent agent sessions per cluster without orchestrator limits.

Sub-second resumption

Idle agents waiting on model responses, external tool calls, or human responses are checkpointed, suspended,
and brought back
in under a second with zero cold-start delay.

Dense multiplexing

Dozens of tasks share worker resources, turning idle waiting time into spare compute capacity so you only pay
when agents are actively thinking and running code.

Generative platform

Generative features built into the platform.

AX integrates generative AI directly into the platform. For example, if you want to set up a workspace just by explaining it in plain English, the environment is prepared automatically before your task starts.

task.yaml
apiVersion: ax.io/v1alpha1
kind: Task
metadata:
  name: data-analysis
spec:
  workspaces:
    - name: python-env
      goal: "Set up a Python 3 development environment"

Generative workspaces

Describe what a ready environment looks like in plain English. AX hands that goal to an agent on first boot
to install toolchains and verify dependencies.

Run anything and everything

Interactive coding agents, long-running agent servers, Jupyter notebooks, headless browser testing, and custom tool runtimes—you name it.

Perfect for research

Spin up massive number of reproducible sandboxes to collect trajectories, run reinforcement
learning loops, and evaluate agents at scale.

For builders & researchers

Built to be the most friendly runtime for developers and researchers.

We want to make dealing with agentic infrastructure easier so you can focus on your work. AX
is
designed with an uncompromising focus on ergonomics, rapid iteration, and joyful workflows for both application
developers and AI researchers.

We aim to keep the runtime minimal and lightweight, while tastefully adding the essential
features everyone needs to build, evaluate, and scale agents.

About

Born from research, built for production.

AX was born at Google when agentic runtime systems research met frontier compute. Over years
of building and operating agentic execution engines, teams across Google recognized that agentic workloads
represent an entirely new computing paradigm: stateful, bursty, long-running actors that compute intensely for a
minute and then wait for model responses, tool responses, or human approval. Traditional orchestrators built for
stateless
microservices or predictable batch jobs become cost-prohibitive when keeping idle sandboxes running, yet lack
native support for sub-second suspend and resume.

Drawing on agentic runtime research from Google DeepMind alongside deep experience in
large-scale isolation, resumption, and scheduling, AX is being built as an open, declarative control plane
purpose-built for agent execution. It abstracts tasks, workspaces, network policies, and models into core
primitives so developers and researchers can run massive fleets of agents without reinventing the underlying
infrastructure. This project heavily relies on Agent
Substrate
but provides agentic abstractions and generative runtime components.


Source: Hacker News

A Necessary History of the Oddest Letter: W

“The letter W is a child of the fall of Rome. In the fifth century CE, the western half of the Roman Empire disintegrated into a patchwork of new kingdoms and new rulers. The reasons behind this collapse of imperial power are complex, but a large role was played by various peoples who had formerly lived outside its borders. The Romans might have looked down on these migrants as ‘barbarians’, but they also increasingly came to rely on them for military support. They were foederati—peoples bound by treaty to fight Rome’s enemies in return for land and food. It was only with the help of the foederati (a Latin word related to English federation) that the Romans were able to see off the threat of Attila the Hun in 451. Yet the more power these regional leaders had, the less authority the emperor and the central state could wield. This culminated in the overthrowing of the last emperor in the west in 476.

Language would have been a part of the divide between Roman and barbarian. By the fourth and fifth centuries, the western empire had become overwhelmingly Latin-speaking. By contrast, the newcomers spoke their own languages, perhaps with a passing knowledge of Latin too. From what we can tell, a great many of these migrants spoke Germanic languages. One of these incoming tongues was the ancestor of the language you are reading right now, English, which arrived in the remains of Roman Britain during this era. Germanic-speaking elites could now be found from southern Spain to the coasts of the North Sea. These new rulers were keen to sell themselves as legitimate successors to the emperors, and there was considerable continuity during this turbulent period.

By keeping up appearances and styling themselves as good Romans, they could dampen the jealousy of the old aristocracy and gain popular support. The new kings did not insist that scribes ought to write official documents in their own Germanic tongue, but eagerly adopted the more prestigious Latin language. This worked fine most of the time, but might occasionally hit a snag. Latin-writing lands were now ruled by men whose names contained un-Latin sounds. A new king might want his scribes to draw up a charter for some great display of generosity, but how were the scribes to spell that king’s name?

One of the troublesome sounds for writers was /w/. This is the common consonant in English water and want, and it would have been present in kingly Germanic names like Clovis, Vitiges and Odoacer. The trouble was, the Latin alphabet now had no letter for this sound.

Nonetheless, W established itself as a standard way to spell the Germanic sound /w/, including in Latin texts produced in England.

In ancient times, you would’ve heard the sound /w/ all around the Mediterranean Sea. Both Latin and Ancient Greek once used the sound, and both the Romans and the Greeks had letters to spell it. This was a sound that they, just like English, had inherited as part of their common Indo-European ancestry. Yet, as we saw in Chapters F and U, it was now foreign to them. In Greek, the consonant and its letter Ϝ had faded away, while in Latin, V had come to stand for the fricative /v/ instead.91 Time and time again, we find the ancient sound /w/ being lost or altered across the Indo-European family of languages. It would later happen in Continental Germanic languages too; in German today, W stands for /v/. The English consonant /w/ is actually a rare survivor, rescued from potential change by its migration to the island of Britain.

Out of the meeting of languages and writing in the new post-classical world, a letter was born to spell the alien /w/. From the sixth century onwards, likely starting in the powerful kingdom of Francia, innovative scribes doubled U. Within Latin texts, we find Germanically-named individuals like the abbot UUandeberctus and King UUaldemarus. The two letters were increasingly written as -one, and at least by the 11th century, they had fused into the letter W as we know it. Note that this was long before the split of V and U into two separate letters, hence some modern disagreement over their offspring’s name. In the English alphabet, it’s called double U. For the French, it’s double vé.

From its origins in Francia, W was exported to nearby lands that also needed it. W appears in early English texts, although not without competition. One alternative, seen in Cædmon’s Hymn in Chapter U, was a single U. Scribes would switch to one U when the following vowel was an /u/. This would avoid awkward-looking sequences of three Us in a row.

This dislike of ‘triple U’ in medieval texts is in fact still active in English spelling today. In the later Middle Ages, scribes would swap a U for an O if it came after W. This was done for the sake of clarity when reading. Even when words had a short /u/ vowel, spellings like wulf, wud and wunder would have been too confusing in the era of manuscript writing, what with its rows of upright quill strokes. This avoidance tactic can explain the modern mismatch between sounds and spelling in wolf, wood and wonder. Nonetheless, W established itself as a standard way to spell the Germanic sound /w/, including in Latin texts produced in England.

Nonetheless, W established itself as a standard way to spell the Germanic sound /w/, including in Latin texts produced in England.For example, the Life of Saint Æthelwold is a tenth-century biography that narrates the holy life of an English bishop. Being based in southern England, its Latin language is crammed full of English place-names containing the sound /w/, like Winchester, Worcester and Wallingford. The saint’s own name is spelled Aðeluuoldus.

Yet, during the same pre-Norman period, a specifically English written culture was also emerging alongside Latin. Its writers clearly had a sense that this was a separate language from Latin, and therefore could have its own spelling practices. While writing in Latin ought to use only Latin letters, they felt that they had more freedom when spelling Old English. Just as they had done with the letter Þ, English writers looked for an alternative to a lengthy W or an ambiguous single U. They reached into the world of runes, and employed Ƿ.

Instead, W has been brought in to tell the reader that the OW in town is a greatly shifted diphthong, no longer a single long vowel as it had once been.

Known as wynn, the letter Ƿ is extremely common in our Old English sources. See on p.301 how it appears twice in the first line in our only surviving copy of the poem Beowulf, in the words hƿæt ‘what’ and ƿe ‘we’.

It was standard spelling in the Wessex tradition, which would have written two and word as tƿa and ƿord. Examples of Ƿ outside the parchment pages of manuscripts show that the letter enjoyed popular use for centuries. A decorated dagger, found in Kent and dated to the ninth or tenth century, informs its viewers:

   Biorhtelm me ƿorte                   ‘Biorhtelm wrought me’

  S[i]gebereht me ah                    ‘S[i]gebereht owns me’

Yet, as you might have noticed, wynn is no longer a part of the English alphabet. It did survive the Norman Conquest, but gradually fizzled out during the Middle English era. It faced considerable opposition from the spelling of French and Latin, which had continued to use W since the sixth century. The pressure to match them meant that it was by W that Ƿ was eventually replaced.

Ever since the reapplication of W to English, the language has put the letter to a great many uses. Some instances of W are more recent in origin. Some even developed out of an original G.

In Old English, the letter G stood for one of a couple of similar sounds, depending on where in the word it came. In the middle of a word, a G represented a velar and fricative sound that was like a weaker /g/. During the Middle English period, this sound shifted into /w/, which also has a velar quality as a sound. This is how an Old English word like fugol ‘bird’ has become fowl, or how the sagu tool is now a saw. The Norse concept of lǫg, the facts of life laid down by fate or society, is behind English law.

These changes of G to W reflect a changed consonant, but elsewhere in spelling, W is used to tell us something about a vowel. It is especially common in words that have undergone the Great Vowel Shift, like town, cow and owl. These go back to tun, cu and ule in Old English, none of which had a G. Instead, W has been brought in to tell the reader that the OW in town is a greatly shifted diphthong, no longer a single long vowel as it had once been.

OW shares this role with OU. The second option for the same vowel appears instead in words like hour, shout and found. There has been a half-hearted rule in English spelling to use OW at the end of a word or syllable, and OU everywhere else. This rule would explain why we don’t write ‘nou’, ‘eyebrou’ and ‘allou’, but rather now, eyebrow and allow. At the end of a word like now, there is an audible /w/ sound, especially if the next word begins with a vowel (e.g. now I think …). However, this rule hasn’t been rigorously applied; we ought to write ‘broun’ and ‘croun’, not brown and crown.

Both OU and OW had good reasons to become the standard spelling for this post-shift vowel, but English failed to make a firm decision in favour of one or the other. It has even exploited the optionality to distinguish different words with a common origin. We spell flower with OW, while we use OU for the best quality or the ‘flower’ of ground grain—that is, flour.

Before we can leave W, there’s a mischievous effect of the letter to be acknowledged.

Consider three words: as, has and was. The third word, I think you will agree, does not rhyme with the previous two, despite their common spelling. Likewise, consider: and, hand and wand. The same lack of rhyme occurs, as it does in the trio arm, harm and warm. Notice the odd one out in ash, bash, cash, dash, gash and wash. If we also compare fan with swan, far with war, or fat with what, then their common denominator becomes clear: there’s something disruptive about the letter W.

To understand this effect, we have to concentrate on a particular quality that sounds in our languages can have. Vowels have featured often in this book, especially with regard to how far forwards, backwards, high or low our tongue is when we pronounce them. These features of tongue position are accompanied by the additional factor of lip rounding—whether or not we purse our lips at the same time.

The key thing to note here is that the consonant /w/ is pronounced with the lips and the back of the tongue. In the case of was, wand, wash and the rest, what has happened is that the /w/ rounded the following vowel, and also dragged it backwards in the mouth. The consonant has shared its rounded lips with the formerly unrounded vowel that comes immediately after it.

Consequently, in many varieties of English today, was, wand and ward have rounded vowels, while unrounded vowels can still be heard in their W-less counterparts, has, hand and hard. The cot-caught merger in North American English (see Chapter O) may be shifting and unrounding the particular vowel in the W-words, but nonetheless, hand still doesn’t rhyme with wand. The fact that this is an effect of adjacent sounds explains why the same changes and divergent vowels have also occurred in quality and quartz. They are not spelled with a W, but they still contain the influential consonant. Quartz doesn’t rhyme with parts, but rather shorts.

We still spell wash and warm as if they rhyme with ash and arm, because until fairly recently, they did. Their rounding is quite modern. It may have started sometime in the 15th century, but for the following four centuries, it remained limited to certain words and contexts. The first instances of W-rounding were likely in very common and unstressed words, like was. When said frequently and quickly, it’s more efficient to progress from a rounded-lipped consonant to a rounded vowel, than to switch off that rounding between the two. The effect was probably not present in the English of Chaucer, nor standard in the later speech of Shakespeare, on the basis of the words that these poets think are rhymes. In his sonnets, Shakespeare pairs was with glass, and warmed with disarmed.

Then were not summer’s distillation left,
A liquid prisoner pent in walls of glass,
Beauty’s effect with beauty were bereft,
Nor it, nor no remembrance what it was.

–Shakespeare, Sonnet 5

The fairest votary took up that fire
Which many legions of true hearts had warmed;
And so the general of hot desire
Was, sleeping, by a virgin hand disarmed.

–Shakespeare, Sonnet 154

Even Lord Byron, composing his narrative poem Childe Harold’s Pilgrimage in the early 19th century, rhymes three words that together sound awkward today.

I stood in Venice, on the Bridge of Sighs,
A palace and a prison on each hand:
I saw from out the wave her structures rise
As from the stroke of the enchanter’s wand:
A thousand years their cloudy wings expand …

–Lord Byron, Childe Harold’s Pilgrimage, Canto IV

Yet again, we have an instance of a reasonable change in sounds, and spellings that have not caught up. We could of course start to write wond instead of wand, or wor instead of war, or even woz for was. Maybe we will one day. For the moment at least, English readers and writers know to be cautious around the English letter W.” (297–306)

__________________________________

This article has been adapted from Why Q Needs U (Blink/Bonnier, U.S. June 2, 2026) by Danny Bate. It is provided courtesy of the publisher.


Source: Hacker News