In the Olympus blog you'll find the latest news about the community, tutorials, helpful resources and much more! React to the news with the emotion stickers and have fun!
Donald Trump will meet with the Chinese president Xi Jinping later today, ahead of a three-day visit that will undoubtedly see artificial intelligence (AI) high on the agenda.
António Guterres, the UN secretary general, has called on the two countries to establish a dialogue on AI similar to the communication between the US and the Soviet Union during the cold war.
“Governments with the greatest AI capabilities have the greatest responsibilities to humanity, to establish channels for dialogue, transparency, trust and cooperation,” Guterres said.
Tech executives including Jeff Bezos of Amazon, Sundar Pichai of Alphabet, Sam Altman of OpenAI, Tim Cook of Apple, Elon Musk of Tesla, Mark Zuckerberg of Meta and Jensen Huang of Nvidia will reportedly attend a state dinner at the White House on Thursday evening.
It comes as two leading progressive US lawmakers, senator Bernie Sanders and representative Greg Casar, were due to unveil legislation later today that would ban artificial superintelligence and create a federal agency to oversee advanced AI.
The bill, provided first to the Associated Press, would also pause advanced AI development until guidelines are implemented while creating the Department of Artificial Intelligence. Multiple employees at leading AI companies are endorsing the bill.
“It doesn’t take a genius to say, ‘slow it down,’” Sanders said in an interview with AP. “Do we really want to develop a super intelligence that when it becomes smarter than human beings could act independently of human control? I don’t think we do.”
Congress has so far done little to rein in the AI industry even as some of its most prominent leaders warn about potentially catastrophic risks.
Xi departed Beijing on Wednesday, state media said, heading to Joint Base Andrews, Maryland, where he will be met by Trump on the tarmac.
“President Xi Jinping left Beijing by special plane for a state visit to the United States” accompanied by his wife Peng Liyuan, his right-hand man Cai Qi, and foreign minister Wang Yi, among others, state broadcaster CCTV said.
In other developments:
Speaking at the UN general assembly, Donald Trump re-issued his threat to “annihilate” Iran and “drive them into Hell” if the regime didn’t capitulate and make a deal.” He also said he was “working closely with the leaders of Russia and Ukraine” to end the disastrous war, defended his administration’s actions in Latin America – including the midnight raid to arrest former Venezuelan president Nicólas Maduro, defended his drone strikes on boats in the Caribbean and the Pacific, accused Mexico of being effectively controlled by the cartels, and said it was “unacceptable” for the US to share a border with a nation under such conditions, pointed to regime change in Cuba and said he would reject any calls for a “globalist scheme” to control artificial intelligence.
Shortly after his UN address, Trump signed a Greenland deal, which Greenland’s prime minister Jens-Frederick Nielsen said “underlines the importance of the Nato alliance.” In tricky, carefully worded paragraphs, she said the agreement “recognises the US defining historical and ongoing contributions to the security and defence of Greenland,” but it also “respects the sovereignty and the territorial integrity of the Kingdom of Denmark, as well as Greenland’s right to self-determination.”
Trump told a CNN reporter at the UN general assembly on Tuesday that he was “surprised” the outlet was covering him, saying it ‘shouldn’t be here’ after he banned the network alongside Politico and MS Now from the White House on Friday.
At a joint press conference on Tuesday, vice-president JD Vance and Mehmet Oz, administrator of the Centers for Medicare and Medicaid Services, announced that the federal government was removing 760,000 enrollees from Affordable Care Act insurance marketplaces.
Sept. 23 (UPI) — The University of California, San Francisco School of Medicine discriminates against White and Asian applicants in favor of their Black and Hispanic counterparts, the Justice Department said amid the Trump administration’s efforts to eradicate diversity, equity and inclusion policies at institutions of higher learning.
The Justice Department has launched investigations into some of the nation’s leading medical and law schools to determine compliance with a 2023 U.S. Supreme Court decision that restricted race-conscious admissions practices.
In its findings released Tuesday concerning UCSF Medical School, federal lawyers alleged that prompting applicants to identify their race in the first two stages of the application process resulted in Black and Hispanic applicants being invited to the third and final interview stage at higher rates than White and Asian applicants with better MCAT scores and undergraduate grade point averages.
The report states that this was the case for the 2023 through 2025 incoming classes. During this period, UCSF Medical School was 4.6 times more likely to admit Hispanic applicants and 12.6 times more likely to admit Black applicants compared to White applicants with the same MCAT score, GPA and socioeconomic characteristics.
“Unfortunately, at UCSF Medical School, MCAT scores and undergraduate GPAs have taken a backseat to race,” Assistant Attorney General Harmeet Dhillon of the Justice Department’s Civil Rights Division said in a statement.
“Aspiring doctors should be admitted based on their qualifications. The Supreme Court has spoken clearly — federally funded medical schools may not admit students based on misguided and illegal notions of diversity.”
UPI has asked UCSF Medical School for comment.
Trump, who ran on removing left-leaning ideology from public and private spaces, has been targeting universities, many of which are some of the most celebrated in the nation, over alleged DEI hiring and admissions policies.
On Jan. 21, 2025, Trump’s second day back in office, the president signed an executive order directing the attorney general to issue guidance to federally funded educational institutions on how to comply with the Supreme Court’s 2023 ruling concerning race-conscious admissions practices at Harvard University and the University of North Carolina.
On March 30, the Justice Department informed UCSF Medical School of its investigation, requesting admissions data and other documents related to its admissions policies and practices.
The Justice Department has announced similar findings concerning other schools, including the University of California, San Diego School of Medicine, Yale University’s medical school, University of California, Berkeley School of Law and several others.
It is calling on UCSF Medical School to contact it by Oct. 2 to enter discussions on implementing a voluntary resolution agreement to ensure its compliance with the Supreme Court ruling.
Scenes from the 81st U.N. General Assembly
President Donald Trump speaks at the United Nations General Assembly at U.N. Headquarters in New York City on September 22, 2026. Photo by John Angelillo/UPI | License Photo
No friction … Marvel’s Wolverine. Photograph: Sony
No friction … Marvel’s Wolverine. Photograph: Sony
Why fans are digging their claws into Wolverine over its patronising gameplay
From highlighted ledges to characters telling you what to do, blockbuster games are increasingly taking the challenge out of playing – and players are pushing back
Marvel’s Wolverine, the latest action game from Sony’s stable of expensive blockbuster-factory studios, has had an unfortunate launch. Reviews were decent, if not stellar, but a series of viral video clips from reviewers and streamers have made it the scapegoat for a patronising flavour of game design that gamers have hated for decades. This clip, from Skill Up’s review, shows the player entering a beautiful wintry landscape before being overwhelmed by visual cues telling them what they’re supposed to be doing there. Sonar pulses point you in the direction of objectives. Targets – which, it bears emphasising, are shaped like targets – get an extra reflective visual effect so you know you’re supposed to pay attention to them. A trail shows you not just where to go, but where the level’s sole hidden collectible is.
It is hand-holding to the point of parody. Indeed, it reminded me of the sniggersome Mario in Unreal Engine satirical videos that makes the social media rounds every few months (“A mushroom? I should take a look at this”). Skill Up’s clip has been followed by hundreds of others, gleefully ripping into Wolverine with ironic savagery. In one much-copied clip, a streamer simply lays down his controller during a motorbike chase sequence and watches as Wolverine speeds on regardless, bouncing off the scenery as the game plays itself.
People really detest this kind of thing. It’s impossible to pinpoint exactly when big-budget games started assuming that players were irredeemably stupid, but throughout the late 00s and 2010s it became more and more common to see obvious GPS trails showing you where to go, and for characters to talk to themselves to remind you what to do. At the same time, there’s been a cyclical discussion over the difficulty of games: every time a notably challenging game comes out, we rehash the same talking points about whether all games should have easy modes. Is it acceptable for games to shut players out if they aren’t prepared to “git gud”?
I am not a game designer, but as a critic who loves challenging games and has thought a lot about them, for me there is a difference between a game’s difficulty and its friction. A game might be difficult for all kinds of reasons, intentional or otherwise: perhaps the controls or camera aren’t great; perhaps there’s a mid-game boss with a bafflingly massive health bar and one devastating unblockable attack; perhaps the balancing is off. But a game’s friction is always an intentional part of its design. Friction is about how much a game helps you achieve its tasks. Games with perfect friction are clear about showing you what to do, but not overbearing about helping you do it: there is intentional distance, distance that you must broach with agency, strategy and skill. Friction leaves space for frustration and learning, and a sense of achievement.
‘A game is only fun if you have to do something’ … 007 First Light. Photograph: IO Interactive A/S
A lot of modern blockbuster game design seems hellbent on removing this friction. Not sure where to go? Here – follow this brightly coloured trail of fart gas. Not sure how to climb that building? We’ve helpfully painted the ledges yellow. Have you stood still for 30 seconds? We’ll have another character talk in your ear and tell you exactly what you need to do. Playing through the opening hours of 007 First Light recently, I was dismayed when Bond inevitably acquired a watch that magically highlights important things in his vicinity with the press of a button.
At its best, game design like this helps players avoid becoming irrevocably stuck. At its worst, it totally removes agency and challenge – you know, the reasons that many of us play games in the first place. Fundamentally, a game is only fun if you have to do something. I don’t play games for passive entertainment. That’s what binge-watching Netflix is for. If you don’t have to invest anything to reap the rewards, the experience just washes over you.
Games can be difficult and have high friction. In Hollow Knight: Silksong, you are constantly figuring out what to do, where the next boss might be, how to open that tantalisingly blocked path. Achieving those goals is then also very challenging, requiring practice and mastery of whip-fast fighting and precise jumping. Or they can be high-friction, low-difficulty: Blue Prince tells you literally nothing about how to make your way to its mansion’s secret 46th room, but it’s not difficult to walk around and try out ideas. An example of a low-friction, high-difficulty game, meanwhile, would be a combat-focused one like Devil May Cry (or most shooters), where the aim is always “kill all the guys to open the next area” but actually killing all those guys is majorly demanding.
Low-friction, low-difficulty games hold very little interest. Many mobile games fall into this category, where you barely have to do anything save prod the screen to provoke a dopamine-shower of flashy visual effects (I’m looking at you, Monopoly Go). It’s unfair to lump Wolverine in with these almost entirely valueless games – but it is instructive that it has been subjected to such ridicule.
The infamously challenging Dark Souls, a dramatically influential game that turned 15 years old this week, seemed to remind the entire games industry that you canactually trust players: that they won’t drop the controller and storm off as soon as they get skewered by a skeleton. That there’s immense and unique pleasure in participatingin a game, unravelling it, conquering it little by little with skill and understanding. I hope that the reaction to Wolverine will serve as another reminder of that fact, for any executives who need to hear it.
Konami’s horror series is on a roll lately. The Silent Hill 2 remake was very well received (with a few exceptions, including our critic), last year’s feminist, small-town Japanese horror story Silent Hill: f was excellent, and now Silent Hill: Townfall has succeeded in sending the series to a new setting: an island off the Scottish coast. Per our reviewer Lewis Gordon: “This is a taut, smart tale, less about psychosexual terror than psychosocial dread… Townfall seems to understand that some things should remain shrouded in fog. The fate of St Amelia naturally festers in our imagination. How very Silent Hill.”
Available on: PS5, Xbox, PC
Estimated playtime: 12-15 hours
Pikachu in spaaaaaace … a toy Pikachu floats inside the International Space Station. Photograph: S Adenot/ESA/NASA/SWNS
Pikachu is now an astronaut, thanks to a collaboration between Pokémon and the European Space Agency as part of the franchise’s seemingly endless series of 30th birthday celebrations. Human astronaut Sophie Adenaut brought a Pikachu plush up to the International Space Station, where it spent 67 days in orbit.
According to Sonic Team boss Takashi Iizuka, Sega was once poised to kill off the whole Sonic franchise. It was the success of 2020’s live-action Sonic movie that saved it. “I had to fight within the organization to keep the franchise as long as possible,” he told One More Game.
The new Resident Evil movie is out, and it’s going down well. Aftermath describes it as like watching your inept stoner roommate play the games; certified horror-game enjoyer Keith Stuart says it’s one of the best video game adaptations he’s seen.
The Tokyo Game Show had to wrap up one day early last week due to an incoming typhoon. Gamesindustry.biz has a slightly depressing rundown of the show, which was heavily dominated by AI.
There have been yet more dramatic changes at Microsoft’s Xbox division. The company is all but closing Halo Studios and moving Halo’s development to Activision, which will also now be in charge of storied British studio Rare. Multi-BAFTA-winning studio Ninja Theory is starting consultations on a proposed closure, having failed to find a funding deal to return itself to independence.
Surreally European … Clair Obscur: Expedition 33. Photograph: Kepler Interactive
Reader Mike asks: “I just finished reading your article about the GameCube, which ironically got me thinking about the PS1 & 2, marvellous machines which I remember fondly. The game I got bundled with my PS2 was ‘Shadow of Memories’ which sticks with me to this day, the story was magnificent and also beautifully rendered, with multiple endings. Are there any similar ‘timeywimey’ games set in beautiful European villages, evoking memories of this game, out there today? Either on Android or PS5?”
This question unearthed a bunch of half-remembered fragments from the early days of the PS2 – I definitely played Shadow of Memories as well. It’s a Konami game about a young German travelling through time, trying to solve his own murder. You mentioned games available on PS5, so here are a few that have some of the same elements. The Life is Strangeseries is set in the US but has a similar time-travel and murder-solving premise; Clair Obscur is a very different kind of game but is as uniquely and surreally European; this year’s Japanese RPG The Adventures of Elliot involves travelling through different eras to prevent a catastrophe. None of these have the same strange adventure-game atmosphere as Shadow of Memories, though – that’s something that felt really specific to the PS2 era. Readers, have you any more suggestions for Mike?
If you have a question for Question Block – or anything else to say about the newsletter – email us at pushingbuttons@theguardian.com.
A UEFI DXE driver to enable Resizable BAR on systems which don’t support it officially. This provides performance benefits and is even required for Intel Arc GPUs to function optimally.
If using an NVIDIA Turing GPU (20 or 16 series) see NvStrapsReBar for enabling Resizable BAR on it.
Requirements
(optional) 4G Decoding enabled. See wiki page Enabling hidden 4G decoding if you can’t find an option for it. Without 4G Decoding you will be limited to 1GB BAR and in some cases 512MB you can try to increase this upto 2GB by reducing TOLUD
(optional) BIOS support for Large BARs. Patches exist to fix most issues relating to this
Usage
Follow the wiki guide Adding FFS module and continue through the steps. It covers adding the module and the additional modifications needed if required.
Once running the modified firmware make sure that 4G decoding is enabled and CSM is off.
Next run ReBarState which can be found in Releases (if you’re on Linux build with CMake) and set the Resizable BAR size. In most cases you should be able to use 32 (unlimited) without issues but you might need to use a smaller BAR size if 32 doesn’t work
If Resizable BAR works for you reply to List of working motherboards so I can add it to the list. Most firmware will accept unsigned/patched modules with Secure Boot on so you won’t have any problems running certain games.
The module is added to the UEFI firmware’s DXE volume so it gets executed on every boot. The ReBarDxe module replaces the function PreprocessController of PciHostBridgeResourceAllocationProtocol with a function that checks for Resizable BAR capability and then sets it to the size from the ReBarState NVRAM variable after running the original function.
The new PreprocessController function later gets called during PCI enumeration by the PciBus module which will detect the new BAR size and allocate it accordingly.
AliExpress X99 Tutorial by Miyconst
Instructions for applying UEFIPatch not included as it isn’t required for these X99 motherboards. You can follow them below.
UEFI Patching
Most UEFI firmwares have problems handling 64-bit BARs so several patches were created to fix these issues. You can use UEFIPatch to apply these patches located in the UEFIPatch folder. See wiki page Using UEFIPatch for more information on using UEFIPatch. Make sure to check that pad files aren’t changed and if they are use the workaround
Working patches
<4GB BAR size limit removal
<16GB BAR size limit removal
<64GB BAR size limit removal
Prevent 64-bit BARs from being downgraded to 32-bit
Increase MMIO space from 16-32GB to full usage of 512GB/39-bit range (Skylake/Kaby Lake/Coffee Lake)
Increase MMIO space from 8-16GB to full usage of 512GB/39-bit range (Haswell/Broadwell). The issue of older patches being limited to 64GB has been fixed.
Increase MMIO space from 16GB to full usage of 64GB/36-bit range (Sandy/Ivy Bridge). Requires DSDT modification on certain motherboards. See wiki page DSDT Patching for more information.
Remove NVRAM whitelist to solve ReBarState GetLastError: 5
Fix USB 3 ports not working in BIOS with 4G Decoding enabled (Ivy Bridge/Haswell/Broadwell)
X79 Above 4G Decoding fix
Build
Use the provided buildffs.py script after cloning inside an edk2 tree to build the DXE driver. ReBarState can be built on Windows or Linux using CMake. See wiki page Building for more information.
FAQ
Will it work on a PCIe Gen2 system ?
Previously it was thought that it won’t work on PCIe Gen2 systems but one user had it work with an i5 2500k.
Can I use Resizable BAR on my system without modifying BIOS ?
You can use Linux with 4G Decoding on, recent versions will automatically resize and allocate GPU BARs. If your BIOS doesn’t have the 4G decoding option (make sure to check hidden) or DSDT is faulty you can then follow the Arch wiki guide for DSDT modification using modifications from DSDT Patching and boot with pci=realloc in your kernel command line. Currently there is no known method to get it on Windows without BIOS modification
I set an unsupported BAR size and my system won’t boot
Clear CMOS and Resizable BAR should be disabled. In some cases it may be necessary to remove the CMOS battery for Resizable BAR to disable.
Will less than optimal BAR sizes still give a performance increase ?
On my system with an i5 3470 and Sapphire Nitro+ RX 580 8GB with Resizable BAR enabled in driver I get an upto 12% FPS increase with 2GB BAR size.
Over the years, Meta has invested in a number of storage service offerings that cater to different use cases and workload characteristics. Along the way, we’ve aimed to reduce and converge the systems in the storage space. At the same time, having a dedicated solution for critical package workload makes everyone happier. Having this in place is necessary for ourdisaster recovery and bootstrap strategy. This realization, coupled with a business need to provide storage for Meta’s build and distribution artifacts, led to the inception of a new object storage service — Delta.
Consider Delta’s positioning in the Meta infrastructure stack (below). It belongs at the very bottom, providing the basic primitive required for the availability and recoverability of the rest of the infrastructure. For bootstrap systems, complexity should be introduced only if it makes the solution more reliable. We’re only minimally concerned with the performance and efficiency of the solution. Another consideration for bootstrap systems involves the bootstrap itself. This process, by which engineers can access a small set of machines and restore the rest of our infrastructure, helps us get the product back up and working for people using it. Lastly, the bootstrap data needs to be backed up for recovery in case disaster strikes.
In this post, we will discuss the goals for Delta, the main concepts that govern Delta’s architecture, Delta’s production use cases, its evolution as a recovery provider, and future work items.
What is Delta?
Delta is a simple, reliable, scalable, low-dependency object storage system. It features only four high-level operations: put, get, delete, and list. Delta trades latency and storage efficiency in favor of simplicity and reliability. Since it’s horizontally scalable, Delta takes on minimal dependencies with appropriate failover strategies for soft dependencies in place.
Delta is not a:
General purpose storage system: Delta’s core tenets are resiliency and data protection. It’s designed specifically to be used by low-dependency systems.
Filesystem: Delta acts as a simple object storage system. It doesn’t intend to expose filesystem semantics like Posix etc.
System optimized for maximum storage efficiency: With resiliency as its primary tenet and focus on critical systems, Delta doesn’t intend to optimize for storage efficiency, latency, or throughput.
Delta’s architecture
Delta has productionizedchain replication, an approach to coordinating clusters of fail-stop storage servers. It intends to support large-scale storage services that exhibit high throughput and availability, without sacrificing strong consistency guarantees.
Before diving deeper into how Delta leverages chain replication to replicate client data, let’s first explore the basics of chain replication.
Chain replication
Fundamentally, chain replication organizes servers in a chain in a linear fashion. Much like a linked list, each chain involves a set of hosts that redundantly store replicas of objects.
Each chain contains a sequence of servers. We call the first server the head and the last one the tail. The figure below shows an example of a chain with four servers. Each write request gets directed to the head server. The update pipelines from the head server to the tail server through the chain. Once all the servers have persisted the update, the tail responds to the write request. Read requests are directed only to tail servers. What a client can read from the tail of the chain replicates across all servers belonging to the chain, guaranteeing strong consistency.
Chain replication vs. quorum replication
Now that we have provided an overview of what chain replication entails, let’s explore how chain replication fares against other widely used replication strategies.
Storage efficiency: Chain replication clearly does not offer the most storage-efficient replication strategy. We store redundant copies of the whole data set on all hosts in a chain. A comparatively efficient approach would involve intelligently replicating fragments of data using erasure coding techniques.
Fault tolerance: In an optimal bucket layout, chain replication can provide similar or better fault tolerance than quorum-based replication mechanisms. Why? A chain with `n` nodes can tolerate failures up to `n – 2` nodes without compromising on availability. On the contrary, for quorum-based replication systems, at least `w` hosts must be available to serve writes. Additionally, `r` hosts must be available to serve reads. Here `w` and `r` represent the write quorum size and read quorum size, respectively.
Performance: In replication strategies (like primary backup), all backup servers can serve reads. This increases read throughput. In the native idea of chain replication, only the chain tail can serve reads. (We optimized this bit and will share details in the apportioned queries section later in this post.) Much like quorum-based replication mechanisms, in chain replication all writes are directed to the primary (the head of the chain). But in chain replication, writes are only responded to after all hosts in the chain have acknowledged the update. Thus, chain replication has higher write latency on average in comparison with quorum-based replication mechanisms.
Quorum consensus: Quorum-based systems need complex consensus and leader election mechanisms to maintain quorum in the system. In contrast, the scope of quorum consensus in a chain-replication based system gets narrowed down to the much simpler, chain-host mapping. For example, the chain head always serves as a leader for processing writes, without the need for an explicit leader election.
Considering the above differences, chain replication clearly fails to offer the most storage-efficient way to replicate data across machines. Additionally, it yields higher average write latency in comparison to quorum-based systems, as we consider writes successful only when all links in a chain have persisted the update. However, it’s very simple while offering similar fault tolerance and consistency guarantees.
The anatomy of a Delta bucket
Now that we have a preliminary understanding of chain replication, let’s talk about how Delta leverages it to replicate data across multiple servers.
Each Delta bucket above includes several chains. Each chain usually consists of four or more servers, which can vary based on the desired replication factor. Each chain itself acts as a replica set and serves a slice of data and traffic. It can be thought of as a logical shard of a client data set. Servers in a particular chain get spread across different failure domains (power, network, etc.). Doing so guarantees durability and availability of client data if servers in one or more failure domains remain unavailable. We maintain a bucket config, the authoritative chain-host mapping for the layout of the bucket. When we add or remove servers and chains from the bucket, the bucket config gets appropriately updated.
When clients access an object within a Delta bucket, a consistent hash of the object name selects the appropriate chain. Writes are always directed to the head of the appropriate chain. It writes the data to the local storage and forwards the write to the next host in the chain. The write is acknowledged only after the last host in the chain has durably stored the data on local media. Reads are always directed to the tail of the appropriate chain. This guarantees that only fully replicated data is visible and readable, thereby guaranteeing strong consistency.
Delta supports horizontal scalability by adding new servers into the bucket and smartly rebalancing chains to the newly added servers without affecting the service’s availability and throughput. As an example, one tactic is to have servers with the most chains transfer some chains to new servers as a way to rebalance the load. New bucket layouts would still follow the desirable failure-domain distribution, etc., and are employed while rebalancing chains and expanding bucket capacity.
Failure and recovery modes in Delta
Failures can stem from hosts going down, networks being partitioned, planned maintenance activities, operational mishaps, or other unintended events. In a suitable implementation of chain replication, we assume servers to be fail-stop. In other words:
Each server halts in response to a failure rather than making erroneous state transitions.
A server’s halted state can be detected by the environment.
Consider a Delta bucket with `n` chains, with each chain comprising >1 host. In the event of any host misbehaving or getting partitioned from the network, the other sibling hosts (upstream/downstream) sharing a chain with the culprit host would be able to detect the suspicious host and report this behavior.
Sibling hosts can detect the erroneous behavior of misbehaving hosts by simple heartbeats or by encountering failures in transmitting acknowledgments/requests up/down the chain. If multiple hosts suspect a particular target host, the latter gets kicked out of all its chains and sent for repair. The bucket config gets updated appropriately.
Some decisions and trade-offs must occur when detecting unhealthy behavior in a host:
Timeout settings: We conducted several performance tests to arrive at the right timeout settings. We need to carefully assess the timeout between individual links in a chain before suspecting a host. We can’t set the timeout too short because servers can always experience transient network issues. We can’t set the timeout too long, either, because doing so would negatively affect operation latency. Not only that, but clients may also timeout while awaiting a valid response.
Suspicious host voting settings: We need to assess how many hosts should vote for a particular one being unhealthy before the suspected host is kicked out of its chains. The limit can’t be one since it’s always possible for two hosts in a chain to vote each other as unhealthy. This would cause both hosts to be disabled from their chains. The limit can’t be too large, either, as this would lead to the unhealthy host staying in the fleet for an extended period and negatively impacting service performance. Each host belongs to multiple chains and is guaranteed frequent connections with upstream and downstream nodes in their chains. As a result, setting the suspicion voting limit to two has worked well for us. Additionally, we have automated the faulty host repair flow. This provides us with the flexibility to configure sensitive thresholds and have more false positives.
Once the faulty host recovers, it can be added back to all the chains that it served prior to getting kicked out of the bucket. New hosts are always added to the rear end of the chains. Upon being added back to the chains, this host must synchronize itself with all the updates on the chain that occurred while it was not part of the chain. This process consists of scanning the objects on the upstream host and copying those not present or that have an obsolete version. Notably, during this interval of reconstruction, the host can still accept new writes from upstream. However, until it’s fully synchronized with the upstream host, it must defer reads to upstream.
We use this process for adding both suspected hosts as well as introducing new capacity to a Delta bucket.
How Delta has evolved over time
Apportioned queries
As evident from the above description, there are a few major inefficiencies with serving reads in chain replication.
The tail, the only node serving both reads and writes, can become a hotspot.
Considering that the tail serves all reads, the tail of the chain limits read throughput.
The basic idea? Each node serves read requests. But before responding to client read requests, it does a crucial check. It verifies whether the requested object’s local copy is clean or has been committed by all the servers in the chain. It may alternatively assess it as dirty, meaning the object does not replicate to all servers in the chain. The server can just return the last committed version of the object back to the client. This ensures that clients get only the object version that has been committed by all servers in the chain, thereby retaining chain replication’s strong consistency guarantees. The tail node would serve as the authority of the latest clean version for a particular object.
As explained above, each non-tail link in the chain makes an additional network call to the chain tail. It does so to fetch the clean version of an object before responding to the client reads. This additional network call to the tail link offers great value. Why? It helps scale the read throughput and chain bandwidth linearly with the chain length. Additionally, these object version check calls are cheaper in comparison with serving actual client reads. Hence, they don’t negatively affect client latency in a significant way.
Automated repair
While hardware failures or network partitions may seem rare, they occur fairly frequently in large clusters. As such, failure detection, host repair, and recovery should be automated to avoid frequent manual interventions.
We built out a control plane service (CPS) responsible for automating Delta’s fleet management. Each instance of the CPS gets configured to monitor a list of Delta buckets. Its primary function includes repairing chains that have missing links.
When repairing chains, the CPS applies a few techniques to achieve maximum efficiency:
When repairing a bucket, the CPS must maintain the failure domain distribution of all chains in the bucket layout. This ensures that the bucket hosts are spread evenly across all failure domains.
The CPS must ensure a uniform chain distribution on all servers. In this way, we avoid a few servers getting overloaded by hosting significantly more chains in contrast to other hosts.
When a chain has a missing link, the CPS would prioritize repairing the chain with the original host rather than a new one. Why is that? Resyncing all chain contents to a new host gets compute-intensive in comparison with syncing partial chain contents to the original host once it returns from repair.
Before enabling a host to a chain, the CPS would perform detailed sanity checks to ensure that a healthy host gets added back to the bucket.
Apart from managing the servers in the bucket, the CPS would also maintain a pool of healthy standby servers. In the event of a chain missing more than 50 percent of its hosts, the CPS would add a fresh standby host to the chain. This ensures that no severely underhosted chain jeopardizes the availability of the bucket. While doing this, the CPS would attempt to apply all the above principles in a best-effort manner.
Global replication
In Delta’s original implementation, when Delta clients wanted to store a blob in multiple regions due to data safety considerations, they made a request to each region. This is clearly not ideal from a user point of view. Delta’s users should not be in the business of tracking object location(s) while being able to control the level of redundancy and geo-distribution preferences. A client app should be able to just put an object to the store once and expect the underlying service fabric to propagate the change everywhere. The same goes for retrievals. In the steady state, the clients can expect to just get an object by submitting a request and letting the service retrieve it from the most optimal source available at the moment.
As our service evolved, we introduced global replication to client regions. This is done in a hybrid fashion — a combination of synchronously replicating blobs to a few regions and asynchronously replicating blobs to the remaining set of regions. Replicating to a few regions synchronously reduces the client latency while ensuring that other regions get the blob in an eventually consistent manner. Imagine that a particular region experiences a network partition or outage. The system would be intelligent enough to exclude that region from participating in geo-replication until it returns. Then that region could be asynchronously backfilled with the missing blobs.
How Delta handles disaster recovery
One of Delta’s key tenets is to be as low-dependency as possible. Along with having minimal dependencies, we also invested in having a reliable disaster recovery story.
As of now, we have integrated with archival services to continuously back up client blobs to cold storage. Additionally, we have developed the ability to continuously restore objects from these archival services. Out-of-box integration with archival services provides us with reliable recoverability guarantees in severely degraded environments. We have several partner teams that have integrated with our service because of our disaster recovery guarantees.
What’s next for Delta?
Looking ahead, we are working on a centralized backup and restore service for core infrastructure services within Meta based on Delta storage. Any reliable stateful service must be able to produce a snapshot of its internal state (backups) and rehydrate the internal state from a snapshot (restore). Our ultimate goal is to position ourselves as a gateway to all archival services and provide a centralized backup offering to our clients. Additionally, we also plan on investing heavily toward improving the reliability of Delta’s disaster preparedness and recovery story. This will help us serve as a reliable and robust recovery provider for our users.
Steve Schwartzberg’s journey from Harvard psychologist to erotic healer, Buddhist and teacher was always haunted by death. Then he was diagnosed with terminal cancer
This study was designed to measure the impact of bogus respondents on opt-in surveys. It also compares three methods for identifying and removing bogus respondents: 1) the use of trap questions, sometimes called “attention checks,” 2) the Sentry prescreening system proprietary to CloudResearch, and 3) matching respondents to a national voter file.
Why did we do this?
Pew Research Center does high-quality research to help the public, the media and decision-makers understand important topics. Our methodological research investigates the current challenges facing the polling industry, including past work on the impact of bogus respondents in opt-in surveys.
We fielded a large online opt-in survey Nov. 14-19, 2024, among 11,114 U.S. adult respondents. Respondents who agreed to provide their name and contact information were matched to a registered voter file after the survey concluded.
We evaluated how each approach for removing bogus cases performed on three data quality metrics: “yea-saying,” or agreeing regardless of what is asked, 2) quality of open-end text responses, and 3) response order effects. We also examined how screening methods substantively affected estimates of voter turnout and vote choice in the 2024 presidential election.
One of the most urgent problems in online opt-in polling is bogus (or fraudulent) respondents. These are survey-takers who make no effort to answer questions truthfully and instead are just looking to finish surveys quickly and collect rewards.
To combat this threat, researchers have developed various ways to identify and purge bogus cases from survey samples. Approaches include 1) trap questions, sometimes called “attention checks,” that genuine respondents should always answer correctly, 2) automated prescreening services, and 3) matching respondents to a registered voter file.
But how well do they work? A new Pew Research Center study finds that:
Overall, purging bogus cases lowers error on most metrics, but a surefire solution remains elusive.
Matching an opt-in sample to voter files slightly increased error by removing mostly good respondents (e.g., those who simply declined to give their name or address).
Trap questions and an automated prescreening service performed similarly, improving data quality somewhat.
All three approaches modestly increased the overestimation of Democratic support in the 2024 election. This appears to be due not to a systematic partisan bias but to bogus respondents’ tendency to say they voted for the winning candidate – in this case, Donald Trump, the Republican.
Matching opt-in samples to voter files may serve other purposes, such as providing data on respondents’ voting history. This study speaks only to whether this is an effective tool for purging bogus cases.
Do all polls have bogus respondents?
No. Bogus respondents are primarily a threat to online opt-in polls, which are recruited through methods like online advertising, self-enrollment and email lists. Polls recruited offline using random sampling (e.g., Pew Research Center’s American Trends Panel) are generally immune to this threat, though they face other challenges.
Studies conducted by Center researchers and others have found that standard data quality checks such as looking for “speeders” who complete surveys too quickly or “straightliners” who consistently select the same answer choice (e.g., always pick the first response option) fail to detect most bogus respondents. The same goes for many trap questions or attention checks, which not only fail to detect bogus respondents but can also confuse legitimate respondents, leading to false positives.
Faced with these challenges, pollsters have developed a variety of new approaches in hopes of better identifying bogus respondents.
In this study, we evaluate three such approaches:
Trap questions about unlikely behaviors
Automated prescreening by a leading fraud-detection company
Matching respondents to a commercial voter file
Trap questions
A difficulty in identifying bogus respondents lies in the fact that, for most questions, we have no way of knowing if a respondent has answered truthfully. To get around this problem, we asked respondents if they had ever engaged in four activities where we can be virtually certain that the true answer is “No.” Under this screening procedure, respondents were coded as bogus if they answered “Yes” to one or more of these questions.
The questions were designed so that diligent respondents would not be confused about how to answer, while bogus respondents trying to appear eligible for surveys targeting specific groups or answering randomly would be more likely to answer “Yes.” Specifically:
We asked respondents if they used any of six different social media platforms, one of which was a made-up platform called Fizzypress. We also asked if they had received any of six different government benefits in 2023, including payments from the United States Railroad Administration (USRA) – a real government agency, but one that ceased to exist in 1920. Both of these were asked early on in the survey and were intended to resemble screening questions looking for users of specific social media platforms or recipients of certain kinds of benefits. Of the full sample, 6% said they used Fizzypress and 9% claimed to have received USRA payments.
Later in the survey, we asked respondents if they had ever visited the International Space Station; 10% said they had. Finally, we asked if they had ever served on a Polar-class icebreaker ship, only two of which were ever built and only one of which remains in active use by the U.S. Coast Guard; 8% answered in the affirmative.
A total of 1,963 cases (18%) answered “Yes” to one or more of these trap questions and were flagged as bogus.
Automated prescreening
CloudResearch’s proprietary Sentry prescreening system is designed to automatically identify problematic respondents before they begin a survey. When a potential respondent first clicks the survey invitation link, they are routed to the Sentry platform and asked a short series of questions designed to elicit problematic survey-taking behaviors like yea-saying and inattentive responding. Open-end answers are checked to confirm that they are responsive to the question asked and not pasted from another source via an automated process.
The system also performs passive checks using metadata about the respondent’s device, location and browser for other signs that they are misrepresenting themselves through technical means.
Respondents who pass these checks are then routed to take the survey. Typically, those who fail are not forwarded to the main survey, but for this study, no respondents were terminated for failing the preescreening checks. We used metadata about which cases failed and why to simulate what would have happened to the survey results if they were excluded at the outset.
Under this procedure, 5,369 cases – nearly half of all completes – failed at least one prescreening check, including 1,879 (35% of failed cases) that failed multiple.
70% of failed cases had a problematic open-end, making this the most frequently failed check by far. This was followed by yea-saying (47%), inattentiveness (12%) and duplicate IP addresses (10%).
7% of these cases failed checks for fraudulent behavior, and only 1% failed other passive technical checks.
Voter file matching
The third approach we tested involves asking respondents to provide their name and address and then looking for a matching record in a commercial voter file, a national database of nearly all registered voters in the United States. If a matching record can be found, the case is considered valid. If no matching record is found, the case is thrown out.
This method assumes that respondents who can be matched are most likely being honest about their identity, and that someone willing to provide detailed contact information that can be validated against official voting records will likely be diligent about answering other survey questions.
This approach has primarily been used by pollsters who work for political campaigns. For cases that are successfully matched, voter files can provide a great deal of information beyond what was asked in the survey, such as a respondent’s registration status and voting history.
But while many voter file vendors attempt to include the unregistered population, Center research has found that a sizable share of this group is not covered by these databases. This can lead to throwing away data for many otherwise-valid respondents who simply aren’t registered to vote. This is not a large drawback for political pollsters, who typically focus on surveying registered voters. But if certain kinds of people who are more likely to be unregistered are underrepresented in a sample or missing altogether, it may lead to biased results if applied to a general population survey.
In this study, we asked respondents if they were willing to provide their name and home address so that they could be matched to the voter file. Out of all respondents, 49% agreed to provide their contact information; of these, 73% were successfully matched to the TargetSmart voter file based on name, address, age and sex. Altogether, 3,977 cases (36% of the full sample) were successfully matched while 7,137 (64%) were screened out under this procedure.
Distinct demographic patterns in flagged cases
Although the overall number of cases screened out by each method varies dramatically, respondents claiming to have certain demographic characteristics were consistently more likely to be flagged. Specifically, trap questions and prescreening were especially likely to flag those claiming to be ages 18 to 29 or Hispanic – the two groups found to be most affected by bogus responding in previous Center studies.
This should not be taken to mean that members of these demographic groups tend to be poor survey-takers. Rather, insincere respondents tend to claim membership in these groups, likely falsely in many cases.
The profile of cases purged using voter file matching was different. This reflects the fact that voter file matching tends to purge cases for benign reasons (e.g., a respondent’s privacy concerns or their not being a registered voter), not harmful ones. Hispanic cases were no more likely to be screened out than non-Hispanic Black cases, and both were only 6 percentage points more likely to be screened out than non-Hispanic White cases. In contrast, Hispanic cases were more likely to be flagged than non-Hispanic White cases by trap questions (+22 points) and prescreening (+19).
Likewise, trap questions and prescreening were both more likely to flag men than women, by margins of 9 and 6 points, respectively. Men and women were screened out by voter file matching at roughly the same rate.
The three methods differed on education:
With trap questions, college graduates (+11) and respondents with a high school education or less (+7) were more likely to be flagged than those with some college education.
With prescreening, high school or less cases (+13) were more likely to be flagged than college graduates, who had the next-highest flag rate.
Voter file matching showed much less differentiation: Rates across education groups did not vary by more than 4 points.
Effect of screening on measures of data quality
What impact do these three screening methods have on data quality? There is no direct way to know what proportion of bogus respondents are successfully identified by each method, nor can we be sure how many valid cases are misidentified as fraudulent. Instead, we can compare various measures of data quality calculated before and after each screening method has been applied and the flagged cases removed. To ensure that results are comparable across screening methods, separate sets of survey weights were created for the unscreened sample and for the set of cases remaining after each method was applied. For details, refer to the methodology.
In this study, we focused on three measures of data quality: 1) the frequency answering “Yes” to yes/no questions, sometimes called “yea-saying,” 2) the quality of text responses to open-ended questions, and 3) the severity of response order effects when the order of answer choices is randomized.
Yea-saying
In a 2023 Pew Research Center study, one indicator of bogus responding was a tendency to answer yes/no questions in the affirmative, regardless of the question. In that study, 8% of online opt-in respondents answered “Yes” to at least 10 of 16 yes/no questions asked, with especially high shares among 18- to 29-year-olds (15%) and Hispanic respondents (19%). In contrast, the corresponding shares among respondents from probability-based panels fell between 1% and 2%. Many of these questions measured rare or uncommon characteristics, and we would expect that virtually no one answering truthfully would say “Yes” to 10 or more.
Our latest survey exhibits largely the same pattern. Prior to any screening, 7% of all adults, 10% of 18- to 29-year-olds and 13% of Hispanic respondents answered “Yes” to at least 10 of 15 yes/no questions (excluding a question measuring Hispanic ethnicity and questions used in any screening procedures).
After screening with trap questions, this behavior is largely eliminated from the remaining respondents: 1% of adults and no more than 2% of any demographic subgroup answered 10 or more yes/no questions in the affirmative. The reduction is almost as large for prescreening, which reduced the shares to 2% for all adults, 4% for young adults and 3% for Hispanic adults.
In contrast, matching to the voter file appears to have made the problem worse. The share among all adults increased slightly to 9%, and the shares among young and Hispanic adults rose to 15% and 17%, respectively.
This surprising result appears to be due to two factors. The first is that respondents who declined to provide their contact information – nearly half of the sample – appear much less likely to be bogus than those who agreed to provide this information. Among those who declined, only 3% answered “Yes” to 10 or more yes/no questions, compared with 12% among those who agreed.
The second is that a nontrivial share of bogus respondents were able to provide contact information that was accurate and detailed enough to result in a successful voter file match, though the information provided may not be their own. To the extent that some bogus respondents were successfully removed, it was not enough to offset the loss of valid respondents who declined to share contact information.
Open-end response quality
Examining text answers to open-ended questions has long been considered a reliable way to identify low-quality respondents. While attentive respondents give answers that are relevant to the question asked, bogus respondents often give nonsensical or gibberish answers.
This survey included an open-ended question that asked, “What is one thing that you would like politicians in Washington, D.C. to know about your own situation when they are writing laws and setting policy? Please share as much detail as you can.”
Each respondent’s answer was reviewed and assigned to one the following categories: 1) relevant, 2) nonresponse, 3) generic positive rating, 4) other non sequitur, 5) gibberish, or 6) probable AI (refer to the methodology for details). Because respondents had been told they could skip any question they did not wish to answer, both relevant and nonresponse answers were considered unproblematic, while the remaining categories were deemed problematic.
Prior to any screening, 88% of answers were classified as unproblematic, including 69% that were relevant and 19% that were nonresponse. The most common kind of problematic answers were other non sequiturs (7%), followed by generic positive ratings (2%), gibberish (2%) and probable AI (1%). It is worth noting that the probable AI responses detected were written in a distinct style that was relatively easy to spot. There may be other, more subtle, AI responses that were not detected.
Among all adults, trap questions and prescreening performed similarly, bringing the overall share of problematic open-ends down to 7% and 5%, respectively. Prescreening also resulted in a higher share of relevant answers than trap questions (80% vs. 74%) and a lower share of nonresponse (15% vs. 19%). This may be because the prescreening process itself includes an open-ended question. Matching to voter files, on the other hand, had virtually no effect on the proportion of problematic answers, though the proportion of relevant answers increased to 73%.
Response order effects
If a respondent is being attentive and answering questions accurately, the answer options they select shouldn’t be affected by the order in which they are presented. But inattentive respondents in online surveys tend to select answer choices that appear toward the top of the list, otherwise known as a “primacy effect.”
Our survey included a question asking if undocumented immigrants should be allowed to remain in the country, after which a random half of respondents were shown answer choices in the following order:
“They should not be allowed to stay in the country legally.”
“There should be a way for them to stay in the country legally, if certain requirements are met.”
For the other half of respondents, this order was reversed. If all respondents were answering diligently, the results from each half of the sample would be largely the same.
Instead, we saw a substantial primacy effect on this question. Without any screening, 43% of adults endorse the view that undocumented immigrants should not be allowed to stay legally when that option is presented first. The share drops to 34% when the order is reversed, a primacy effect of +9 percentage points.
The primacy effects were even larger for those subgroups most affected by bogus responding, at +13 points for men, +15 points for 18- to 29-year-olds and +17 points for Hispanic adults.
Screening with trap questions reduced the primacy effect by roughly half, to +4 points for all adults, +6 for men, +7 for young adults and +9 for Hispanic adults.
Prescreening performed similarly, reducing the primacy effect to between +5 and +10 points.
Voter file matching, on the other hand, resulted in even larger primacy effects than when there was no screening at all, with magnitudes increasing to +14 for all adults, +18 for men, +29 for young adults and +25 for Hispanic adults.
As with the frequency of giving “Yes” answers, voter file matching’s negative impact on data quality appears to be due to valid respondents choosing not to provide contact information for matching, with those cases showing a primacy effect of only +2 points.
Including the question about the status of undocumented immigrants, there were a total of 13 questions with randomized response options that were asked of all respondents. Nearly all of these showed the same pattern. Without any screening, the average primacy effect on these questions was +7 percentage points for all adults, and as high as +11 and +12 points for young adults and Hispanic adults, respectively. Trap questions reduced the average primacy effect by 4 points for all adults and by 7 points for young and Hispanic adults, while prescreening did almost as well.
How does screening affect questions about voter turnout and vote choice?
The prevalence of online opt-in samples in election polling raises the question of how different screening methods – voter file matching in particular – impact estimates related to voter turnout and vote choice. Fielded shortly after the 2024 U.S. presidential election, this survey asked respondents if they voted, and if so, for whom.
Without any screening, an estimated 84% of self-reported registered voters said they voted in the 2024 presidential election. Purging cases with trap questions and prescreening both increased that share by 3 points, to 87%, while voter file matching increased it by 1 point, to 85%. This pattern is more pronounced among young adults and Hispanic adults: For these groups, trap questions and prescreening increased estimated turnout by between 6 and 9 percentage points, while voter file matching produced a 1-point increase among young adults and a 2-point decrease among Hispanic adults.
Although an increase in voter turnout would appear to make these estimates less accurate relative to a higher quality estimate of roughly 77%, it is notable that, for this question, the order of response options was not randomized and the response option for having voted was presented last.1 This kind of pattern is what we would expect to see if bogus respondents, who are more likely to select answer choices toward the top of the list, are being screened out. The fact that the effect is smaller for voter file matching is also consistent with that method’s tendency to disproportionately exclude valid respondents.
All three screening methods increased estimated support for Harris
When it comes to presidential vote (for which response option order was randomized), the weighted raw sample showed Donald Trump and Kamala Harris tied, each with 48% of the vote – accurate to Harris’ true vote share, but underestimating Trump’s by 2 points. Trap questions slightly shifted the result from a tie to a 3-point advantage for Harris. The effect was larger for prescreening and voter file matching, which shifted Harris’ margin to +6 and +7 points, respectively.
These patterns suggest that Harris supporters are overrepresented among valid respondents. But prior to screening, this bias was largely offset by the presence of bogus respondents, who disproportionately said they supported Trump. When bogus respondents were screened out, the bias in favor of Harris became more apparent.
Because the answer choices for Trump and Harris were randomized, this result also implies that bogus respondents were not simply choosing the first option or choosing randomly. There appears to be a subset of bogus respondents, specifically those with the most glaring data quality problems, who were much more likely to say that they voted for Trump, regardless of whether that option was presented first or second. For example, in the unscreened sample, cases coded as having a problematic open-end supported Trump over Harris by a margin of 29 percentage points, while cases that gave 10 or more “Yes” answers favored Trump by 30 points.
Cases with one or both of these data quality problems made up anywhere from 11% of screenouts from voter file matching up to 57% of screenouts from trap questions.
This should not be taken to mean that bogus respondents will always be systematically biased in favor of Republican candidates. In a previous Center benchmarking study where respondents were asked about the 2020 presidential election, opt-in respondents who gave 10 or more “Yes” answers claimed to have voted for Joe Biden, the Democrat, over Trump by margins ranging from 34 to 51 points across three different opt-in samples. It seems plausible that, when asked about past elections, many of these respondents are simply choosing the candidate who won.
These effects are even larger for subgroups in which bogus respondents are most common. Among 18- to 29-year-old voters, the survey initially showed Trump ahead by 2 points prior to screening. All three screening approaches shifted the margin in favor of Harris, putting her ahead by 3 points with trap questions, 6 points with prescreening and 1 point with voter file matching. Among Hispanic voters, trap questions slightly decreased Harris’ lead from 9 points to 8. Prescreening and voter file matching had larger effects, increasing Harris’ lead among Hispanic voters to 16 and 19 points, respectively.
CORRECTION (Aug. 28, 2026): The following sentence was updated to reflect the correct percentages for all U.S. adults, young adults (ages 18 to 29) and Hispanic adults: “The share among all adults increased slightly to 9%, and the shares among young and Hispanic adults rose to 15% and 17%, respectively.” Figures were correct in the corresponding graphic; this change does not affect other findings reported.
Intel ships programmable SHAVE cores inside its NPUs, but the public stack
exposes only graph-level programming. npunlock reconstructs the missing path
from custom C code to a runnable NPU kernel.
The current implementation has been verified on Windows x64 with Meteor Lake /
NPU3720.
Latest breakthrough — 2026-09-23: One native graph can execute independent FP32-unary and FP16-binary custom branches; explicit ACT-group preflight handles the compiler’s branch reordering. Evidence and limits.
Quick example
This complete FP32 GELU example embeds the C kernel in Python, places it in an
NPU graph, and checks the result against NumPy. The bundled npunlock/npu3720_kernel.h target header supplies the NPU3720 invocation and
tensor-address helpers. The tested MoviTools toolchain makes most conventional libm functions available to kernels without including <math.h>; this
example calls tanhf directly. See the mlibm.a symbol inventory for the observed candidates.
importnumpyasnpimportnpunlockasnpunpu.configure(movi_dll_dir=r"C:pathtoMVC_DEPEND")
gelu_c: bytes=b"""#define MLIBM_DEFINE_LINK_COMPAT 1#include <npunlock/npu3720_kernel.h>void controlled_act(unsigned layerParams) { act_abi_invocation invocation; ACT_ABI_LOAD_INVOCATION32_OR_RETURN(layerParams, invocation); const float *in = ACT_ABI_INPUT_PTR32(const float, invocation, 0u); float *out = ACT_ABI_OUTPUT_PTR32(float, invocation, 1u); const float SQRT_2_DIV_PI = 0.7978845608028654f; for (unsigned i = 0; i < invocation.element_count; ++i) { float x = in[i]; float w = x + 0.044715f * x * x * x; w = tanhf(w * SQRT_2_DIV_PI); out[i] = 0.5f * x * (1.0f + w); }}"""N=2048x=npu.input("x", shape=(1, N), dtype="f32")
y=npu.custom(
x,
source=gelu_c,
carrier="Abs",
_name="y",
)
program=npu.compile(npu.Graph(inputs=[x], outputs=[y], name="gelu_f32_example"))
input_value=np.linspace(-4, 4, N, dtype=np.float32).reshape(1, -1)
output=program.run({"x": input_value})["y"]
reference=0.5*input_value* (
1.0+np.tanh(
np.sqrt(2.0/np.pi)
* (input_value+0.044715*input_value**3)
)
)
print(f"maximum absolute error: {np.max(np.abs(output-reference)):g}")
Intel’s normal NPU software accepts graphs made from operations its compiler
supports; it does not expose a public workflow for supplying a C implementation
for an operation. The NPU’s ACT-SHAVE processors are programmable and run
software kernels. npunlock makes those processors usable for compatible
custom graph operations while retaining Intel’s compiler and driver for the
surrounding graph and hardware execution.
Requirements
Windows x64
Meteor Lake / Intel NPU3720
an installed Intel NPU driver for the device
Python 3.10 or newer
CMake 3.24 or newer and an installed MSVC toolchain for source installation
the extracted MoviTools MVC_DEPEND toolchain for custom C compilation
OpenVINO is not required as a runtime, Python package, or compiler frontend. npunlock does emit OpenVINO-format IR for the installed Intel driver.
Install
npunlock is currently installed from a source checkout:
python -m pip install .
The build bundles npunlock.dll and npunlock_worker.exe inside the Python
package, so normal Python use does not require a separate native path.
Get MoviTools
Custom C compilation uses Intel/Movidius MoviTools, which npunlock does not
redistribute or download.
A MoviTools package verified to work was found in a legacy Lenovo driver pack. See Getting MoviTools for the official download,
hash, extraction command, and expected layout.
Extract the MVC_DEPEND payload from Lenovo’s older
Intel NPU driver package 31.0.100.1688, but remember, DO NOT install or downgrade to
that driver. All we need is the bundled MoviTools.
Run an example
Point npunlock at the extracted MVC_DEPEND root and run GELU:
The example runs on the NPU and reports its maximum error against a NumPy
reference.
What currently works
compile user-written C into ACT-SHAVE machine code
run custom kernels inside Intel NPU graphs
static dense FP16 unary and two-input custom kernels
a verified unary FP32 path
one graph containing independent FP32-unary and FP16-binary custom branches
nonlinear math such as GELU and tanhf
reusable NumPy-compatible host/NPU shared input and output buffers
Python, CLI, and native C APIs
Current limitations
Support is experimental and currently limited to Windows x64, Meteor Lake /
NPU3720, static shapes, compatible ACT carriers, and known tensor layouts.
Connected mixed-precision conversion groups are not yet patch-discoverable;
the verified mixed-precision example uses independent branches. Other NPU
generations have not been verified. See Current limitations for the full compatibility boundary.
Help test Linux and newer NPUs
Have an NPU3720 Linux system or a newer Intel NPU? Contributions are welcome.
Two routes look especially promising but remain untested:
a patched NPU3720 graph produced on Windows may run on Linux because the NPU
firmware executes the custom machine code; building SHAVE code on Linux would
additionally require a way to load the Windows MoviTools DLLs;
newer NPUs may execute the existing 3720xx SHAVE image, or an older OEM
driver package for that generation may provide matching MoviTools components.
Both need hardware validation, driver/firmware version records, and output
comparison against a host oracle. If you can test either path, feedback, failure
reports, and code contributions are welcome. See Porting to Linux and newer NPUs for the hypotheses, caveats,
and a suggested test plan.
Documentation
Getting MoviTools — obtain the compiler toolchain without installing the legacy driver
Python API — construct, compile, and execute graphs from Python
A note on AI use: I did use AI while building this project–for scaffolding, repetitive implementation work, converting my reverse-engineered results into organized documentation, and fixing my English. The reverse engineering, experiments, debugging, and technical conclusions came from hands-on work. If that doesn’t bother you, there’s a pretty deep and surprisingly satisfying rabbit hole ahead.
License
npunlock is licensed under the Apache License 2.0. MoviTools and
the Intel/Movidius libraries are external proprietary dependencies and are not
covered or redistributed by this repository.
Traditional VPNs see your identity and your browsing history
Traditional VPN
Obscura only sees your IP address and never your browsing history
Obscura
Exit Node
What makes Obscura different from existing VPNs?
Unlike VPNs with a “no-logs” policy, Obscura is provably private by design.
Even “no-logs” VPNs see both your identity and your internet activity, meaning you have to blindly trust their pinky-promise for privacy. This is exactly why some privacy-conscious folks will tell you not to use a VPN at all.
Obscura is different – we never see your decrypted internet packets. It’s simply impossible for us to log your internet activity, even if we were compelled to, or if our servers were compromised. You can even verify this yourself.
Obscura’s stealth protocol is much harder to block.
Our unique stealth protocol is designed to blend in with regular internet traffic. It does so by leveraging QUIC – the same technology that powers HTTP/3 – making it far harder for censors or network filters to detect or block.
Obscura only sees your connecting IP address and necessary payment info – never your actual internet traffic.
We physically can’t decrypt your internet traffic (see how) and never log your connecting IP address.
For even more privacy, we actively support privacy-friendly payment methods like Bitcoin over Lightning and Monero. Learn more about accepted payment methods here.
How do I know Obscura does what it claims?
I like the way you think! 😎
Don’t trust — verify. Our app’s entire source code is on GitHub for you to verify that we do what we say.
We also plan to provide reproducible builds of our app, meaning anyone can confirm that the app you download matches the code we publish.
Additionally, our app displays your current exit hop’s WireGuard public key on its “Location” page. You can check this key against what Mullvad publishes here to ensure that you’re connected via a genuine Mullvad exit hop!
How much does Obscura cost, what payment methods are accepted?
Obscura is just $8/month.
You can top-up your account using:
Credit Card (via Stripe)
Bitcoin over Lightning (for more privacy)
Monero (for more privacy)
For convenient renewals, you can also subscribe with your Credit Card (via Stripe). Note that Stripe may need your email for subscriptions, but Obscura never stores it.
How many devices can I use Obscura on?
Each Obscura account has 5 simultaneous connection slots. This is different from ”5 devices”.
If you’re connecting with our Obscura app…
You can log in on as many devices as you’d like.
A device only uses 1 slot while it is actively connected (see the “Connection” tab in the app).
If you’re connecting with our WireGuard configs…
Each config reserves 1 slot as long as it exists on your account, even when the device is not connected.
To free a slot, delete the config from your account here.
Examples
Only using our Obscura app: Unlimited sign-ins; up to 5 devices connected at the same time.
Only using WireGuard configs: You can keep up to 5 configs total (i.e., 5 devices).
A Mix: If you have 1 WireGuard config on your account, it uses 1 of 5 slots, leaving 4 slots open for the Obscura app to connect from any signed-in devices.
How does Obscura differ from my VPN’s multihop option?
With a typical multihop VPN, the same provider controls every hop—so they can still link your identity to your traffic.
Obscura solves this by using a fully-independent exit hop (currently Mullvad). Ensuring that our servers never see your actual traffic, and the exit hop never sees your identity.
How does Obscura compare to Tor?
We have immense respect for the Tor project (and encourage you to support it), but its volunteer-run network can be slow and susceptible to DDoS issues, making it infeasible for everyday use.
Obscura uses two dedicated, high-performance hops for maximum speed and reliability – meaning you get many of Tor’s privacy benefits without sacrificing everyday usability.
What server locations are available?
We’re always working to add more exit locations. We currently have the following:
North America
Toronto, Canada 🇨🇦
Vancouver, Canada 🇨🇦
United States 🇺🇸
Ashburn, VA
Chicago, IL
Dallas, TX
Denver, CO
Los Angeles, CA
Miami, FL
New York, NY
San Jose, CA
Seattle, WA
Latin America
São Paulo, Brazil 🇧🇷
Santiago, Chile 🇨🇱
Bogotá, Colombia 🇨🇴
Querétaro, Mexico 🇲🇽
Europe
Prague, Czech Republic 🇨🇿
Paris, France 🇫🇷
Frankfurt, Germany 🇩🇪
Milan, Italy 🇮🇹
Palermo, Italy 🇮🇹
Amsterdam, Netherlands 🇳🇱
Warsaw, Poland 🇵🇱
Madrid, Spain 🇪🇸
Stockholm, Sweden 🇸🇪
Zurich, Switzerland 🇨🇭
Kyiv, Ukraine 🇺🇦
London, United Kingdom 🇬🇧
Asia
Hong Kong 🇭🇰
Tokyo, Japan 🇯🇵
Singapore 🇸🇬
Istanbul, Turkey 🇹🇷
Africa
Johannesburg, South Africa 🇿🇦
Oceania
Auckland, New Zealand 🇳🇿
Sydney, Australia 🇦🇺
What relay servers does Obscura offer?
We will continue to add more relay (first-hop) servers to Obscura. Currently we have the following:
North America
United States 🇺🇸
Atlanta, GA
Chicago, IL
Dallas, TX
Denver, CO
Los Angeles, CA
Miami, FL
New York, NY
San Jose, CA
Seattle, WA
Washington D.C.
Latin America
Querétaro, Mexico 🇲🇽
Santiago, Chile 🇨🇱
São Paulo, Brazil 🇧🇷
Europe
Amsterdam, Netherlands 🇳🇱
Frankfurt, Germany 🇩🇪
London, United Kingdom 🇬🇧
Madrid, Spain 🇪🇸
Stockholm, Sweden 🇸🇪
Warsaw, Poland 🇵🇱
Asia
Hong Kong 🇭🇰
Osaka, Japan 🇯🇵
Tokyo, Japan 🇯🇵
Singapore 🇸🇬
Africa
Johannesburg, South Africa 🇿🇦
Oceania
Melbourne, Australia 🇦🇺
Perth, Australia 🇦🇺
Sydney, Australia 🇦🇺
How does the macOS app work? Does it have kernel-level access?
Our macOS app installs a Network Extension, which is a fully
sandboxed process with no kernel-level access to your system.
You can verify this by looking at our source code.
Give me the technical details!
Obscura’s servers relay the WireGuard tunnel between your device and Mullvad’s servers. Essentially, WireGuard-over-QUIC.
This looks something like:
<user> --> Obscura Server --> Mullvad Server --> <internet>
This ensures that no single party has the information to leak or correlate your traffic and your identity, since:
Obscura’s servers are unable to decrypt the WireGuard packets it relays, since they are encrypted to a Mullvad server’s WireGuard pubkey.
Mullvad’s servers never see your connecting IP since Obscura’s servers are effectively doing NAT.
Your device connects to Obscura’s servers over QUIC, meaning:
Obscura connections are harder to detect or block since they look like regular internet traffic (HTTP/3 uses QUIC for transport).
Note that since this uses the original WireGuard protocol, you won’t benefit from Obscura’s QUIC-based obfuscation for outsmarting internet censorship. However, your connection still benefits from improved privacy with our Two-Party Relay architecture.
FoxDev Studio opens the projects, forms and tables you already have and runs them the way you remember: no rewrite, no conversion, no export step.
Point it at a folder you have not opened in years, and what is there is what you get.
Projects, forms, class libraries, menus and reports open straight from the files you already have. Nothing is migrated first, and nothing is left behind.
Forms, classes, menus and reports are edited where you expect them to be, alongside the project manager, the Command Window and a debugger that stops on your line.
Tables, indexes, memos and databases are read and written in place. What sits on disk afterwards is the same kind of file it was before.
The system calls, the automation objects and the old add-in libraries your application leans on keep working, so the parts nobody wants to touch stay untouched.
The IDE, on Visual FoxPro's own sample projects. Every picture is a real session.
Small differences are what break an old application: a number printed a column too wide, an event arriving a moment late, an error with the wrong number on it. So behaviour here is settled by asking Visual FoxPro itself and matching its answer, rather than by reading a reference page and hoping.
What comes out of that is a runtime written from scratch: quick to start, self-contained, and straightforward about the corners it has not reached yet.
Visual FoxPro stopped at version 9, and at 32 bits. This is the same language, rebuilt on a foundation that has not been frozen since 2007. Four pieces are what make that possible.
Visual FoxPro is a 32-bit program, and that decides more than it appears to. It is why a table stops at two gigabytes, why a memo file stops at two gigabytes, and why a big report runs out of memory on a machine with plenty to spare. The limits are signed 32-bit numbers buried in the file handling, not a licensing decision anybody made.
FoxDev Studio is 64-bit throughout. Every file offset is 64-bit and a table is never read into memory at all, so the same .dbf that used to stop dead carries on into the hundreds of gigabytes. One thing to know before you lean on it: a table grown past two gigabytes will not open in Visual FoxPro again. If you still work in both, that is a one-way door.
Visual FoxPro compiled your program to p-code and shipped a runtime to execute it. This is the same arrangement, made again: a compiler and a bytecode interpreter written in Rust and compiled to WebAssembly, so one machine runs your code wherever the application runs. The editor checks what you type through that very compiler, so what it underlines and what the runtime refuses cannot drift apart.
A running program is a fiber. When it needs something from the world outside (a message box, a modal form, the next record) it does not call out and block; it yields, the work is done while the machine is off the stack, and the answer is handed back. That is why MESSAGEBOX() stops your program without freezing the window behind it, why READ EVENTS waits without spinning, and whySetFocus can fire GotFocus, and Init can run while a form is still being built, in the order FoxPro always did it.
How the machine works and what it runs.
A running form is a live tree of objects with the properties you would expect, and the interface is drawn straight from that tree by React. Each object watches only itself, so THISFORM.lblGreeting.Caption = cMsg repaints one label rather than the whole form. On a dense screen that is the difference between instant and sluggish.
The same tree is what the designer edits, one step earlier. There is no second model kept in step with the first, which is the usual place a form and its designer start telling different stories.
An .fll is a 32-bit image, and every process in a 64-bit application is 64-bit, so nothing inside the application itself could ever open one. Rather than tell you it is impossible, SET LIBRARY TO starts a small 32-bit process whose only job is to hold your library, and the runtime talks to it. The calls are synchronous, because a program may call into a library halfway through an expression, and an answer that arrived later would not be an answer. It is measured against real libraries: the encryption library, FoxTools, and libraries built from Microsoft's own API samples.
Nothing 64-bit needs the bridge: DECLARE … DLL reaches a modern library in the same process, and automation objects are reached the way they always were. The old road stays open; it just is not the only one any more.
Everything you have written still means what it always meant. FoxScript only adds on top: a block you can hand to something else to run later, and a way to answer a web request from the code that already knows your business.
Every line of that runs on the same runtime your forms do. The queries, the cursor and the library call are ordinary FoxPro; the lambda and the server are what FoxScript adds: no second language, no service to stand up beside it.The keywords and the HTTP APIare each written down in full.
The nightly is rebuilt from every push to main and published as a pre-release on GitHub. Unsigned, so the first launch asks you to confirm.
Each part of the product is written down, including the parts that are not there yet.
Work already mapped out, roughly in the order it is being built.
Today, Waymo is excited to announce our transit rewards program, which we will initially offer in the San Francisco Bay Area, with future cities to follow. This first-of-its-kind initiative rewards riders with Waymo Cash when they connect their Waymo rides with public transit and use their Visa card. We’ll start with employee access before gradually rolling out to the public in the coming weeks.
Public transit networks and shared autonomous vehicles are part of a complementary mobility ecosystem. As transit agencies modernize their payment systems to accept tap-to-pay debit and credit cards, multimodal trips are becoming more convenient than ever.
“Governor Newsom has challenged all of us to build a transportation system that meets Californians where they are, and that means making the trip to the station as effortless as the train ride itself,” said California Transportation Secretary Toks Omishakin. “CalSTA welcomes innovations that reward people for choosing transit and make a multimodal trip feel like one seamless journey.”
Recent surveys show that over 50% of Waymo riders in our most mature service areas (San Francisco, Los Angeles, and Phoenix) use public transit, so we’re proud to invest directly in this connectivity to make everyday commutes easier and more affordable. The transit rewards program works alongside the regional transfer discount programs and tap-to-pay model launched in the San Francisco Bay Area last December. The program applies to all 27 San Francisco Bay Area transit agencies that accept contactless Visa tap-to-pay.
“San José has always believed innovation should make everyday life work better,” said San José Mayor Matt Mahan. “This partnership does exactly that by using new technology to strengthen public transit, not replace it, making regional travel more seamless and giving people more ways to get where they need to go.”
To participate, riders will link their Visa card in the Waymo app. When riders in the SF Bay Area use their Visa card to take a Waymo trip and public transit within 2 hours of each other, we will automatically issue $2.85 in Waymo Cash, which is the cost of a bus fare in San Francisco.
Alongside the launch of our transit rewards program, we’re also announcing a partnership with Caltrain, which accepts tap-to-pay. Waymo will lease 40 dedicated Caltrain station parking spaces to better serve transit riders with staged vehicles. This is a model we hope to replicate with additional transit agencies, ultimately enabling more seamless connections between Waymo’s service and public transit.
This launch is the next step toward making our vision of shared mobility a reality. By taking the learnings from our transit credit pilot programs in San Francisco and Los Angeles, adding Waymo’s service to the LA Metro Mobility Wallet program, and kicking off a strategic microtransit partnership with Via in Chandler, AZ, we’ve been steadily learning what it takes to offer a successful transit integration product and are pleased to formally launch transit rewards.
“Shared mobility is a team effort, and our autonomous ride-hailing service helps expand clean transportation options for everyone,” said Adam Lenz, Head of Sustainability & Environment at Waymo. “We’re thrilled to partner with local transit agencies to fill multimodal gaps and make daily commutes easier through our new transit rewards program.”
Our goal is to expand this transit rewards program and transit agency partnerships to more cities over time, ultimately making shared, all-electric transportation more accessible and affordable for everyone.
Data-only attacks, those that do not affect a program’s control flow, have long been considered too sophisticated and niche to pose a practical threat. With our research, however, we have built a tool that automatically generates them with surprising ease. We explain how such attacks work, and why our tool, Einstein, calls upon both researchers and vendors alike to rethink their mitigation strategies.
Suppose you are a hacker and you just found a bug that allows you to overwrite data in a victim program. Such a scenario is not uncommon: Microsoft, Google, and Mozilla report that about 70% of their security bugs are indeed such memory safety bugs [1, 2, 3]. The question then becomes, as a hacker, how do you weaponize this bug into a real exploit?
In the past, it would have been relatively straightforward: you could, for example, use the bug to conduct a control-flow hijacking attack, overwriting code pointers in the program [4], forcing it to execute your own malicious code. However, due to decades of research (resulting in defenses such as DEP, CFI, CPI, etc.), it is now very difficult to divert a program’s control flow away from the code that it intends to execute. Hence, weaponizing the bug in such a way is now often infeasible in practice.
In our recently published paper at USENIX Security 2024 [5], we present a practical approach to an entirely different method of exploitation: letting the program execute all of its intended code (e.g., any benign functions, system calls, etc.), but with malicious data. These so-called data-only attacks have been known for quite some time [6], but were assumed to be too application-specific or complex to pose any practical threat [7]. In our work, we show that such assumptions are not justified. In particular, we implemented a scalable and automated solution, Einstein, that demonstrates that building data-only attacks is easy — well within reach of low-effort attackers. In this article, we will discuss the insights that allow Einstein to automatically generate such exploits with surprising ease, and the implications of our findings on software vendors.
An example data-only attack
Let us first walk through one of the classic data-only attacks described in the literature, which exploits a victim web server [6] (simplified for clarity). At start up, the server reads its configuration file to initialize its data. One such configuration option is the CGI-BIN path, which is the directory it uses to execute external programs. In our example, the server sets its cgi_bin_path variable to “/usr/local/server/cgi-bin“. We assume that the server has a program in its CGI-BIN directory, sort_script, that a client can use to sort numbers. Moreover, the server has a memory safety bug that allows a malicious client to overflow some buffer and overwrite, for instance, the contents of the cgi_bin_path variable to “/bin“:
Figure 1: An example memory safety bug that allows the attacker to modify the CGI-BIN path.
Of course, the low-level details of the memory safety bug could differ from this (e.g., it could be a use-after-free rather than a buffer overflow), but nonetheless, the question arises: how could an attacker weaponize such a bug? To answer this question, we will show first how the server interacts with a benign client, then how it interacts with a malicious client:
Figure 2: An example data-only attack that corrupts a server’s CGI-BIN path to execute arbitrary code.
After the server initializes (Fig. 2a), it begins processing requests. A benign client interacts with it as follows (Fig. 2b):
➊ The client sends a “POST sort-script” request with the unsorted numbers “2 1 3” in the request body.
➋ The server concatenates the CGI-BIN path and the request’s path to determine the program to be executed, “/usr/local/server/cgi-bin/sort-script”.
➌ The server executes the program, i.e., the sort script, and passes in “2 1 3” as its input. It does so by invoking the execve system call, which instructs the operating system to run the script on behalf of the server.
➍ The script sorts the numbers and outputs “1 2 3”, which the server forwards to the client in its HTTP response.
Let us now sketch how a malicious client could exploit this (Fig. 2c):
➎ The client exploits the bug to set the cgi_bin_path to the string “/bin” (Fig. 1).
➏ The client sends a “POST /sh” request with “touch /tmp/attacker-was-here” in the request body.
➐ The server concatenates the CGI-BIN path and the request’s path to determine the program to be executed, “/bin/sh”.
➑ The server executes the program, i.e., the system shell, and passes in “touch /tmp/attacker-was-here” as its input. It does so by invoking the execve system call, which instructs the operating system to run the shell on behalf of the server.
➒ The shell creates the file /tmp/attacker-was-here.
There are two important takeaways here.
First, the victim server does not execute any malicious code provided by the client; all harmful actions are triggered by malicious data. The attack effectively modifies only the arguments of the execve syscall. Other than that, the benign and malicious executions are equivalent — when handling a request, the victim server performs the same steps, and executes the same functions, albeit with different arguments.
Second, this attack is very powerful, as it allows the attacker to execute arbitrary programs on the victim machine. In our example, the client only creates the file /tmp/attacker-was-here, but any shell command is possible, e.g., to install a malicious program or to exfiltrate data.
Why data-only attacks are considered difficult
Despite the discovery of data-only attacks almost two decades ago, conventional wisdom says they rarely pose a practical threat, either because they are too application-specific or too complex.
Application-specific. As pointed out by the authors of the example attack, building such an attack “require[s] sophisticated knowledge about program semantics”. In other words, an attacker has to become so familiar with the server’s inner-workings — either through reverse-engineering its code, or studying its protocols, etc. — that they know that out of all the program’s data, the cgi_bin_path variable specifically is security-critical, and that a POST request that is malformed in a very specific way can exploit it. In all likelihood, this kind of labor-intensive, application-specific analysis is prohibitively expensive, and hence, according to conventional wisdom, such data-only attacks are too niche to pose a practical threat.
Complex. Recent approaches to building data-only attacks foray into complex territory, under the assumption that simpler attacks — such as the example attack — are not generally at reach. In particular, they assume the need either to solve complex data-flow constraints using heavyweight analyses, or to deviate the victim program away from the code it intends to execute to circumvent a variety of defenses. Several approaches even go so far as to construct highly complicated, Turing-complete machines — something real-world attackers rarely need.
In our work, we show that these assumptions are all false:
Exploitation requires neither extensive knowledge of the program semantics, nor the solving of complex data-flow constraints, nor the diversion of the control flow in a complicated (or even any) way.
Einstein: “As simple as possible, but not simpler”
Inspired by the quote attributed to Albert Einstein, we present a simple (but not too simple) data-only attack exploitation pipeline, named Einstein, that builds attacks with surprising ease. It generates data-only attacks using an application-agnostic technique, proving that such attacks are well within reach of low-effort attackers.
Application-agnostic. Rather than attempting to understand application-specific semantics (e.g., the corner cases of the HTTP protocol), Einstein targets a universal interface used by any program to communicate with the operating system kernel: its syscalls. In particular, we track the data that ends up in syscall arguments, determining whether an attacker can corrupt them to e.g., execute arbitrary code via execve or modify files in the filesystem via write.
Simple. Moreover, Einstein abstracts away unnecessary complexities and, instead, targets the exploits that are not only the simplest to identify, but also the most promising for an attacker. In particular, Einstein automatically generates exploits for the security-sensitive syscalls along a program’s (already valid) runtime path, and whose arguments are (simply) copied verbatim from attacker-controllable data. As detailed later, this simple approach can automatically generate a surprisingly large number of practical data-only exploits in popular real-world programs.
How Einstein builds the example attack
To explain how Einstein works, we walk through each step of how it builds the example attack and how it crafts the arguments of a security sensitive system call. We assume that the attacker has access to a program that is equivalent to the one deployed by their prospective victim, so they can run the server locally for analysis. Einstein takes the victim program as input, and operates in two stages: first, it generates candidate exploits; and second, it confirms whether each candidate exploit is indeed a working exploit. For an explanation of the finer points of the design beyond the scope of this example — e.g., how Einstein tracks unbounded data, chains together multiple syscalls, etc. — please refer to our paper [5].
Candidate exploit generation. To generate candidate exploits, Einstein tracks all attacker-corruptible data at runtime, determining which can influence the arguments of security-sensitive syscalls. To facilitate this, we first start the server with Einstein’s binary-level instrumentation (Fig. 3a, ➊). The instrumentation adds support for dynamic taint analysis, which allows us to track any “tainted” program data at runtime [8]. The server starts up, initializes its cgi_bin_path, and starts waiting for requests. Einstein models an attacker exploiting the memory safety bug by uniquely tainting any data that it could potentially corrupt, e.g., the string “/usr/local/server/cgi-bin”, but also all other data within reach of it. Additionally, we record the tainted data in a memory snapshot (➋).
Next, Einstein continues executing the program and tracks how the tainted data propagates throughout the program’s execution as the server handles a workload consisting of benign requests (➌). For instance, it sends the “POST /sort-script” request from Fig. 2b. Then, while handling the request, the server passes the tainted string as an argument to the execve syscall. Einstein identifies this flow of attacker-controllable data into a security-sensitive syscall, and records information about it, such as the arguments and their taintedness (➍).
Then, Einstein determines that execve’s pathname and argv parameters are not only tainted with an identifier that corresponds to cgi_bin_path, but they are in fact identical to cgi_bin_path. We refer to this kind of (very) straightforward data flow as an identity data flow. Einstein builds a candidate exploit for the identity data flow by generating (address, value) pairs that specify that the memory write bug could exploit the execve by overwriting the cgi_bin_path from “/usr/local/server/cgi-bin” to “/bin” (➎).
Figure 3a: First, Einstein generates a candidate exploit.
Exploit confirmation. Even though we have identified an identity data flow from attacker data into a syscall argument, we are not guaranteed that it is necessarily exploitable. For example, before executing the file, the server may double-check that it is within some preset, hard-coded directory, thereby mitigating the vulnerability. Hence, we confirm whether each candidate exploit is indeed a working exploit.
To confirm the exploit, we first restart the server (Fig. 3b, ➏). Then, at the point where an attacker may exploit the memory write bug, Einstein overwrites the data that is specified by the candidate exploit, changing cgi_bin_path to “/bin” (➐). Next, we send a workload to exploit the gadget — in this case, a “POST /sh” request with a shell command to create a file (➑). Finally, Einstein confirms that the file is indeed created, thereby confirming that the candidate exploit is indeed a working exploit (➒).
Figure 3b: Second, Einstein confirms the candidate exploit to be a working exploit.
Evaluation
We have seen how Einstein can automatically build the example attack from almost two decades ago. Now, let us see how it performs against popular servers today. Although we follow previous work by targeting server applications, we note that other types of applications (e.g., those with more limited user-input interaction and few obviously dangerous syscalls) may well be at risk.
We target the web servers httpd, lighttpd, and nginx; and the database servers postgres and redis — all of which have been shown to be at risk of weaponizable memory write bugs. To generate workloads for the target servers, we use their test suites. Moreover, because nginx is a common target for exploitation case studies, we confirm the candidate exploits that Einstein generates for nginx.
We first refer to Table 1, which presents the number of attacker-tainted syscalls per target program, and the percentage of those that have an identity data flow from attacker data. We observe that an attacker may corrupt many security-sensitive syscall arguments, and (with the exception of postgres) the high rate of identity flows (84–98%) allows us to generate candidate exploits for the vast majority of them. We refer to our paper [5] for a full breakdown per syscall argument. Despite the test suites generating relatively low code coverage (27–49%), Einstein still uncovers many security-sensitive issues. Further work into increasing coverage would undoubtedly yield even more dire results.
Target program
Security-sensitive syscalls with tainted arguments (% with an identity data flow from attacker data)
Code coverage
httpd
1834 (97%)
27.3%
lighttpd
92 (98%)
27.8%
nginx
1623 (82%)
49.1%
postgres
2105 (27%)
46.5%
redis
218 (84%)
33.6%
Table 1: Number of attacker-tainted syscalls per target program.
Next, we refer to Table 2, which presents the number of confirmed exploits for nginx. We observe that the exploits offer many primitives to attacker: a vulnerable execve gives us a Code-Execution primitive, vulnerable file-configuring syscalls (e.g., openat) combined with vulnerable file-write syscalls (e.g., write) give us 17 Write-What-Where primitives, vulnerable socket-configuring syscalls (e.g., connect) combined with vulnerable socket-write syscalls (e.g., sendmsg) give us 41 Send-What-Where primitives, etc. We refer to our paper for a description of two such exploits that bypass state-of-the-art mitigations [5]. Despite conventional wisdom dictating that data-only attacks are too complex or too niche to be practical, we can conclude that an attacker indeed has a diverse set of primitives at their disposal against popular server programs, even today.
Attack Primitive
Count
Code-Execution
1
Write-What-Where
17
Write-What
375
Write-Where
79
Send-What-Where
41
Send-What
372
Send-Where
59
Total
944
Table 2: Confirmed exploits for nginx.
Conclusion
We presented Einstein, a data-only attack exploitation pipeline, which automatically builds exploits against popular servers and bypasses state-of-the-art mitigations. The question, then, is how do we properly mitigate such attacks? To answer this, let us consider two features that a proper mitigation would provide: (1) comprehensiveness, i.e., mitigating an attack surface entirely; and (2) practicality, i.e., requiring little effort to deploy, and, therefore, being more amenable to practical adoption.
For control-flow hijacking attacks in the old days, the parameters were relatively well-defined: they would typically overwrite a code pointer, which would then corrupt an indirect branch (e.g., a return instruction). Hence, the mitigations for such attacks (e.g., DEP, CFI, CPI) can generally afford to be both comprehensive and practical.
On the other hand, data-only attacks present a unique challenge: they may overwrite any data (not just code pointers), which may then corrupt any operation (not just indirect branches, or even syscalls, for that matter). Hence, the mitigations for such attacks can generally afford only to be either comprehensive or practical. That is, comprehensive defenses (e.g., memory safety, DFI) are impractical, because they either incur poor performance or require onerous changes in the software or hardware. Meanwhile, practical defenses (e.g., memory error scanning, selective DFI, syscall filtering) are noncomprehensive, because, as Einstein demonstrates, they leave part of the attack surface vulnerable.
Generating exploits that trivially bypass the practical defenses, Einstein highlights that vendors should strongly consider deploying one or more of the comprehensive defenses. Vendors may also use Einstein to mitigate vulnerabilities on a case-by-case basis, however, a general solution — e.g., by making the comprehensive defenses more practical — poses a pressing direction for future research.
Article Categories:
Security
Last updated July 1, 2024
Authors:
Brian Johannesmeyer is a PhD candidate with VUSec. His research focuses on using program analysis techniques to uncover security vulnerabilities. When he’s not revolutionizing the world of computer security, you can either find him perfecting his taco recipe or camping in the mountains of southern Arizona (or sometimes both, at the same time) [webpage].
Herbert Bos is Full Professor at Vrije Universiteit Amsterdam where he co-leads the VUSec Systems Security group. He is very proud of his current and former students who are all much cleverer than he is. Also, he loves the Beatles.
Cristiano Giuffrida is an Associate Professor at Vrije Universiteit Amsterdam where he co-leads the VUSec Systems Security group. His research interests span across several aspects of computer systems, with a focus on systems security.
Asia Slowinska is a researcher with one foot in industry, and the other in academia. As an Assistant Professor at Vrije Universiteit Amsterdam, she conducts research into systems security. Owing to her unique combination of both industrial and academic experience, she hopes to foster further communication and collaboration, and eventually make the world a safer place.