Preloader

News

Six-year-old breaks women's world Rubik's Cube record – video

Six-year-old breaks women’s world Rubik’s Cube record – video

Six-year-old Lian Yunzhi broke the women’s 3×3 average world record twice at two World Cube Association-certified competitions in Wuhan and Guangzhou, according to state broadcaster CCTV on Monday. She posted averages of 4.52 and 4.27 seconds, after solving the Rubik’s Cube five times. She is the first female speedcuber to average less than 4.5 seconds.

Explore more on these topics


Source: Technology

Show HN: Npunlock – Run custom C kernels for Intel NPUs

npunlock

Intel ships programmable SHAVE cores inside its NPUs, but the public stack
exposes only graph-level programming. npunlock reconstructs the missing path
from custom C code to a runnable NPU kernel.

How npunlock adds custom C kernels to an Intel NPU graph

The current implementation has been verified on Windows x64 with Meteor Lake /
NPU3720.

Latest breakthrough — 2026-09-23: One native graph can execute independent FP32-unary and FP16-binary custom branches; explicit ACT-group preflight handles the compiler’s branch reordering. Evidence and limits.

Quick example

This complete FP32 GELU example embeds the C kernel in Python, places it in an
NPU graph, and checks the result against NumPy. The bundled
npunlock/npu3720_kernel.h target header supplies the NPU3720 invocation and
tensor-address helpers. The tested MoviTools toolchain makes most conventional
libm functions available to kernels without including <math.h>; this
example calls tanhf directly. See the
mlibm.a symbol inventory for the observed candidates.

import numpy as np
import npunlock as npu

npu.configure(movi_dll_dir=r"C:pathtoMVC_DEPEND")

gelu_c: bytes = b"""
#define MLIBM_DEFINE_LINK_COMPAT 1
#include <npunlock/npu3720_kernel.h>

void controlled_act(unsigned layerParams) {
    act_abi_invocation invocation;
    ACT_ABI_LOAD_INVOCATION32_OR_RETURN(layerParams, invocation);
    const float *in = ACT_ABI_INPUT_PTR32(const float, invocation, 0u);
    float *out = ACT_ABI_OUTPUT_PTR32(float, invocation, 1u);
    const float SQRT_2_DIV_PI = 0.7978845608028654f;
    for (unsigned i = 0; i < invocation.element_count; ++i) {
        float x = in[i];
        float w = x + 0.044715f * x * x * x;
        w = tanhf(w * SQRT_2_DIV_PI);
        out[i] = 0.5f * x * (1.0f + w);
    }
}
"""

N = 2048
x = npu.input("x", shape=(1, N), dtype="f32")
y = npu.custom(
    x,
    source=gelu_c,
    carrier="Abs",
    _name="y",
)

program = npu.compile(npu.Graph(inputs=[x], outputs=[y], name="gelu_f32_example"))

input_value = np.linspace(-4, 4, N, dtype=np.float32).reshape(1, -1)
output = program.run({"x": input_value})["y"]
reference = 0.5 * input_value * (
    1.0
    + np.tanh(
        np.sqrt(2.0 / np.pi)
        * (input_value + 0.044715 * input_value**3)
    )
)
print(f"maximum absolute error: {np.max(np.abs(output - reference)):g}")

The same code is available as the runnable
FP32 GELU example. See also the
FP16 GELU and
multi-layer two-input examples,
plus a
mixed-precision graph with unary and binary custom branches.

Why npunlock?

Intel’s normal NPU software accepts graphs made from operations its compiler
supports; it does not expose a public workflow for supplying a C implementation
for an operation. The NPU’s ACT-SHAVE processors are programmable and run
software kernels. npunlock makes those processors usable for compatible
custom graph operations while retaining Intel’s compiler and driver for the
surrounding graph and hardware execution.

Requirements

  • Windows x64
  • Meteor Lake / Intel NPU3720
  • an installed Intel NPU driver for the device
  • Python 3.10 or newer
  • CMake 3.24 or newer and an installed MSVC toolchain for source installation
  • the extracted MoviTools MVC_DEPEND toolchain for custom C compilation

OpenVINO is not required as a runtime, Python package, or compiler frontend.
npunlock does emit OpenVINO-format IR for the installed Intel driver.

Install

npunlock is currently installed from a source checkout:

python -m pip install .

The build bundles npunlock.dll and npunlock_worker.exe inside the Python
package, so normal Python use does not require a separate native path.

Get MoviTools

Custom C compilation uses Intel/Movidius MoviTools, which npunlock does not
redistribute or download.

A MoviTools package verified to work was found in a legacy Lenovo driver pack. See Getting MoviTools for the official download,
hash, extraction command, and expected layout.

Extract the MVC_DEPEND payload from Lenovo’s older
Intel NPU driver package 31.0.100.1688, but remember, DO NOT install or downgrade to
that driver
. All we need is the bundled MoviTools.

Run an example

Point npunlock at the extracted MVC_DEPEND root and run GELU:

$env:NPUNLOCK_MOVITOOLS_DIR = 'C:pathtoMVC_DEPEND'
python examplesexample_gelu.py

The example runs on the NPU and reports its maximum error against a NumPy
reference.

What currently works

  • compile user-written C into ACT-SHAVE machine code
  • run custom kernels inside Intel NPU graphs
  • static dense FP16 unary and two-input custom kernels
  • a verified unary FP32 path
  • one graph containing independent FP32-unary and FP16-binary custom branches
  • nonlinear math such as GELU and tanhf
  • reusable NumPy-compatible host/NPU shared input and output buffers
  • Python, CLI, and native C APIs

Current limitations

Support is experimental and currently limited to Windows x64, Meteor Lake /
NPU3720, static shapes, compatible ACT carriers, and known tensor layouts.
Connected mixed-precision conversion groups are not yet patch-discoverable;
the verified mixed-precision example uses independent branches. Other NPU
generations have not been verified. See
Current limitations for the full compatibility boundary.

Help test Linux and newer NPUs

Have an NPU3720 Linux system or a newer Intel NPU? Contributions are welcome.
Two routes look especially promising but remain untested:

  • a patched NPU3720 graph produced on Windows may run on Linux because the NPU
    firmware executes the custom machine code; building SHAVE code on Linux would
    additionally require a way to load the Windows MoviTools DLLs;
  • newer NPUs may execute the existing 3720xx SHAVE image, or an older OEM
    driver package for that generation may provide matching MoviTools components.

Both need hardware validation, driver/firmware version records, and output
comparison against a host oracle. If you can test either path, feedback, failure
reports, and code contributions are welcome. See
Porting to Linux and newer NPUs for the hypotheses, caveats,
and a suggested test plan.

Documentation

Warning

A note on AI use: I did use AI while building this project–for scaffolding, repetitive implementation work, converting my reverse-engineered results into organized documentation, and fixing my English. The reverse engineering, experiments, debugging, and technical conclusions came from hands-on work. If that doesn’t bother you, there’s a pretty deep and surprisingly satisfying rabbit hole ahead.

License

npunlock is licensed under the Apache License 2.0. MoviTools and
the Intel/Movidius libraries are external proprietary dependencies and are not
covered or redistributed by this repository.


Source: Hacker News

Six-year-old Chinese girl ‘sets new women’s Rubik’s Cube world record’

Six-year-old breaks women’s world Rubik’s Cube record – video

Six-year-old Chinese girl ‘sets new women’s Rubik’s Cube world record’

Lian Yunzhi solves 3×3 puzzle on several occasions in average time of under 4.5 seconds, according to state media

A six-year-old Chinese girl has set a new women’s world record for solving a Rubik’s Cube in just over four seconds, state media reported.

“Speedcuber” Lian Yunzhi broke the women’s 3×3 average world record twice this month at two World Cube Association certified competitions in Wuhan and Guangzhou, according to state broadcaster CCTV on Monday.

She posted averages of 4.52 and 4.27 seconds, after solving the Rubik’s Cube five times.

The girl, who is from Chengdu, is “the first female speedcuber to average under 4.5 seconds”, CCTV said.

A video posted by the broadcaster showed her wearing a red gingham bow on her head as she deftly turned and twisted a Rubik’s Cube.

After solving the puzzle she was seen placing the cube on the table, clapping her hands and pumping a fist in the air.

Nicknamed “Zhizhi”, the six-year-old was introduced to the Rubik’s Cube after being enrolled in a class – alongside various other extracurricular activities such as dance and music lessons, her mother, Lin Yangxi, told local media outlets.

Lin “gradually realised” that cubing was an “excellent activity for fostering a child’s comprehensive abilities” and shifted her focus towards it, she told Chengdu’s Red Star News.

Her daughter has now mastered more than 1,300 Rubik’s Cube solving algorithms, practising daily for two to three hours, Lin said.

Photos from Red Star News showed the girl’s trophies displayed on a shelf in her home.

“Solving a Rubik’s Cube improves focus, memory, logical thinking, spatial awareness, and helps develop a healthy mindset toward winning and losing,” Lin said.

“It’s not just about blindly memorising formulas, but using logical reasoning to find optimal solutions,” she added.


Source: Technology

No Easy Fix for Bogus Respondents in Online Opt-In Polls

No Easy Fix for Bogus Respondents in Online Opt-In Polls

In a test of three screening methods, voter file matching increased error by removing valid cases

About this research

This study was designed to measure the impact of bogus respondents on opt-in surveys. It also compares three methods for identifying and removing bogus respondents: 1) the use of trap questions, sometimes called “attention checks,” 2) the Sentry prescreening system proprietary to CloudResearch, and 3) matching respondents to a national voter file.

Why did we do this?

Pew Research Center does high-quality research to help the public, the media and decision-makers understand important topics. Our methodological research investigates the current challenges facing the polling industry, including past work on the impact of bogus respondents in opt-in surveys.

Learn more about Pew Research Center and our methodological research.

How did we do this?

We fielded a large online opt-in survey Nov. 14-19, 2024, among 11,114 U.S. adult respondents. Respondents who agreed to provide their name and contact information were matched to a registered voter file after the survey concluded.

We evaluated how each approach for removing bogus cases performed on three data quality metrics: “yea-saying,” or agreeing regardless of what is asked, 2) quality of open-end text responses, and 3) response order effects. We also examined how screening methods substantively affected estimates of voter turnout and vote choice in the 2024 presidential election.

Here are the questions used for this report the survey methodology.

One of the most urgent problems in online opt-in polling is bogus (or fraudulent) respondents. These are survey-takers who make no effort to answer questions truthfully and instead are just looking to finish surveys quickly and collect rewards.

To combat this threat, researchers have developed various ways to identify and purge bogus cases from survey samples. Approaches include 1) trap questions, sometimes called “attention checks,” that genuine respondents should always answer correctly, 2) automated prescreening services, and 3) matching respondents to a registered voter file.

But how well do they work? A new Pew Research Center study finds that:

A table showing Matching online opt-in respondents to voter file did not reduce error
  • Overall, purging bogus cases lowers error on most metrics, but a surefire solution remains elusive.
  • Matching an opt-in sample to voter files slightly increased error by removing mostly good respondents (e.g., those who simply declined to give their name or address).
  • Trap questions and an automated prescreening service performed similarly, improving data quality somewhat.
  • All three approaches modestly increased the overestimation of Democratic support in the 2024 election. This appears to be due not to a systematic partisan bias but to bogus respondents’ tendency to say they voted for the winning candidate – in this case, Donald Trump, the Republican.

Matching opt-in samples to voter files may serve other purposes, such as providing data on respondents’ voting history. This study speaks only to whether this is an effective tool for purging bogus cases.

Do all polls have bogus respondents?

No. Bogus respondents are primarily a threat to online opt-in polls, which are recruited through methods like online advertising, self-enrollment and email lists. Polls recruited offline using random sampling (e.g., Pew Research Center’s American Trends Panel) are generally immune to this threat, though they face other challenges.

Related: Do AI and bogus respondents threaten polling’s future?

Methods for removing bogus respondents

Studies conducted by Center researchers and others have found that standard data quality checks such as looking for “speeders” who complete surveys too quickly or “straightliners” who consistently select the same answer choice (e.g., always pick the first response option) fail to detect most bogus respondents. The same goes for many trap questions or attention checks, which not only fail to detect bogus respondents but can also confuse legitimate respondents, leading to false positives.

A table showing Pros and cons of 3 approaches for removing bogus respondents from opt-in samples

Faced with these challenges, pollsters have developed a variety of new approaches in hopes of better identifying bogus respondents.

In this study, we evaluate three such approaches:

  • Trap questions about unlikely behaviors
  • Automated prescreening by a leading fraud-detection company
  • Matching respondents to a commercial voter file

Trap questions

A difficulty in identifying bogus respondents lies in the fact that, for most questions, we have no way of knowing if a respondent has answered truthfully. To get around this problem, we asked respondents if they had ever engaged in four activities where we can be virtually certain that the true answer is “No.” Under this screening procedure, respondents were coded as bogus if they answered “Yes” to one or more of these questions.

The questions were designed so that diligent respondents would not be confused about how to answer, while bogus respondents trying to appear eligible for surveys targeting specific groups or answering randomly would be more likely to answer “Yes.” Specifically:

  • We asked respondents if they used any of six different social media platforms, one of which was a made-up platform called Fizzypress. We also asked if they had received any of six different government benefits in 2023, including payments from the United States Railroad Administration (USRA) – a real government agency, but one that ceased to exist in 1920. Both of these were asked early on in the survey and were intended to resemble screening questions looking for users of specific social media platforms or recipients of certain kinds of benefits. Of the full sample, 6% said they used Fizzypress and 9% claimed to have received USRA payments.
  • Later in the survey, we asked respondents if they had ever visited the International Space Station; 10% said they had. Finally, we asked if they had ever served on a Polar-class icebreaker ship, only two of which were ever built and only one of which remains in active use by the U.S. Coast Guard; 8% answered in the affirmative.

A total of 1,963 cases (18%) answered “Yes” to one or more of these trap questions and were flagged as bogus.

Automated prescreening

CloudResearch’s proprietary Sentry prescreening system is designed to automatically identify problematic respondents before they begin a survey. When a potential respondent first clicks the survey invitation link, they are routed to the Sentry platform and asked a short series of questions designed to elicit problematic survey-taking behaviors like yea-saying and inattentive responding. Open-end answers are checked to confirm that they are responsive to the question asked and not pasted from another source via an automated process.

The system also performs passive checks using metadata about the respondent’s device, location and browser for other signs that they are misrepresenting themselves through technical means.

Respondents who pass these checks are then routed to take the survey. Typically, those who fail are not forwarded to the main survey, but for this study, no respondents were terminated for failing the preescreening checks. We used metadata about which cases failed and why to simulate what would have happened to the survey results if they were excluded at the outset.

Under this procedure, 5,369 cases – nearly half of all completes – failed at least one prescreening check, including 1,879 (35% of failed cases) that failed multiple.

  • 70% of failed cases had a problematic open-end, making this the most frequently failed check by far. This was followed by yea-saying (47%), inattentiveness (12%) and duplicate IP addresses (10%).
  • 7% of these cases failed checks for fraudulent behavior, and only 1% failed other passive technical checks.

Voter file matching

The third approach we tested involves asking respondents to provide their name and address and then looking for a matching record in a commercial voter file, a national database of nearly all registered voters in the United States. If a matching record can be found, the case is considered valid. If no matching record is found, the case is thrown out.

This method assumes that respondents who can be matched are most likely being honest about their identity, and that someone willing to provide detailed contact information that can be validated against official voting records will likely be diligent about answering other survey questions.

This approach has primarily been used by pollsters who work for political campaigns. For cases that are successfully matched, voter files can provide a great deal of information beyond what was asked in the survey, such as a respondent’s registration status and voting history.

But while many voter file vendors attempt to include the unregistered population, Center research has found that a sizable share of this group is not covered by these databases. This can lead to throwing away data for many otherwise-valid respondents who simply aren’t registered to vote. This is not a large drawback for political pollsters, who typically focus on surveying registered voters. But if certain kinds of people who are more likely to be unregistered are underrepresented in a sample or missing altogether, it may lead to biased results if applied to a general population survey.

In this study, we asked respondents if they were willing to provide their name and home address so that they could be matched to the voter file. Out of all respondents, 49% agreed to provide their contact information; of these, 73% were successfully matched to the TargetSmart voter file based on name, address, age and sex. Altogether, 3,977 cases (36% of the full sample) were successfully matched while 7,137 (64%) were screened out under this procedure.

Distinct demographic patterns in flagged cases

A bar chart showing Across demographic groups, prescreening and voter file matching flag far more cases than trap questions as potentially bogus

Although the overall number of cases screened out by each method varies dramatically, respondents claiming to have certain demographic characteristics were consistently more likely to be flagged. Specifically, trap questions and prescreening were especially likely to flag those claiming to be ages 18 to 29 or Hispanic – the two groups found to be most affected by bogus responding in previous Center studies.

This should not be taken to mean that members of these demographic groups tend to be poor survey-takers. Rather, insincere respondents tend to claim membership in these groups, likely falsely in many cases.

The profile of cases purged using voter file matching was different. This reflects the fact that voter file matching tends to purge cases for benign reasons (e.g., a respondent’s privacy concerns or their not being a registered voter), not harmful ones. Hispanic cases were no more likely to be screened out than non-Hispanic Black cases, and both were only 6 percentage points more likely to be screened out than non-Hispanic White cases. In contrast, Hispanic cases were more likely to be flagged than non-Hispanic White cases by trap questions (+22 points) and prescreening (+19).

Likewise, trap questions and prescreening were both more likely to flag men than women, by margins of 9 and 6 points, respectively. Men and women were screened out by voter file matching at roughly the same rate.

The three methods differed on education:

  • With trap questions, college graduates (+11) and respondents with a high school education or less (+7) were more likely to be flagged than those with some college education.
  • With prescreening, high school or less cases (+13) were more likely to be flagged than college graduates, who had the next-highest flag rate.
  • Voter file matching showed much less differentiation: Rates across education groups did not vary by more than 4 points.

Effect of screening on measures of data quality

What impact do these three screening methods have on data quality? There is no direct way to know what proportion of bogus respondents are successfully identified by each method, nor can we be sure how many valid cases are misidentified as fraudulent. Instead, we can compare various measures of data quality calculated before and after each screening method has been applied and the flagged cases removed. To ensure that results are comparable across screening methods, separate sets of survey weights were created for the unscreened sample and for the set of cases remaining after each method was applied. For details, refer to the methodology.

In this study, we focused on three measures of data quality: 1) the frequency answering “Yes” to yes/no questions, sometimes called “yea-saying,” 2) the quality of text responses to open-ended questions, and 3) the severity of response order effects when the order of answer choices is randomized.

Yea-saying

In a 2023 Pew Research Center study, one indicator of bogus responding was a tendency to answer yes/no questions in the affirmative, regardless of the question. In that study, 8% of online opt-in respondents answered “Yes” to at least 10 of 16 yes/no questions asked, with especially high shares among 18- to 29-year-olds (15%) and Hispanic respondents (19%). In contrast, the corresponding shares among respondents from probability-based panels fell between 1% and 2%. Many of these questions measured rare or uncommon characteristics, and we would expect that virtually no one answering truthfully would say “Yes” to 10 or more.

A bar chart showing Trap questions nearly eliminate yea-saying while voter file matching makes it worse

Our latest survey exhibits largely the same pattern. Prior to any screening, 7% of all adults, 10% of 18- to 29-year-olds and 13% of Hispanic respondents answered “Yes” to at least 10 of 15 yes/no questions (excluding a question measuring Hispanic ethnicity and questions used in any screening procedures).

After screening with trap questions, this behavior is largely eliminated from the remaining respondents: 1% of adults and no more than 2% of any demographic subgroup answered 10 or more yes/no questions in the affirmative. The reduction is almost as large for prescreening, which reduced the shares to 2% for all adults, 4% for young adults and 3% for Hispanic adults.

In contrast, matching to the voter file appears to have made the problem worse. The share among all adults increased slightly to 9%, and the shares among young and Hispanic adults rose to 15% and 17%, respectively.

This surprising result appears to be due to two factors. The first is that respondents who declined to provide their contact information – nearly half of the sample – appear much less likely to be bogus than those who agreed to provide this information. Among those who declined, only 3% answered “Yes” to 10 or more yes/no questions, compared with 12% among those who agreed.

The second is that a nontrivial share of bogus respondents were able to provide contact information that was accurate and detailed enough to result in a successful voter file match, though the information provided may not be their own. To the extent that some bogus respondents were successfully removed, it was not enough to offset the loss of valid respondents who declined to share contact information.

Open-end response quality

Examining text answers to open-ended questions has long been considered a reliable way to identify low-quality respondents. While attentive respondents give answers that are relevant to the question asked, bogus respondents often give nonsensical or gibberish answers.

A bar chart showing Voter file matching had minimal effect on open-end response quality

This survey included an open-ended question that asked, “What is one thing that you would like politicians in Washington, D.C. to know about your own situation when they are writing laws and setting policy? Please share as much detail as you can.”

Each respondent’s answer was reviewed and assigned to one the following categories: 1) relevant, 2) nonresponse, 3) generic positive rating, 4) other non sequitur, 5) gibberish, or 6) probable AI (refer to the methodology for details). Because respondents had been told they could skip any question they did not wish to answer, both relevant and nonresponse answers were considered unproblematic, while the remaining categories were deemed problematic.

Prior to any screening, 88% of answers were classified as unproblematic, including 69% that were relevant and 19% that were nonresponse. The most common kind of problematic answers were other non sequiturs (7%), followed by generic positive ratings (2%), gibberish (2%) and probable AI (1%). It is worth noting that the probable AI responses detected were written in a distinct style that was relatively easy to spot. There may be other, more subtle, AI responses that were not detected.

Among all adults, trap questions and prescreening performed similarly, bringing the overall share of problematic open-ends down to 7% and 5%, respectively. Prescreening also resulted in a higher share of relevant answers than trap questions (80% vs. 74%) and a lower share of nonresponse (15% vs. 19%). This may be because the prescreening process itself includes an open-ended question. Matching to voter files, on the other hand, had virtually no effect on the proportion of problematic answers, though the proportion of relevant answers increased to 73%.

Response order effects

If a respondent is being attentive and answering questions accurately, the answer options they select shouldn’t be affected by the order in which they are presented. But inattentive respondents in online surveys tend to select answer choices that appear toward the top of the list, otherwise known as a “primacy effect.”

Our survey included a question asking if undocumented immigrants should be allowed to remain in the country, after which a random half of respondents were shown answer choices in the following order:

  • “They should not be allowed to stay in the country legally.”
  • “There should be a way for them to stay in the country legally, if certain requirements are met.”

For the other half of respondents, this order was reversed. If all respondents were answering diligently, the results from each half of the sample would be largely the same.

A table showing Trap questions and prescreening reduced primacy effects on immigration question; voter file matching made them worse

Instead, we saw a substantial primacy effect on this question. Without any screening, 43% of adults endorse the view that undocumented immigrants should not be allowed to stay legally when that option is presented first. The share drops to 34% when the order is reversed, a primacy effect of +9 percentage points.

The primacy effects were even larger for those subgroups most affected by bogus responding, at +13 points for men, +15 points for 18- to 29-year-olds and +17 points for Hispanic adults.

  • Screening with trap questions reduced the primacy effect by roughly half, to +4 points for all adults, +6 for men, +7 for young adults and +9 for Hispanic adults.
  • Prescreening performed similarly, reducing the primacy effect to between +5 and +10 points.
  • Voter file matching, on the other hand, resulted in even larger primacy effects than when there was no screening at all, with magnitudes increasing to +14 for all adults, +18 for men, +29 for young adults and +25 for Hispanic adults.

As with the frequency of giving “Yes” answers, voter file matching’s negative impact on data quality appears to be due to valid respondents choosing not to provide contact information for matching, with those cases showing a primacy effect of only +2 points.

A table showing Voter file matching made primacy effects worse

Including the question about the status of undocumented immigrants, there were a total of 13 questions with randomized response options that were asked of all respondents. Nearly all of these showed the same pattern. Without any screening, the average primacy effect on these questions was +7 percentage points for all adults, and as high as +11 and +12 points for young adults and Hispanic adults, respectively. Trap questions reduced the average primacy effect by 4 points for all adults and by 7 points for young and Hispanic adults, while prescreening did almost as well.

How does screening affect questions about voter turnout and vote choice?

The prevalence of online opt-in samples in election polling raises the question of how different screening methods – voter file matching in particular – impact estimates related to voter turnout and vote choice. Fielded shortly after the 2024 U.S. presidential election, this survey asked respondents if they voted, and if so, for whom.

Without any screening, an estimated 84% of self-reported registered voters said they voted in the 2024 presidential election. Purging cases with trap questions and prescreening both increased that share by 3 points, to 87%, while voter file matching increased it by 1 point, to 85%. This pattern is more pronounced among young adults and Hispanic adults: For these groups, trap questions and prescreening increased estimated turnout by between 6 and 9 percentage points, while voter file matching produced a 1-point increase among young adults and a 2-point decrease among Hispanic adults.

Although an increase in voter turnout would appear to make these estimates less accurate relative to a higher quality estimate of roughly 77%, it is notable that, for this question, the order of response options was not randomized and the response option for having voted was presented last.1 This kind of pattern is what we would expect to see if bogus respondents, who are more likely to select answer choices toward the top of the list, are being screened out. The fact that the effect is smaller for voter file matching is also consistent with that method’s tendency to disproportionately exclude valid respondents.

All three screening methods increased estimated support for Harris

When it comes to presidential vote (for which response option order was randomized), the weighted raw sample showed Donald Trump and Kamala Harris tied, each with 48% of the vote – accurate to Harris’ true vote share, but underestimating Trump’s by 2 points. Trap questions slightly shifted the result from a tie to a 3-point advantage for Harris. The effect was larger for prescreening and voter file matching, which shifted Harris’ margin to +6 and +7 points, respectively.

These patterns suggest that Harris supporters are overrepresented among valid respondents. But prior to screening, this bias was largely offset by the presence of bogus respondents, who disproportionately said they supported Trump. When bogus respondents were screened out, the bias in favor of Harris became more apparent.

Because the answer choices for Trump and Harris were randomized, this result also implies that bogus respondents were not simply choosing the first option or choosing randomly. There appears to be a subset of bogus respondents, specifically those with the most glaring data quality problems, who were much more likely to say that they voted for Trump, regardless of whether that option was presented first or second. For example, in the unscreened sample, cases coded as having a problematic open-end supported Trump over Harris by a margin of 29 percentage points, while cases that gave 10 or more “Yes” answers favored Trump by 30 points.

Cases with one or both of these data quality problems made up anywhere from 11% of screenouts from voter file matching up to 57% of screenouts from trap questions.

This should not be taken to mean that bogus respondents will always be systematically biased in favor of Republican candidates. In a previous Center benchmarking study where respondents were asked about the 2020 presidential election, opt-in respondents who gave 10 or more “Yes” answers claimed to have voted for Joe Biden, the Democrat, over Trump by margins ranging from 34 to 51 points across three different opt-in samples. It seems plausible that, when asked about past elections, many of these respondents are simply choosing the candidate who won.

These effects are even larger for subgroups in which bogus respondents are most common. Among 18- to 29-year-old voters, the survey initially showed Trump ahead by 2 points prior to screening. All three screening approaches shifted the margin in favor of Harris, putting her ahead by 3 points with trap questions, 6 points with prescreening and 1 point with voter file matching. Among Hispanic voters, trap questions slightly decreased Harris’ lead from 9 points to 8. Prescreening and voter file matching had larger effects, increasing Harris’ lead among Hispanic voters to 16 and 19 points, respectively.

CORRECTION (Aug. 28, 2026): The following sentence was updated to reflect the correct percentages for all U.S. adults, young adults (ages 18 to 29) and Hispanic adults: “The share among all adults increased slightly to 9%, and the shares among young and Hispanic adults rose to 15% and 17%, respectively.” Figures were correct in the corresponding graphic; this change does not affect other findings reported.

  1. 77% voter turnout is based on an analysis of the validated voter dataset from Pew Research Center’s 2024 postelection survey.↩

Source: Hacker News

Abandoning Scientific Linux Was a Mistake

A recent CERN announcement caught my attention. By the end of 2026, more than 2,200 industrial computers and embedded systems around CERN’s accelerator complex are expected to be running Debian 13.

CERN is not abandoning the Red Hat ecosystem. AlmaLinux and RHEL remain the main supported Linux distributions across much of the organization; Debian support is currently limited to accelerator front-end systems.

Still, the move took me back to a decision made more than a decade ago: CERN and Fermilab’s gradual abandonment of Scientific Linux. With hindsight, I think it was a mistake.

Not because Scientific Linux was technically superior to CentOS, or because maintaining another Linux distribution was free. And not because CERN or Fermilab could somehow have controlled what Red Hat or IBM later chose to do.

The mistake was treating Scientific Linux mainly as duplicated engineering work. It was infrastructure.

What Scientific Linux was solving

Scientific Linux grew out of a practical problem in High Energy Physics. Large experiments span laboratories, universities, computing centers, and countries. If one site builds against one version of glibc, another uses something slightly different, and a third runs an entirely different packaging environment, things become painful very quickly.

The scientific community also needed an unusual combination: a Linux distribution that would remain stable for years, work with enterprise software, be freely redistributable, and run across institutions without requiring a commercial license for every machine.

Red Hat Enterprise Linux provided the stability and long lifecycle, and Red Hat published the source needed to rebuild it. Scientific Linux turned that source into a community resource. Fermilab announced the distribution in 2003, and CERN joined soon afterward. It eventually spread far beyond those two laboratories. Universities, research institutions, experiments, companies, and even systems aboard the International Space Station used it or distributions derived from it.

Scientific Linux was never simply “RHEL with a different wallpaper.” It gave the scientific community an institutionally independent implementation of the Enterprise Linux platform.

Then CentOS looked like the obvious answer

When Red Hat and CentOS joined forces in 2014, moving to CentOS looked perfectly reasonable. CentOS already offered what many Scientific Linux users wanted: a freely available Enterprise Linux rebuild backed by a much larger general-purpose community.

CERN began moving its next major release from Scientific Linux CERN to CERN CentOS 7. Scientific Linux 5 and 6 remained supported, but the future platform at CERN would be based on CentOS. The rationale was compelling. Why should CERN and Fermilab spend scarce engineering time rebuilding a Linux distribution when CentOS was already doing essentially the same work?

Why maintain an HEP-specific distribution when the scientific community could converge on a larger common platform? If Red Hat itself was supporting the CentOS project, that seemed to make the platform more sustainable, not less. In 2019, Fermilab took the argument to its logical conclusion or its literal “end.” There would be no Scientific Linux 8. Fermilab would deploy CentOS 8 instead and work with CERN and other laboratories to improve CentOS for high-energy physics.

Taken on its own, this was not an irrational decision. But it removed something that was hard to see on a spreadsheet: a credible exit.

A distribution has option value

This is what I think we underestimated.

The cost of maintaining Scientific Linux was visible. People had to rebuild packages, test updates, maintain repositories, produce releases, handle security updates, and support users.

The benefit of an independent distribution was harder to quantify. As long as CentOS behaved exactly as CERN and Fermilab expected, Scientific Linux looked redundant. That is how redundancy always looks when nothing has failed yet.

Scientific Linux gave the scientific community its own implementation of a RHEL-compatible computing environment. More importantly, it kept alive the people, processes, infrastructure, and institutional knowledge needed to maintain one.

That capability had option value. You might not need to exercise the option this year, or even this decade. It becomes valuable when the assumptions underneath your primary platform change. Those assumptions changed surprisingly quickly.

[!NOTE]

CERN uses a similar argument to justify the FCC: the knowledge required to build large machines such as accelerators and detectors is valuable and must be preserved. If Europe does not build the next machine after the LHC, that expertise may be lost, the next large machine may be built in China, and Europe may lose its technological leadership in the field. The same argument applies to software infrastructure.

The CentOS assumption did not last

In December 2020, Red Hat changed the role of CentOS.

CentOS Linux, the traditional downstream rebuild of released RHEL versions was discontinued in favor of CentOS Stream, which sits ahead of RHEL rather than behind it. This was more than a change in release cadence.

Organizations had standardized on CentOS because they wanted a freely distributable approximation of the current RHEL release. The product they had standardized on effectively ceased to exist.

CERN and Fermilab immediately had to reconsider their Linux strategy. CERN noted that the shorter CentOS Stream lifecycle was incompatible with some scientific use cases. The laboratories evaluated Stream, RHEL licensing arrangements, and the emerging Enterprise Linux rebuilds. By 2022, CERN and Fermilab were recommending AlmaLinux as the standard distribution for experiments.

The circularity is hard to miss.

Scientific Linux had been retired partly because maintaining another RHEL rebuild seemed unnecessary while CentOS existed. A few years later, the community needed an independent RHEL-compatible distribution again, so CERN and Fermilab adopted another independently governed RHEL-compatible distribution.

AlmaLinux is a good project. This is not a criticism of it. The requirement never disappeared; only our implementation of it did.

Then the ground moved again

In 2023, Red Hat changed how RHEL-related sources were publicly distributed. CentOS Stream became the sole public repository for that source material, replacing the previous git.centos.org publication model.

This did not make RHEL “closed source,” despite how the change was sometimes described. That characterization would be inaccurate. But it made the dependency structure of the Enterprise Linux rebuild ecosystem much more obvious.

By then, the lesson should have been familiar: technical compatibility with an upstream product is not the same as independence from the organization that controls it.

Scientific Linux had strategic value that was never properly accounted for. Its existence meant CERN, Fermilab, and other institutions were not merely consumers of an ecosystem. Together, they could reproduce a critical part of that ecosystem themselves.

Once discarded, that capability is much harder to recreate than to keep alive.

And now CERN is moving part of the accelerator complex to Debian

The latest chapter makes this history particularly interesting. CERN’s accelerator controls group first tried to remain within the Red Hat ecosystem. According to the presentation at MiniDebConf Winterthur 2026, the original plan involved CentOS Stream, with Debian prepared as a fallback because, in the presenters’ words, they could not afford new surprises.

Then came a more physical problem. RHEL 9 raised its x86-64 baseline to x86-64-v2. RHEL 10 moved to x86-64-v3.

Replacing old hardware may be inconvenient but manageable for an ordinary server fleet. Accelerator controls are not an ordinary server fleet.

These computers interface with custom electronics, specialized boards, legacy buses, real-time systems, and equipment installed throughout a vast accelerator complex. The hardware can remain operational for decades because replacing one computer may mean redesigning the electronics connected to it.

CERN’s 2023 risk analysis estimated that staying in the Red Hat ecosystem could cost about 5.4 million CHF. It could also require redesigning around eleven boards, hiring more engineers and technicians, reorganizing racks, recabling systems, and recommissioning equipment.

The solution was refreshingly simple: do not replace millions of francs’ worth of functioning hardware to satisfy an operating system’s CPU baseline. Replace the operating system.

CERN chose Debian. By the end of 2026, more than 2,200 industrial computers and embedded systems around the accelerator complex are expected to run Debian 13.

The irony is difficult to miss

What interests me most is not that CERN chose Debian, but why Debian was valuable.

The accelerator team specifically praised its continued support for older and less common architectures and its ability, as a community-led distribution, to compete with Red Hat.

CERN is also sponsoring Freexian to strengthen Debian’s long-term-support ecosystem. There is an important principle here: when an institution depends on infrastructure supplied by a community project, contributing resources to keep that project independent and healthy may cost less than consolidating everything around one vendor ecosystem in pursuit of short-term efficiency.

That sounds remarkably close to what Scientific Linux used to provide.

Scientific Linux was maintained by a surprisingly small group

This is another reason I question whether ending Scientific Linux produced the savings we assumed. The distribution was not maintained by hundreds of engineers.

Its project history lists only a handful of major developers in its later years, with Fermilab providing the main sponsorship, build infrastructure, bandwidth, and website. That does not make the work trivial. Release engineering, security updates, testing, package rebuilds, infrastructure, and user support all take real time.

But those costs should be compared with the value they produced. Scientific Linux provided a stable computing platform to an enormous scientific ecosystem for roughly two decades. Preserving that capability should not have been evaluated solely in terms of the engineer-hours saved by adopting CentOS.

The relevant question was: what would it cost to lose the ability to operate independently if the surrounding ecosystem changed?

The past decade gives us at least a partial answer.

Did Scientific Linux keep Red Hat honest?

There is a tempting, stronger argument: that Scientific Linux and other independent rebuilds constrained Red Hat, and that removing one of the major alternatives helped concentrate power around CentOS and RHEL.

I think there is something to it, but I would not present it as historical fact. We cannot know the counterfactual.

Scientific Linux probably would not have prevented Red Hat from changing CentOS. IBM completed its acquisition of Red Hat in July 2019, after Fermilab had already announced that there would be no Scientific Linux 8. It would therefore be misleading to reduce this to a story about IBM killing something CERN and Fermilab should have anticipated.

What I do believe is that independent alternatives change incentives. A vendor behaves differently when major customers and institutions have a credible exit than when leaving requires them to rebuild years of infrastructure.

This is not specific to Red Hat. It is basic dependency management. Competition matters even when nobody switches; the possibility of switching matters too. Scientific Linux provided that possibility.

And because major scientific institutions maintained it rather than a commercial Linux vendor, its incentives were unusually well aligned with long-lived scientific infrastructure.

The real mistake was confusing standardization with dependency

Standardization is good.

Particle physics could not operate efficiently if every university and laboratory invented its own incompatible computing platform. CERN and Fermilab were right to want common interfaces, compatible binaries, common packaging, and predictable operating environments. But standardization and monoculture are not the same thing.

The strategy went wrong when we assumed that standardization required abandoning an independent implementation of the standard. Scientific Linux could have remained boring.

In fact, boring was exactly what we needed from it. It did not need to compete with Fedora on innovation or Ubuntu on desktop adoption. It needed to remain a reproducible, institutionally controlled Enterprise Linux platform: somewhere the scientific community could go if the commercial ecosystem moved in a direction incompatible with scientific computing.

That is valuable infrastructure even when almost nobody notices it. Especially then.

The lesson is larger than Scientific Linux

I do not think the answer today is to resurrect Scientific Linux. The ecosystem has moved on. AlmaLinux exists. Rocky Linux exists. Debian is proving increasingly useful at CERN. Containers and modern software distribution have reduced the importance of the host operating system for many workloads.

The lesson I take from Scientific Linux is about how research institutions value infrastructure. We are good at calculating the immediate cost of maintaining something ourselves. We are much worse at calculating the long-term cost of losing the ability to maintain it.

When deciding whether to retire institutional open-source infrastructure, the calculation should include more than maintenance hours. It should account for governance, concentration risk, the cost of migration if an upstream project changes direction, the institutional knowledge being discarded, and the value of a credible alternative that you hope never to need.

The CERN accelerator team’s final advice in its Debian presentation is probably the best summary:

“Distribution portability is a very good thing.”

It took more than a decade, several Enterprise Linux strategy changes, a new community rebuild, and now the migration of thousands of accelerator computers to arrive there. Scientific Linux was already teaching us that lesson twenty years ago.

We just stopped listening.


Source: Hacker News

Jev in 25 Lines of Python

Jev in 25 lines of Python

My name is Jev

Everyone and their mom is talking about Jev. Jev this, Jev that. Everyone on Twitter is all over Jev, how it’s the next frontier of large language models and the AI paradigm. We don’t really think so. So here’s Jev in 25 lines of Python.

Load the model.

# /// script
# requires-python = ">=3.12"
# dependencies = ["huggingface-hub", "llama-cpp-python", "numpy"]
# ///

import numpy
from llama_cpp import Llama

# Really, you can use any GGUF model from https://huggingface.co/models?library=gguf

model = Llama.from_pretrained(
    repo_id="Qwen/Qwen3-0.6B-GGUF",
    filename="Qwen3-0.6B-Q8_0.gguf",
    n_ctx=512,
    logits_all=True,
    verbose=False,
)

Load the prompt and define your choices.

labels = ["A", "B", "C"]
choices = ["Legitimate", "Spam", "Phishing"]
email = "Payroll asks for your password on a non-company sign-in page."
options = "n".join(
    f"{label}. {choice}" for label, choice in zip(labels, choices, strict=True)
)
prompt = f"""<|im_start|>system
Choose one option.<|im_end|>
<|im_start|>user
Email: {email}nn{options}<|im_end|>
<|im_start|>assistant
<think>nn</think>nn"""
model.eval(tokens=model.tokenize(text=prompt.encode(), add_bos=False, special=True))

Massage the logits into probabilities.

logits = model.scores[model.n_tokens - 1]
token_ids = [model.tokenize(text=label.encode(), add_bos=False)[0] for label in labels]
choice_logits = numpy.asarray([logits[token_id] for token_id in token_ids])
logprobs = choice_logits - numpy.logaddexp.reduce(choice_logits)
probabilities = numpy.exp(logprobs)

for name, scores in (
    ("Logits", choice_logits),
    ("Log probabilities", logprobs),
    ("Probabilities", probabilities),
):
    values = numpy.round(scores.astype(float), 3).tolist()
    print(f"{name}:", dict(zip(choices, values, strict=True)))

# Logits: {'Legitimate': 26.254, 'Spam': 27.262, 'Phishing': 29.614}
# Log probabilities: {'Legitimate': -3.482, 'Spam': -2.474, 'Phishing': -0.122}
# Probabilities: {'Legitimate': 0.031, 'Spam': 0.084, 'Phishing': 0.885}

There. That’s Jev.

But no, you don’t understand Jev!

Yeah, we know.

But yes. This is Jev.

  • It classifies: it gets a prompt with choices and outputs probabilities.
  • It’s fast.
  • It’s local.
  • You don’t send your data anywhere else.

And we like not sending your data anywhere else. Check out NobodyWho.

(note: this is a parody blog post, see these links for better/more complete open implementations of Jev: OpenJev, openjev-sglang, and OpenJev on DiffusionGemma.)


Everything NobodyWho do is open-source, please leave a
star on Github
to support us ❤️

Published Sep 22, 2026 by Duarte O.Carmo


Technical


Source: Hacker News

Delta: Highly available, strongly consistent storage using chain replication (2022)

Over the years, Meta has invested in a number of storage service offerings that cater to different use cases and workload characteristics. Along the way, we’ve aimed to reduce and converge the systems in the storage space. At the same time, having a dedicated solution for critical package workload makes everyone happier. Having this in place is necessary for our disaster recovery and bootstrap strategy. This realization, coupled with a business need to provide storage for Meta’s build and distribution artifacts, led to the inception of a new object storage service — Delta.

Consider Delta’s positioning in the Meta infrastructure stack (below). It belongs at the very bottom, providing the basic primitive required for the availability and recoverability of the rest of the infrastructure. For bootstrap systems, complexity should be introduced only if it makes the solution more reliable. We’re only minimally concerned with the performance and efficiency of the solution. Another consideration for bootstrap systems involves the bootstrap itself. This process, by which engineers can access a small set of machines and restore the rest of our infrastructure, helps us get the product back up and working for people using it. Lastly, the bootstrap data needs to be backed up for recovery in case disaster strikes.

In this post, we will discuss the goals for Delta, the main concepts that govern Delta’s architecture, Delta’s production use cases, its evolution as a recovery provider, and future work items.

What is Delta?

Delta is a simple, reliable, scalable, low-dependency object storage system. It features only four high-level operations: put, get, delete, and list. Delta trades latency and storage efficiency in favor of simplicity and reliability. Since it’s horizontally scalable, Delta takes on minimal dependencies with appropriate failover strategies for soft dependencies in place.

Delta is not a:

  • General purpose storage system: Delta’s core tenets are resiliency and data protection. It’s designed specifically to be used by low-dependency systems.
  • Filesystem: Delta acts as a simple object storage system. It doesn’t intend to expose filesystem semantics like Posix etc.
  • System optimized for maximum storage efficiency: With resiliency as its primary tenet and focus on critical systems, Delta doesn’t intend to optimize for storage efficiency, latency, or throughput.

Delta’s architecture

Delta has productionized chain replication, an approach to coordinating clusters of fail-stop storage servers. It intends to support large-scale storage services that exhibit high throughput and availability, without sacrificing strong consistency guarantees.

Before diving deeper into how Delta leverages chain replication to replicate client data, let’s first explore the basics of chain replication.

Chain replication

Fundamentally, chain replication organizes servers in a chain in a linear fashion. Much like a linked list, each chain involves a set of hosts that redundantly store replicas of objects.

Each chain contains a sequence of servers. We call the first server the head and the last one the tail. The figure below shows an example of a chain with four servers. Each write request gets directed to the head server. The update pipelines from the head server to the tail server through the chain. Once all the servers have persisted the update, the tail responds to the write request. Read requests are directed only to tail servers. What a client can read from the tail of the chain replicates across all servers belonging to the chain, guaranteeing strong consistency.

Chain replication vs. quorum replication

Now that we have provided an overview of what chain replication entails, let’s explore how chain replication fares against other widely used replication strategies.

  • Storage efficiency: Chain replication clearly does not offer the most storage-efficient replication strategy. We store redundant copies of the whole data set on all hosts in a chain. A comparatively efficient approach would involve intelligently replicating fragments of data using erasure coding techniques.
  • Fault tolerance: In an optimal bucket layout, chain replication can provide similar or better fault tolerance than quorum-based replication mechanisms. Why? A chain with `n` nodes can tolerate failures up to `n – 2` nodes without compromising on availability. On the contrary, for quorum-based replication systems, at least `w` hosts must be available to serve writes. Additionally, `r` hosts must be available to serve reads. Here `w` and `r` represent the write quorum size and read quorum size, respectively.
  • Performance: In replication strategies (like primary backup), all backup servers can serve reads. This increases read throughput. In the native idea of chain replication, only the chain tail can serve reads. (We optimized this bit and will share details in the apportioned queries section later in this post.) Much like quorum-based replication mechanisms, in chain replication all writes are directed to the primary (the head of the chain). But in chain replication, writes are only responded to after all hosts in the chain have acknowledged the update. Thus, chain replication has higher write latency on average in comparison with quorum-based replication mechanisms.
  • Quorum consensus: Quorum-based systems need complex consensus and leader election mechanisms to maintain quorum in the system. In contrast, the scope of quorum consensus in a chain-replication based system gets narrowed down to the much simpler, chain-host mapping. For example, the chain head always serves as a leader for processing writes, without the need for an explicit leader election.

Considering the above differences, chain replication clearly fails to offer the most storage-efficient way to replicate data across machines. Additionally, it yields higher average write latency in comparison to quorum-based systems, as we consider writes successful only when all links in a chain have persisted the update. However, it’s very simple while offering similar fault tolerance and consistency guarantees.

The anatomy of a Delta bucket

Now that we have a preliminary understanding of chain replication, let’s talk about how Delta leverages it to replicate data across multiple servers.

Each Delta bucket above includes several chains. Each chain usually consists of four or more servers, which can vary based on the desired replication factor. Each chain itself acts as a replica set and serves a slice of data and traffic. It can be thought of as a logical shard of a client data set. Servers in a particular chain get spread across different failure domains (power, network, etc.). Doing so guarantees durability and availability of client data if servers in one or more failure domains remain unavailable. We maintain a bucket config, the authoritative chain-host mapping for the layout of the bucket. When we add or remove servers and chains from the bucket, the bucket config gets appropriately updated.

When clients access an object within a Delta bucket, a consistent hash of the object name selects the appropriate chain. Writes are always directed to the head of the appropriate chain. It writes the data to the local storage and forwards the write to the next host in the chain. The write is acknowledged only after the last host in the chain has durably stored the data on local media. Reads are always directed to the tail of the appropriate chain. This guarantees that only fully replicated data is visible and readable, thereby guaranteeing strong consistency.

Delta supports horizontal scalability by adding new servers into the bucket and smartly rebalancing chains to the newly added servers without affecting the service’s availability and throughput. As an example, one tactic is to have servers with the most chains transfer some chains to new servers as a way to rebalance the load. New bucket layouts would still follow the desirable failure-domain distribution, etc., and are employed while rebalancing chains and expanding bucket capacity.

Failure and recovery modes in Delta

Failures can stem from hosts going down, networks being partitioned, planned maintenance activities, operational mishaps, or other unintended events. In a suitable implementation of chain replication, we assume servers to be fail-stop. In other words:

  • Each server halts in response to a failure rather than making erroneous state transitions.
  • A server’s halted state can be detected by the environment.

Consider a Delta bucket with `n` chains, with each chain comprising >1 host. In the event of any host misbehaving or getting partitioned from the network, the other sibling hosts (upstream/downstream) sharing a chain with the culprit host would be able to detect the suspicious host and report this behavior.

Sibling hosts can detect the erroneous behavior of misbehaving hosts by simple heartbeats or by encountering failures in transmitting acknowledgments/requests up/down the chain. If multiple hosts suspect a particular target host, the latter gets kicked out of all its chains and sent for repair. The bucket config gets updated appropriately. 

Some decisions and trade-offs must occur when detecting unhealthy behavior in a host:

  • Timeout settings: We conducted several performance tests to arrive at the right timeout settings. We need to carefully assess the timeout between individual links in a chain before suspecting a host. We can’t set the timeout too short because servers can always experience transient network issues. We can’t set the timeout too long, either, because doing so would negatively affect operation latency. Not only that, but clients may also timeout while awaiting a valid response.
  • Suspicious host voting settings: We need to assess how many hosts should vote for a particular one being unhealthy before the suspected host is kicked out of its chains. The limit can’t be one since it’s always possible for two hosts in a chain to vote each other as unhealthy. This would cause both hosts to be disabled from their chains. The limit can’t be too large, either, as this would lead to the unhealthy host staying in the fleet for an extended period and negatively impacting service performance. Each host belongs to multiple chains and is guaranteed frequent connections with upstream and downstream nodes in their chains. As a result, setting the suspicion voting limit to two has worked well for us. Additionally, we have automated the faulty host repair flow. This provides us with the flexibility to configure sensitive thresholds and have more false positives.

Once the faulty host recovers, it can be added back to all the chains that it served prior to getting kicked out of the bucket. New hosts are always added to the rear end of the chains. Upon being added back to the chains, this host must synchronize itself with all the updates on the chain that occurred while it was not part of the chain. This process consists of scanning the objects on the upstream host and copying those not present or that have an obsolete version. Notably, during this interval of reconstruction, the host can still accept new writes from upstream. However, until it’s fully synchronized with the upstream host, it must defer reads to upstream.

We use this process for adding both suspected hosts as well as introducing new capacity to a Delta bucket.

How Delta has evolved over time

Apportioned queries

As evident from the above description, there are a few major inefficiencies with serving reads in chain replication.

  • The tail, the only node serving both reads and writes, can become a hotspot.
  • Considering that the tail serves all reads, the tail of the chain limits read throughput.

In order to mitigate these limitations, we can let all nodes in the chain serve reads via the idea of chain replication with apportioned queries.

The basic idea? Each node serves read requests. But before responding to client read requests, it does a crucial check. It verifies whether the requested object’s local copy is clean or has been committed by all the servers in the chain. It may alternatively assess it as dirty, meaning the object does not replicate to all servers in the chain. The server can just return the last committed version of the object back to the client. This ensures that clients get only the object version that has been committed by all servers in the chain, thereby retaining chain replication’s strong consistency guarantees. The tail node would serve as the authority of the latest clean version for a particular object.

As explained above, each non-tail link in the chain makes an additional network call to the chain tail. It does so to fetch the clean version of an object before responding to the client reads. This additional network call to the tail link offers great value. Why? It helps scale the read throughput and chain bandwidth linearly with the chain length. Additionally, these object version check calls are cheaper in comparison with serving actual client reads. Hence, they don’t negatively affect client latency in a significant way. 

Automated repair

While hardware failures or network partitions may seem rare, they occur fairly frequently in large clusters. As such, failure detection, host repair, and recovery should be automated to avoid frequent manual interventions.

We built out a control plane service (CPS) responsible for automating Delta’s fleet management. Each instance of the CPS gets configured to monitor a list of Delta buckets. Its primary function includes repairing chains that have missing links.

When repairing chains, the CPS applies a few techniques to achieve maximum efficiency:

  • When repairing a bucket, the CPS must maintain the failure domain distribution of all chains in the bucket layout. This ensures that the bucket hosts are spread evenly across all failure domains.
  • The CPS must ensure a uniform chain distribution on all servers. In this way, we avoid a few servers getting overloaded by hosting significantly more chains in contrast to other hosts.
  • When a chain has a missing link, the CPS would prioritize repairing the chain with the original host rather than a new one. Why is that? Resyncing all chain contents to a new host gets compute-intensive in comparison with syncing partial chain contents to the original host once it returns from repair.
  • Before enabling a host to a chain, the CPS would perform detailed sanity checks to ensure that a healthy host gets added back to the bucket.
  • Apart from managing the servers in the bucket, the CPS would also maintain a pool of healthy standby servers. In the event of a chain missing more than 50 percent of its hosts, the CPS would add a fresh standby host to the chain. This ensures that no severely underhosted chain jeopardizes the availability of the bucket. While doing this, the CPS would attempt to apply all the above principles in a best-effort manner.

Global replication

In Delta’s original implementation, when Delta clients wanted to store a blob in multiple regions due to data safety considerations, they made a request to each region. This is clearly not ideal from a user point of view. Delta’s users should not be in the business of tracking object location(s) while being able to control the level of redundancy and geo-distribution preferences. A client app should be able to just put an object to the store once and expect the underlying service fabric to propagate the change everywhere. The same goes for retrievals. In the steady state, the clients can expect to just get an object by submitting a request and letting the service retrieve it from the most optimal source available at the moment.

As our service evolved, we introduced global replication to client regions. This is done in a hybrid fashion — a combination of synchronously replicating blobs to a few regions and asynchronously replicating blobs to the remaining set of regions. Replicating to a few regions synchronously reduces the client latency while ensuring that other regions get the blob in an eventually consistent manner. Imagine that a particular region experiences a network partition or outage. The system would be intelligent enough to exclude that region from participating in geo-replication until it returns. Then that region could be asynchronously backfilled with the missing blobs.

How Delta handles disaster recovery

One of Delta’s key tenets is to be as low-dependency as possible. Along with having minimal dependencies, we also invested in having a reliable disaster recovery story.

As of now, we have integrated with archival services to continuously back up client blobs to cold storage. Additionally, we have developed the ability to continuously restore objects from these archival services. Out-of-box integration with archival services provides us with reliable recoverability guarantees in severely degraded environments. We have several partner teams that have integrated with our service because of our disaster recovery guarantees.

What’s next for Delta?

Looking ahead, we are working on a centralized backup and restore service for core infrastructure services within Meta based on Delta storage. Any reliable stateful service must be able to produce a snapshot of its internal state (backups) and rehydrate the internal state from a snapshot (restore). Our ultimate goal is to position ourselves as a gateway to all archival services and provide a centralized backup offering to our clients. Additionally, we also plan on investing heavily toward improving the reliability of Delta’s disaster preparedness and recovery story. This will help us serve as a reliable and robust recovery provider for our users.


Source: Hacker News

ReBarUEFI: Resizable BAR for almost any UEFI system

ReBarUEFI

GitHub Actions ReBarDxe
GitHub Actions ReBarState
Downloads

A UEFI DXE driver to enable Resizable BAR on systems which don’t support it officially. This provides performance benefits and is even required for Intel Arc GPUs to function optimally.

screenshot showing cpu-z, gpu-z and amd software

If using an NVIDIA Turing GPU (20 or 16 series) see NvStrapsReBar for enabling Resizable BAR on it.

Requirements

  • (optional) 4G Decoding enabled. See wiki page Enabling hidden 4G decoding if you can’t find an option for it. Without 4G Decoding you will be limited to 1GB BAR and in some cases 512MB you can try to increase this upto 2GB by reducing TOLUD
  • (optional) BIOS support for Large BARs. Patches exist to fix most issues relating to this

Usage

Follow the wiki guide Adding FFS module and continue through the steps. It covers adding the module and the additional modifications needed if required.

Once running the modified firmware make sure that 4G decoding is enabled and CSM is off.

Next run ReBarState which can be found in Releases (if you’re on Linux build with CMake) and set the Resizable BAR size. In most cases you should be able to use 32 (unlimited) without issues but you might need to use a smaller BAR size if 32 doesn’t work

If Resizable BAR works for you reply to List of working motherboards so I can add it to the list. Most firmware will accept unsigned/patched modules with Secure Boot on so you won’t have any problems running certain games.

If you have any issues after enabling Resizable BAR see Common Issues (and fixes)

How it works

The module is added to the UEFI firmware’s DXE volume so it gets executed on every boot. The ReBarDxe module replaces the function PreprocessController of PciHostBridgeResourceAllocationProtocol with a function that checks for Resizable BAR capability and then sets it to the size from the ReBarState NVRAM variable after running the original function.

The new PreprocessController function later gets called during PCI enumeration by the PciBus module which will detect the new BAR size and allocate it accordingly.

AliExpress X99 Tutorial by Miyconst

Resizable BAR on LGA 2011-3 X99

Instructions for applying UEFIPatch not included as it isn’t required for these X99 motherboards. You can follow them below.

UEFI Patching

Most UEFI firmwares have problems handling 64-bit BARs so several patches were created to fix these issues. You can use UEFIPatch to apply these patches located in the UEFIPatch folder. See wiki page Using UEFIPatch for more information on using UEFIPatch. Make sure to check that pad files aren’t changed and if they are use the workaround

Working patches

  • <4GB BAR size limit removal
  • <16GB BAR size limit removal
  • <64GB BAR size limit removal
  • Prevent 64-bit BARs from being downgraded to 32-bit
  • Increase MMIO space from 16-32GB to full usage of 512GB/39-bit range (Skylake/Kaby Lake/Coffee Lake)
  • Increase MMIO space from 8-16GB to full usage of 512GB/39-bit range (Haswell/Broadwell). The issue of older patches being limited to 64GB has been fixed.
  • Increase MMIO space from 16GB to full usage of 64GB/36-bit range (Sandy/Ivy Bridge). Requires DSDT modification on certain motherboards. See wiki page DSDT Patching for more information.
  • Remove NVRAM whitelist to solve ReBarState GetLastError: 5
  • Fix USB 3 ports not working in BIOS with 4G Decoding enabled (Ivy Bridge/Haswell/Broadwell)
  • X79 Above 4G Decoding fix

Build

Use the provided buildffs.py script after cloning inside an edk2 tree to build the DXE driver. ReBarState can be built on Windows or Linux using CMake. See wiki page Building for more information.

FAQ

Will it work on a PCIe Gen2 system ?

Previously it was thought that it won’t work on PCIe Gen2 systems but one user had it work with an i5 2500k.

Can I use Resizable BAR on my system without modifying BIOS ?

You can use Linux with 4G Decoding on, recent versions will automatically resize and allocate GPU BARs. If your BIOS doesn’t have the 4G decoding option (make sure to check hidden) or DSDT is faulty you can then follow the Arch wiki guide for DSDT modification using modifications from DSDT Patching and boot with pci=realloc in your kernel command line. Currently there is no known method to get it on Windows without BIOS modification

I set an unsupported BAR size and my system won’t boot

Clear CMOS and Resizable BAR should be disabled. In some cases it may be necessary to remove the CMOS battery for Resizable BAR to disable.

Will less than optimal BAR sizes still give a performance increase ?

On my system with an i5 3470 and Sapphire Nitro+ RX 580 8GB with Resizable BAR enabled in driver I get an upto 12% FPS increase with 2GB BAR size.

Credit

  • @dsanke, @cursemex, @dripsnek, @val3nt33n, @Mak3rde and @romulus2k4 for testing/helping develop patches

  • The Linux kernel especially the amdgpu driver

  • EDK2 for the base that all OEM UEFI follows

  • Ghidra which was used to patch UEFI modules to workaround artificial limitations

  • @vit9696 for the NVRAM whitelist patches

  • @ZOXZX for helping with the X79 Above 4G patches

  • @NikolajSchlej for developing UEFITool/UEFIPatch

  • QEMU/OVMF made testing hooking way easier although it didn’t have any resizable BAR devices so the only way I could test it was on my actual PC.


Source: Hacker News

Why fans are digging their claws into Wolverine over its patronising gameplay

A game still of Wolverine with claws extended attacks an armed soldier amid flying sparks in an industrial setting

No friction … Marvel’s Wolverine. Photograph: Sony

No friction … Marvel’s Wolverine. Photograph: Sony

Why fans are digging their claws into Wolverine over its patronising gameplay

From highlighted ledges to characters telling you what to do, blockbuster games are increasingly taking the challenge out of playing – and players are pushing back

Don’t get Pushing Buttons delivered to your inbox? Sign up here

Marvel’s Wolverine, the latest action game from Sony’s stable of expensive blockbuster-factory studios, has had an unfortunate launch. Reviews were decent, if not stellar, but a series of viral video clips from reviewers and streamers have made it the scapegoat for a patronising flavour of game design that gamers have hated for decades. This clip, from Skill Up’s review, shows the player entering a beautiful wintry landscape before being overwhelmed by visual cues telling them what they’re supposed to be doing there. Sonar pulses point you in the direction of objectives. Targets – which, it bears emphasising, are shaped like targets – get an extra reflective visual effect so you know you’re supposed to pay attention to them. A trail shows you not just where to go, but where the level’s sole hidden collectible is.

It is hand-holding to the point of parody. Indeed, it reminded me of the sniggersome Mario in Unreal Engine satirical videos that makes the social media rounds every few months (“A mushroom? I should take a look at this”). Skill Up’s clip has been followed by hundreds of others, gleefully ripping into Wolverine with ironic savagery. In one much-copied clip, a streamer simply lays down his controller during a motorbike chase sequence and watches as Wolverine speeds on regardless, bouncing off the scenery as the game plays itself.

People really detest this kind of thing. It’s impossible to pinpoint exactly when big-budget games started assuming that players were irredeemably stupid, but throughout the late 00s and 2010s it became more and more common to see obvious GPS trails showing you where to go, and for characters to talk to themselves to remind you what to do. At the same time, there’s been a cyclical discussion over the difficulty of games: every time a notably challenging game comes out, we rehash the same talking points about whether all games should have easy modes. Is it acceptable for games to shut players out if they aren’t prepared to “git gud”?

I am not a game designer, but as a critic who loves challenging games and has thought a lot about them, for me there is a difference between a game’s difficulty and its friction. A game might be difficult for all kinds of reasons, intentional or otherwise: perhaps the controls or camera aren’t great; perhaps there’s a mid-game boss with a bafflingly massive health bar and one devastating unblockable attack; perhaps the balancing is off. But a game’s friction is always an intentional part of its design. Friction is about how much a game helps you achieve its tasks. Games with perfect friction are clear about showing you what to do, but not overbearing about helping you do it: there is intentional distance, distance that you must broach with agency, strategy and skill. Friction leaves space for frustration and learning, and a sense of achievement.

‘A game is only fun if you have to do something’ … 007 First Light. Photograph: IO Interactive A/S

A lot of modern blockbuster game design seems hellbent on removing this friction. Not sure where to go? Here – follow this brightly coloured trail of fart gas. Not sure how to climb that building? We’ve helpfully painted the ledges yellow. Have you stood still for 30 seconds? We’ll have another character talk in your ear and tell you exactly what you need to do. Playing through the opening hours of 007 First Light recently, I was dismayed when Bond inevitably acquired a watch that magically highlights important things in his vicinity with the press of a button.

At its best, game design like this helps players avoid becoming irrevocably stuck. At its worst, it totally removes agency and challenge – you know, the reasons that many of us play games in the first place. Fundamentally, a game is only fun if you have to do something. I don’t play games for passive entertainment. That’s what binge-watching Netflix is for. If you don’t have to invest anything to reap the rewards, the experience just washes over you.

Games can be difficult and have high friction. In Hollow Knight: Silksong, you are constantly figuring out what to do, where the next boss might be, how to open that tantalisingly blocked path. Achieving those goals is then also very challenging, requiring practice and mastery of whip-fast fighting and precise jumping. Or they can be high-friction, low-difficulty: Blue Prince tells you literally nothing about how to make your way to its mansion’s secret 46th room, but it’s not difficult to walk around and try out ideas. An example of a low-friction, high-difficulty game, meanwhile, would be a combat-focused one like Devil May Cry (or most shooters), where the aim is always “kill all the guys to open the next area” but actually killing all those guys is majorly demanding.

Low-friction, low-difficulty games hold very little interest. Many mobile games fall into this category, where you barely have to do anything save prod the screen to provoke a dopamine-shower of flashy visual effects (I’m looking at you, Monopoly Go). It’s unfair to lump Wolverine in with these almost entirely valueless games – but it is instructive that it has been subjected to such ridicule.

The infamously challenging Dark Souls, a dramatically influential game that turned 15 years old this week, seemed to remind the entire games industry that you can actually trust players: that they won’t drop the controller and storm off as soon as they get skewered by a skeleton. That there’s immense and unique pleasure in participating in a game, unravelling it, conquering it little by little with skill and understanding. I hope that the reaction to Wolverine will serve as another reminder of that fact, for any executives who need to hear it.

What to play

Psychosocial dread … Silent Hill: Townfall. Photograph: Screen Burn

Konami’s horror series is on a roll lately. The Silent Hill 2 remake was very well received (with a few exceptions, including our critic), last year’s feminist, small-town Japanese horror story Silent Hill: f was excellent, and now Silent Hill: Townfall has succeeded in sending the series to a new setting: an island off the Scottish coast. Per our reviewer Lewis Gordon: “This is a taut, smart tale, less about psychosexual terror than psychosocial dread… Townfall seems to understand that some things should remain shrouded in fog. The fate of St Amelia naturally festers in our imagination. How very Silent Hill.”

Available on: PS5, Xbox, PC
Estimated playtime:
12-15 hours

skip past newsletter promotion


What to read

Pikachu in spaaaaaace … a toy Pikachu floats inside the International Space Station. Photograph: S Adenot/ESA/NASA/SWNS
  • Pikachu is now an astronaut, thanks to a collaboration between Pokémon and the European Space Agency as part of the franchise’s seemingly endless series of 30th birthday celebrations. Human astronaut Sophie Adenaut brought a Pikachu plush up to the International Space Station, where it spent 67 days in orbit.

  • According to Sonic Team boss Takashi Iizuka, Sega was once poised to kill off the whole Sonic franchise. It was the success of 2020’s live-action Sonic movie that saved it. “I had to fight within the organization to keep the franchise as long as possible,” he told One More Game.

  • The new Resident Evil movie is out, and it’s going down well. Aftermath describes it as like watching your inept stoner roommate play the games; certified horror-game enjoyer Keith Stuart says it’s one of the best video game adaptations he’s seen.

  • The Tokyo Game Show had to wrap up one day early last week due to an incoming typhoon. Gamesindustry.biz has a slightly depressing rundown of the show, which was heavily dominated by AI.

  • There have been yet more dramatic changes at Microsoft’s Xbox division. The company is all but closing Halo Studios and moving Halo’s development to Activision, which will also now be in charge of storied British studio Rare. Multi-BAFTA-winning studio Ninja Theory is starting consultations on a proposed closure, having failed to find a funding deal to return itself to independence.

What to click

Question Block

Surreally European … Clair Obscur: Expedition 33. Photograph: Kepler Interactive

Reader Mike asks: “I just finished reading your article about the GameCube, which ironically got me thinking about the PS1 & 2, marvellous machines which I remember fondly. The game I got bundled with my PS2 was ‘Shadow of Memories’ which sticks with me to this day, the story was magnificent and also beautifully rendered, with multiple endings. Are there any similar ‘timeywimey’ games set in beautiful European villages, evoking memories of this game, out there today? Either on Android or PS5?”

This question unearthed a bunch of half-remembered fragments from the early days of the PS2 – I definitely played Shadow of Memories as well. It’s a Konami game about a young German travelling through time, trying to solve his own murder. You mentioned games available on PS5, so here are a few that have some of the same elements. The Life is Strange series is set in the US but has a similar time-travel and murder-solving premise; Clair Obscur is a very different kind of game but is as uniquely and surreally European; this year’s Japanese RPG The Adventures of Elliot involves travelling through different eras to prevent a catastrophe. None of these have the same strange adventure-game atmosphere as Shadow of Memories, though – that’s something that felt really specific to the PS2 era. Readers, have you any more suggestions for Mike?

If you have a question for Question Block – or anything else to say about the newsletter – email us at pushingbuttons@theguardian.com.


Source: Technology