Preloader

Technology

Cloudflare/Security-Audit-Skill

security-audit

A coding-agent skill that turns your agent into a security auditor. It orchestrates isolated agents through reconnaissance, coverage-led hunting, candidate validation, structured output, independent record verification, and target-neutral reporting.

This is the skill that seeded Cloudflare’s vulnerability discovery harness, described in Build your own vulnerability harness. The harness grew into a multi-stage, fleet-wide system; this skill is the single-repo starting point it evolved from.

What it does

The skill runs a structured audit in six phases:

  1. Reconnaissance — map architecture, trust boundaries, input surfaces, prior evidence, and deterministic coverage in architecture.md and coverage-ledger.json.
  2. Coverage-led hunting — assign isolated hunters from ledger units, record their checks, and use coverage critics to find gaps.
  3. Candidate validation — give every unique candidate to a fresh verifier that tries to disprove it.
  4. Structured output — write confirmed, needs_validation, and rejected records to findings.json and validate them against report-schema.json.
  5. Independent record verification — fresh agents verify final source claims. Material replacements receive another independent verifier.
  6. Target-neutral reporting — derive REPORT.md, FINDINGS-DETAIL.md, and NEEDS-VALIDATION.md from the verified records and coverage ledger.

The parent runs validate-coverage-ledger.cjs after creating the ledger and after each later ledger update. It runs validate-findings.cjs in Phase 4 and again after every Phase 5 replacement.

The verdicts are distinct: confirmed has a complete source trace and bounded observed result, needs_validation has an exact unresolved fact and no severity, and rejected records a disproved candidate.

Multiple runs against the same repo are additive. The skill uses prior ledgers and findings to target gaps, revalidate changed source, and carry forward current-source evidence without treating stale or unresolved work as covered.

Files

File Purpose
SKILL.md Setup, core principles, platform terminology, workflow overview, and audit anti-patterns
RECONNAISSANCE.md Phase 1 reconnaissance prompts and synthesis instructions
HUNTING.md Phase 2 orchestration, hunting methodology, and validation rules
ATTACK-CLASSES.md Core, wildcard, and obvious-things attack prompts
MEMORY-SAFETY-AND-BINARY.md Memory-safety, binary, and kernel hunting classes for native targets
AI-AND-LLM.md Prompt-injection, agent/tool, and output-handling hunting classes for LLM-backed targets
WEB-PROTOCOL-AND-AUTH.md HTTP request-framing, cache, and authentication-protocol hunting classes for HTTP-protocol and auth targets
CLIENT-SIDE.md DOM-injection, messaging-trust, UI-redress, and prototype-pollution hunting classes for client-side/browser targets
SUPPLY-CHAIN-AND-RELEASE.md Dependency, CI, release, signing, update, plugin, and extension hunting classes
CLOUD-AND-DEPLOYMENT.md IAM, infrastructure-as-code, container, serverless, ingress, and runtime-configuration hunting classes
PROTOCOLS-RPC-AND-MESSAGING.md RPC, serialization, queue, broker, webhook, and streaming-protocol hunting classes
RESOURCE-EXHAUSTION-AND-AVAILABILITY.md Shared resource, quota, queue, worker, and operator-spend hunting classes
DATA-ISOLATION-AND-LIFECYCLE.md Tenant isolation, cache, search, export, backup, migration, deletion, and restore hunting classes
DESKTOP-MOBILE-AND-LOCAL-IPC.md Native app, deep-link, webview, exported-component, helper, daemon, and local-IPC hunting classes
VALIDATION-AND-REPORTING.md Phases 3–6 candidate validation, structured output, record verification, and reporting
report-schema.json JSON schema for all three findings.json verdicts
validate-findings.cjs Zero-dependency validator for findings.json in Phases 4 and 5
validate-findings.test.cjs Findings-validator tests and producer-compatible fixture checks
validate-coverage-ledger.cjs Zero-dependency validator for coverage-ledger.json in Phases 1–5
validate-coverage-ledger.test.cjs Coverage-ledger validator tests

Installation

Install the skill with the Skills CLI:

npx skills add https://github.com/cloudflare/security-audit-skill 
  --skill security-audit

Use --global for a user-level installation:

npx skills add https://github.com/cloudflare/security-audit-skill 
  --skill security-audit 
  --global

Run npx skills --help for agent-selection and non-interactive options.

Usage

Start your coding agent in (or pointed at) the codebase you want to audit, then ask it to do a security audit:

security audit this codebase
find security vulnerabilities in ./src
do a security review, output to ~/audits/my-project

The skill activates automatically when the request matches its trigger (security audit, find vulnerabilities, pen-test the code, etc.). A direct codebase audit or pen-test request uses full audit mode. Security questions and focused vulnerability work use guidance mode unless you request report artifacts. In full audit mode, an unspecified output directory defaults to ~/security-audit-skill/<repo-name>/run-<N>. The workflow writes inside the target repository only when you explicitly select a directory that version control ignores.

Requirements

  • A coding agent with a model that supports tool use and parallel sub-agents
  • Node.js for the zero-dependency findings and coverage-ledger validators
  • An OS-enforced sandbox for target-controlled builds, tests, processes, browsers, emulators, fuzzers, and fixtures. It must disable external networking, use a sanitized allowlisted environment, enforce resource limits, and allow writes only to assigned scratch paths. Without these controls, the workflow keeps the lead as needs_validation instead of executing target code.

Design principles

  • Only confirm established boundary failures. Keep a source-grounded blocked lead as needs_validation with its exact unresolved fact.
  • Adversarial validation. The agent that checks a finding is never the agent that found it.
  • Severity requires impact. Likelihood x impact, not deviation from a checklist.
  • Defense-in-depth gaps are not vulnerabilities. If Layer A prevents the attack, the absence of Layer B is a hardening note.
  • Multiple runs improve coverage. In our test runs, a single run found roughly half of the vulnerabilities that repeated runs found in total.

Contact

Questions, feedback, or comparing notes on AI-driven security tooling: security-ai-research@cloudflare.com

License

MIT — see LICENSE.


Source: Hacker News

CrowdSec Source Code Leak

On September 16, CrowdSec was informed of a source code leak involving our GitHub repository, which occurred in May 2026. Our team verified and confirmed the report. CrowdSec source code consists of two parts: a private one and another that hosts our Free Open Source Software (i.e., the Security Engine), which is public by design and therefore out of scope. The private part, though, contains the source code for our SaaS console, some AWS Cloud routines, some connectors, and automations. 

The news headline claiming 300 different repositories is accurate (when you include the 130+ public ones), though that number mostly reflects the code’s subdivision rather than a specific volume. We do not confirm any “other file contained” or “internal development material”, since all the code is published in these repositories. The API related information is the token used by the CI/CD component itself. (see below)

No client data, login/password, name, organization, or anything else was leaked, and CrowdSec doesn’t store PII or client logs; the impact is limited to CrowdSec. Our team quickly hunted for any token, credential, or sensitive leak that could enable lateral movement but found none so far.

The code contained in these private repositories has value but cannot really harm CrowdSec, since our efficiency depends on our network effect and size, which code alone can’t replicate. We regularly audited the SaaS source code, and its leakage shouldn’t pose an immediate threat either. Most of the leaked code has evolved significantly over those four months, but we will closely monitor for any abnormal activity. Also, using it outside of CrowdSec seems unlikely because it only interacts with our data and tools and cannot really be leveraged in another context. 

We will keep you updated as we continue investigating, but the Tanstack compromise is very likely to have been the leak vector (more about it here), as in the case of the Mistral AI case. This component was used in our organization in May and appears to have been backdoored to extract an API key with authorization to read the private codebase. The leak was only exploitable during a short timeframe in May 2026.

We nevertheless immediately rotated all required tokens & credentials to prevent further incidents.

The team would like to thank Fuites Infos for their timely, professional outreach in reporting the issue.


Source: Hacker News

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

We were a little impatient.

Qwen 3.8 27B was released on August 14, 2026. Four days later, on August 18, we published our first set of GGUFs. We called them ShapeLearn-Lite for a reason: they were produced using a much smaller optimization budget, fewer checks, and much less waiting.

Now the full ShapeLearn models are done, and we have benchmarked them alongside the original Lite set and competing quants.

The good news: ShapeLearn-Lite held up pretty well. We will come back to that later in “ShapeLearn-Lite, in retrospect”.

The better news: the full ShapeLearn models are even better.


Quick start with llama.cpp

The MTP draft head is bundled in every GGUF. DFlash2 uses a separate 1.1 GB draft model. Both commands use GPU-5. Swap the tag for any other model in the release.

MTP Embedded draft. Works with image inputs.

llama-server 
  -hf byteshape/Qwen3.8-27B-GGUF:Qwen3.8-27B-IQ4_XS-3.84bpw 
  --mmproj-auto 
  --spec-type draft-mtp --spec-draft-n-max 3

DFlash2 External draft. Fastest option, text only.

llama-server 
  -hf byteshape/Qwen3.8-27B-GGUF:Qwen3.8-27B-IQ4_XS-3.84bpw 
  -hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M 
  --spec-type draft-dflash --spec-draft-n-max 7 
  --no-mmproj

DFlash2 needs llama.cpp b10658 or newer. Ready-to-run commands for every model, with the recommended sampling settings, are in the run tool and on the model card.

TL;DR

  • Full ShapeLearn moves the measured quality-speed frontier beyond Lite. All five models in the new release sit on the frontier in each of our six GPU comparisons.
  • GPU-5 is our default recommendation wherever it fits, reaching 99.63% of BF16’s aggregate benchmark score. If it does not fit with the context you need, GPU-4 is still very competitive: it reaches 98.72% of BF16 at a much smaller size (11.0 GB instead of 13.1 GB), and it is faster.
  • ShapeLearn-Lite also performed better than its KLD ranking suggested: three of its six models sit on the frontier in the Lite-versus-Unsloth Dynamic v3 comparison.
  • Speculative Decoding with MTP or DFlash2 increases throughput across every ShapeLearn model and GPU tested. DFlash2 is usually faster but requires more memory and does not support image inputs with llama.cpp. Choose DFlash2 for maximum text-only throughput when memory allows, and MTP when VRAM or multimodal support matters more.

Full ShapeLearn moves the frontier

We are releasing the full ShapeLearn run for Qwen 3.8 27B.

Within this release, larger models yield higher aggregate scores, while smaller models deliver higher throughput. That ordering holds across all six GPUs tested. Because this is a dense model and memory transfers are the bottleneck, lower BPW translates more directly into higher TPS than it does for MoEs.

The per-GPU comparisons also include AtomicChat, Bartowski, ISTA-DASLab, and Unsloth Dynamic v3. Bartowski’s latest models were released after our testing and are not included. Full ShapeLearn is labelled ByteShape in the figures.

All five ShapeLearn models remain on the measured frontier, with GPU-5 achieving the highest aggregate score among the plotted quants. Other teams also contribute competitive points. Notably, ISTA-DASLab’s excellent model (the yellow “d” on the graph below) also sits on the frontier.

By “frontier,” we mean that no other plotted model is both faster and more accurate.

96 GB: RTX Pro 6000

RTX PRO 6000, the GPU with the most memory, lets us show the full range of models tested.

RTX Pro 6000: tokens per second vs quality (full ShapeLearn and competing quants)
RTX Pro 6000: tokens per second vs quality (full ShapeLearn and competing quants)
Tap Show Legend below for model details.

RTX Pro 6000: tokens per second vs quality (full ShapeLearn and competing quants)
Hover over the bubbles, or click Show Legend below, for model details.
Show Legend
# Model Acc TPS BPW
GPU-1 IQ2_XXS-2.56bpw 0.9304 116.11 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 108.01 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 105.44 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 101.11 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 90.42 3.84
A UD-IQ2_S 0.8633 114.81 2.49
B UD-Q2_K_XL 0.9572 106.95 2.82
C UD-IQ3_XXS 0.9359 100.51 3.14
D UD-IQ3_S 0.9555 95.53 3.47
E UD-Q3_K_XL 0.9760 90.88 3.80
F UD-IQ4_XS 0.9920 86.73 4.13
G UD-Q4_K_S 0.9877 82.21 4.46
H UD-Q4_K_M 0.9703 78.58 4.79
I UD-Q4_K_XL 0.9871 74.75 5.12
J UD-Q5_K_S 0.9878 71.00 5.44
K UD-Q5_K_M 0.9897 67.71 5.77
L UD-Q5_K_XL 0.9905 65.52 6.10
a GSQ-RCO-IQ2_XS 0.8647 112.77 2.50
b GSQ-RCO-IQ2_S 0.9364 108.22 2.75
c GSQ-RCO-IQ3_XXS 0.9438 103.39 3.00
d GSQ-RCO-IQ3_S 0.9943 94.77 3.50
a IQ2_XXS 0.7986 112.21 2.72
b IQ2_S 0.9296 106.33 2.99
c Q2_K 0.9616 96.08 3.45
d IQ3_XXS 0.9594 92.43 3.68
e IQ3_XS 0.9582 88.36 3.89
f IQ3_M 0.9667 86.17 4.06
a AD-IQ2_XXS 0.8385 116.28 2.58
b AD-IQ2_XS 0.9296 108.67 2.85
c AD-IQ2_S 0.9061 100.28 3.22
d AD-IQ3_XXS 0.9644 95.29 3.50
e AD-IQ3_S 0.9730 88.52 4.04

GPU-5 is our default wherever you can fit it: it reaches 90.4 tok/s at 99.63% of the BF16 baseline.

32 GB: RTX 5090

The RTX 5090 tells a similar story, leading to the same recommendations.

RTX 5090: tokens per second vs quality (full ShapeLearn and competing quants)
RTX 5090: tokens per second vs quality (full ShapeLearn and competing quants)
Tap Show Legend below for model details.

RTX 5090: tokens per second vs quality (full ShapeLearn and competing quants)
Hover over the bubbles, or click Show Legend below, for model details.
Show Legend
# Model Acc TPS BPW
GPU-1 IQ2_XXS-2.56bpw 0.9304 119.13 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 110.78 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 108.08 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 103.57 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 93.66 3.84
A UD-IQ2_S 0.8633 115.34 2.49
B UD-Q2_K_XL 0.9572 108.12 2.82
C UD-IQ3_XXS 0.9359 102.87 3.14
D UD-IQ3_S 0.9555 97.87 3.47
E UD-Q3_K_XL 0.9760 93.46 3.80
F UD-IQ4_XS 0.9920 89.69 4.13
G UD-Q4_K_S 0.9877 85.28 4.46
H UD-Q4_K_M 0.9703 81.63 4.79
I UD-Q4_K_XL 0.9871 77.51 5.12
J UD-Q5_K_S 0.9878 73.60 5.44
K UD-Q5_K_M 0.9897 69.87 5.77
L UD-Q5_K_XL 0.9905 67.59 6.10
a GSQ-RCO-IQ2_XS 0.8647 114.11 2.50
b GSQ-RCO-IQ2_S 0.9364 110.19 2.75
c GSQ-RCO-IQ3_XXS 0.9438 105.67 3.00
d GSQ-RCO-IQ3_S 0.9943 95.49 3.50
a IQ2_XXS 0.7986 115.77 2.72
b IQ2_S 0.9296 109.39 2.99
c Q2_K 0.9616 99.29 3.45
d IQ3_XXS 0.9594 95.83 3.68
e IQ3_XS 0.9582 91.14 3.89
f IQ3_M 0.9667 89.08 4.06
a AD-IQ2_XXS 0.8385 119.53 2.58
b AD-IQ2_XS 0.9296 112.39 2.85
c AD-IQ2_S 0.9061 103.46 3.22
d AD-IQ3_XXS 0.9644 98.22 3.50
e AD-IQ3_S 0.9730 91.72 4.04

Once again GPU-5 is our default choice, reaching 93.7 tok/s. Choose GPU-4 for slightly more context length or slightly better TPS.

24 GB: RTX 4090 and RTX 3090

Both 24 GB cards fit all five ShapeLearn models. We plot them separately because their throughput differs, but the ordering is the same on both.

RTX 4090

The RTX 4090 keeps the same pattern: GPU-5 is the default, reaching 59.2 tok/s.

RTX 4090: tokens per second vs quality (full ShapeLearn and competing quants)
RTX 4090: tokens per second vs quality (full ShapeLearn and competing quants)
Tap Show Legend below for model details.

RTX 4090: tokens per second vs quality (full ShapeLearn and competing quants)
Hover over the bubbles, or click Show Legend below, for model details.
Show Legend
# Model Acc TPS BPW
GPU-1 IQ2_XXS-2.56bpw 0.9304 78.67 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 72.89 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 71.12 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 67.77 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 59.16 3.84
A UD-IQ2_S 0.8633 78.23 2.49
B UD-Q2_K_XL 0.9572 72.79 2.82
C UD-IQ3_XXS 0.9359 67.24 3.14
D UD-IQ3_S 0.9555 63.02 3.47
E UD-Q3_K_XL 0.9760 59.13 3.80
F UD-IQ4_XS 0.9920 55.62 4.13
G UD-Q4_K_S 0.9877 52.44 4.46
H UD-Q4_K_M 0.9703 49.77 4.79
I UD-Q4_K_XL 0.9871 47.08 5.12
J UD-Q5_K_S 0.9878 44.80 5.44
K UD-Q5_K_M 0.9897 42.45 5.77
L UD-Q5_K_XL 0.9905 40.84 6.10
a GSQ-RCO-IQ2_XS 0.8647 77.80 2.50
b GSQ-RCO-IQ2_S 0.9364 73.34 2.75
c GSQ-RCO-IQ3_XXS 0.9438 69.26 3.00
d GSQ-RCO-IQ3_S 0.9943 62.41 3.50
a IQ2_XXS 0.7986 75.28 2.72
b IQ2_S 0.9296 71.27 2.99
c Q2_K 0.9616 63.74 3.45
d IQ3_XXS 0.9594 61.03 3.68
e IQ3_XS 0.9582 58.20 3.89
f IQ3_M 0.9667 56.16 4.06
a AD-IQ2_XXS 0.8385 79.51 2.58
b AD-IQ2_XS 0.9296 74.19 2.85
c AD-IQ2_S 0.9061 67.40 3.22
d AD-IQ3_XXS 0.9644 63.60 3.50
e AD-IQ3_S 0.9730 57.38 4.04

RTX 3090

Older, but still fast in these measurements.

RTX 3090: tokens per second vs quality (full ShapeLearn and competing quants)
RTX 3090: tokens per second vs quality (full ShapeLearn and competing quants)
Tap Show Legend below for model details.

RTX 3090: tokens per second vs quality (full ShapeLearn and competing quants)
Hover over the bubbles, or click Show Legend below, for model details.
Show Legend
# Model Acc TPS BPW
GPU-1 IQ2_XXS-2.56bpw 0.9304 53.03 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 51.20 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 50.75 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 49.49 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 45.79 3.84
A UD-IQ2_S 0.8633 52.82 2.49
B UD-Q2_K_XL 0.9572 50.37 2.82
C UD-IQ3_XXS 0.9359 48.16 3.14
D UD-IQ3_S 0.9555 46.39 3.47
E UD-Q3_K_XL 0.9760 46.28 3.80
F UD-IQ4_XS 0.9920 46.12 4.13
G UD-Q4_K_S 0.9877 44.49 4.46
H UD-Q4_K_M 0.9703 43.30 4.79
I UD-Q4_K_XL 0.9871 41.37 5.12
J UD-Q5_K_S 0.9878 39.37 5.44
K UD-Q5_K_M 0.9897 37.38 5.77
L UD-Q5_K_XL 0.9905 36.16 6.10
a GSQ-RCO-IQ2_XS 0.8647 51.56 2.50
b GSQ-RCO-IQ2_S 0.9364 50.07 2.75
c GSQ-RCO-IQ3_XXS 0.9438 48.66 3.00
d GSQ-RCO-IQ3_S 0.9943 47.19 3.50
a IQ2_XXS 0.7986 53.63 2.72
b IQ2_S 0.9296 51.21 2.99
c Q2_K 0.9616 45.34 3.45
d IQ3_XXS 0.9594 47.46 3.68
e IQ3_XS 0.9582 44.68 3.89
f IQ3_M 0.9667 43.42 4.06
a AD-IQ2_XXS 0.8385 53.95 2.58
b AD-IQ2_XS 0.9296 51.91 2.85
c AD-IQ2_S 0.9061 48.24 3.22
d AD-IQ3_XXS 0.9644 47.10 3.50
e AD-IQ3_S 0.9730 47.22 4.04

GPU-4 reaches 49.5 tok/s, compared with 45.8 tok/s for GPU-5. Moving to the larger model costs about 7.5% in throughput, while the aggregate score rises from 98.72% to 99.63% of BF16. That makes GPU-5 the default here as well.

16 GB: RTX 4080 and RTX 5060 Ti

With a tighter VRAM budget, the pragmatic choice is to leave room for the context you need, not just the model weights. These plots contain fewer competing configurations, but all five ShapeLearn models are represented.

RTX 4080

On the RTX 4080, GPU-4 reaches 52.4 tok/s, while GPU-5 reaches 45.7 tok/s.

RTX 4080: tokens per second vs quality (full ShapeLearn and competing quants)
RTX 4080: tokens per second vs quality (full ShapeLearn and competing quants)
Tap Show Legend below for model details.

RTX 4080: tokens per second vs quality (full ShapeLearn and competing quants)
Hover over the bubbles, or click Show Legend below, for model details.
Show Legend
# Model Acc TPS BPW
GPU-1 IQ2_XXS-2.56bpw 0.9304 62.01 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 56.84 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 55.14 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 52.43 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 45.74 3.84
A UD-IQ2_S 0.8633 61.70 2.49
B UD-Q2_K_XL 0.9572 56.70 2.82
C UD-IQ3_XXS 0.9359 52.32 3.14
D UD-IQ3_S 0.9555 49.06 3.47
a GSQ-RCO-IQ2_XS 0.8647 60.47 2.50
b GSQ-RCO-IQ2_S 0.9364 57.37 2.75
c GSQ-RCO-IQ3_XXS 0.9438 54.19 3.00
d GSQ-RCO-IQ3_S 0.9943 48.42 3.50
a IQ2_XXS 0.7986 58.69 2.72
b IQ2_S 0.9296 55.39 2.99
c Q2_K 0.9616 48.91 3.45
d IQ3_XXS 0.9594 46.99 3.68
e IQ3_XS 0.9582 44.71 3.89
a AD-IQ2_XXS 0.8385 62.55 2.58
b AD-IQ2_XS 0.9296 57.82 2.85
c AD-IQ2_S 0.9061 52.43 3.22
d AD-IQ3_XXS 0.9644 49.15 3.50

RTX 5060 Ti

On the RTX 5060 Ti, the corresponding figures are 33.1 tok/s and 29.1 tok/s.

RTX 5060 Ti: tokens per second vs quality (full ShapeLearn and competing quants)
RTX 5060 Ti: tokens per second vs quality (full ShapeLearn and competing quants)
Tap Show Legend below for model details.

RTX 5060 Ti: tokens per second vs quality (full ShapeLearn and competing quants)
Hover over the bubbles, or click Show Legend below, for model details.
Show Legend
# Model Acc TPS BPW
GPU-1 IQ2_XXS-2.56bpw 0.9304 38.05 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 35.41 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 34.50 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 33.05 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 29.15 3.84
A UD-IQ2_S 0.8633 38.02 2.49
B UD-Q2_K_XL 0.9572 35.42 2.82
C UD-IQ3_XXS 0.9359 32.80 3.14
D UD-IQ3_S 0.9555 30.97 3.47
a GSQ-RCO-IQ2_XS 0.8647 37.62 2.50
b GSQ-RCO-IQ2_S 0.9364 35.80 2.75
c GSQ-RCO-IQ3_XXS 0.9438 33.95 3.00
d GSQ-RCO-IQ3_S 0.9943 30.72 3.50
a IQ2_XXS 0.7986 36.33 2.72
b IQ2_S 0.9296 34.76 2.99
c Q2_K 0.9616 30.79 3.45
d IQ3_XXS 0.9594 29.69 3.68
e IQ3_XS 0.9582 28.24 3.89
a AD-IQ2_XXS 0.8385 38.42 2.58
b AD-IQ2_XS 0.9296 36.05 2.85
c AD-IQ2_S 0.9061 32.61 3.22
d AD-IQ3_XXS 0.9644 31.01 3.50

GPU-5 remains the default on both cards when the model, KV cache, and runtime buffers fit within your memory budget. When they do not, GPU-4 is still very competitive: almost 99% of BF16 at a much smaller size, and faster. A model appearing in these measurements does not establish that every context length or serving configuration will fit.

ShapeLearn-Lite, in retrospect

ShapeLearn-Lite uses a smaller optimization budget than full ShapeLearn. It let us get Qwen 3.8 27B onto 12 GB to 24 GB GPUs within a few days.

We released after targeted sanity checks and started the full evaluation afterwards. The full ShapeLearn models were ready before the benchmarking was finished. Evaluating both sets, along with the competing models, is what took most of the time.

Then Unsloth released its Dynamic v3 models. At similar sizes, several had lower KLD than Lite in our measurements. On KLD alone, Lite looked less competitive.

KLD looked decisive

KLD measures divergence between a quantized model’s predicted token distributions and the BF16 reference under a particular evaluation setup. It is useful for diagnosing substantial changes, but lower divergence does not automatically mean better task performance.

We measure KLD on a dataset of about 5 million tokens of prompt and response pairs, drawn from several benchmarks, including long-context and agentic tasks. We also changed how KLD is computed, so that it is closer to what we expect KLD to measure:

  • KLD is measured on response tokens only, not on prompt tokens. We do not want to measure how well a model can generate prompts.
  • KLD only considers the tokens that have a chance of being sampled during generation, the top-20, top-40, or top-60 tokens at each position. The tail tokens never get sampled, so they do not contribute.
  • Requests have clear boundaries. Each prompt and response pair is scored as its own request, not as part of one long concatenated stream.

KL divergence versus model size for ShapeLearn-Lite and Unsloth
KL divergence versus model size for ShapeLearn-Lite and Unsloth
Tap Show Legend below for model details.

KL divergence versus model size for ShapeLearn-Lite and Unsloth
Hover over the bubbles, or click Show Legend below, for model details.
Show Legend
# Model KLD Size (GB) BPW
Lite-1 IQ3_S-3.44bpw 0.035875 10.79 3.44
Lite-2 IQ4_XS-3.67bpw 0.028296 11.51 3.68
Lite-3 IQ4_XS-4.00bpw 0.018249 12.52 4.00
Lite-4 IQ4_XS-4.40bpw 0.009901 13.78 4.40
Lite-5 Q5_K_S-4.72bpw 0.007578 14.78 4.72
Lite-6 Q5_K_M-5.60bpw 0.003297 17.53 5.60
i UD-IQ1_S 0.389550 5.76 1.84
ii UD-IQ1_M 0.261876 6.26 2.00
iii UD-IQ2_XXS 0.181493 6.76 2.16
iv UD-IQ2_S 0.108374 7.79 2.49
v UD-Q2_K_XL 0.065200 8.81 2.81
vi UD-IQ3_XXS 0.040407 9.84 3.14
vii UD-IQ3_S 0.028759 10.87 3.47
viii UD-Q3_K_XL 0.019844 11.90 3.80
ix UD-IQ4_XS 0.011992 12.93 4.13
x UD-Q4_K_S 0.008827 13.96 4.46
xi Q4_0 0.019264 14.69 4.69
xii UD-Q4_K_M 0.007054 14.99 4.79
xiii UD-Q4_K_XL 0.005210 16.01 5.11
xiv Q4_1 0.009603 16.06 5.13
xv UD-Q5_K_S 0.003619 17.04 5.44
xvi UD-Q5_K_M 0.002839 18.07 5.77
xvii UD-Q5_K_XL 0.002432 19.10 6.10
xviii UD-Q6_K 0.001771 20.13 6.43
xix UD-Q6_K_M 0.001426 21.16 6.76
xx UD-Q6_K_L 0.001150 22.19 7.09
xxi UD-Q6_K_XL 0.000985 23.22 7.42
xxii UD-Q8_K_L 0.000726 25.78 8.23
xxiii Q8_0 0.000648 26.62 8.50
xxiv UD-Q8_K_XL 0.000503 28.76 9.19

For example, Unsloth’s UD-IQ3_S (vii) has about 20% lower KLD than the similarly sized smallest Lite model (Lite-1): 0.028759 versus 0.035875. Yet its aggregate benchmark score is lower: 95.55% versus 97.33% of BF16.

If lower KLD were sufficient to rank these models by task performance, the benchmark ordering should have followed it.

It did not.

The point is not that KLD is useless. It is that a fidelity ranking is not a task-performance ranking. This is the distinction explored in our KLD evaluation blog. Our related paper on KLD and quantization fidelity metrics was also recently accepted to the EMNLP Industry Track.

Lite held up

Naturally, we made more plots.

Here, we show the RTX Pro 6000 because it can accommodate the full comparison. Each model’s benchmark score is reused across the GPU plots; the measured throughput and the set of displayed models change.

RTX Pro 6000: ShapeLearn, ShapeLearn-Lite, and Unsloth Dynamic v3
RTX Pro 6000: ShapeLearn, ShapeLearn-Lite, and Unsloth Dynamic v3
Tap Show Legend below for model details.

RTX Pro 6000: ShapeLearn, ShapeLearn-Lite, and Unsloth Dynamic v3
Hover over the bubbles, or click Show Legend below, for model details.
Show Legend
# Model Acc TPS BPW
GPU-1 IQ2_XXS-2.56bpw 0.9304 116.11 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 108.01 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 105.44 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 101.11 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 90.42 3.84
Lite-1 IQ3_S-3.44bpw 0.9733 98.74 3.45
Lite-2 IQ4_XS-3.67bpw 0.9802 94.48 3.68
Lite-3 IQ4_XS-4.00bpw 0.9880 89.85 4.00
Lite-4 IQ4_XS-4.40bpw 0.9856 85.03 4.40
Lite-5 Q5_K_S-4.72bpw 0.9909 79.75 4.72
Lite-6 Q5_K_M-5.60bpw 0.9919 70.28 5.60
A UD-IQ2_S 0.8633 114.81 2.49
B UD-Q2_K_XL 0.9572 106.95 2.82
C UD-IQ3_XXS 0.9359 100.51 3.14
D UD-IQ3_S 0.9555 95.53 3.47
E UD-Q3_K_XL 0.9760 90.88 3.80
F UD-IQ4_XS 0.9920 86.73 4.13
G UD-Q4_K_S 0.9877 82.21 4.46
H UD-Q4_K_M 0.9703 78.58 4.79
I UD-Q4_K_XL 0.9871 74.75 5.12
J UD-Q5_K_S 0.9878 71.00 5.44
K UD-Q5_K_M 0.9897 67.71 5.77
L UD-Q5_K_XL 0.9905 65.52 6.10

Leaving the full ShapeLearn models aside for a moment, three of the six ShapeLearn-Lite models sit on the Lite-versus-Unsloth frontier: the three smallest Lite models, the lighter orange bubbles labelled 1-3.

Of the twelve Unsloth v3 models shown, three also sit on that frontier: UD-IQ2_S (A), UD-Q2_K_XL (B), and UD-IQ4_XS (F). UD-IQ4_XS (F) is a strong higher-quality point, while Lite earns its places in the middle of the range.

Add the five full ShapeLearn models back in (the darker orange bubbles), and they take over the entire frontier.

Lite was never meant to be the final result. It still held its own where it mattered.

Speculative Decoding

We also evaluated MTP and DFlash2 with the new models, using 3 draft tokens for MTP and 7 draft tokens for DFlash2. Both methods increased throughput for all five ShapeLearn models on all six GPUs tested.

DFlash2 was faster than MTP in almost all cases. Across the full lineup, DFlash2 reached 1.34-2.10x the baseline next-token prediction (NTP) throughput, while MTP reached 1.28-1.66x.

We measured with the sampling parameters Qwen recommends for thinking mode, over a diverse set of agentic coding, mathematics, and general-knowledge requests. The speedups would likely be larger under greedy decoding, but temperature-based sampling better reflects real usage.

The figure below shows NTP, MTP, and DFlash2 throughput for each GPU. The quality axis is the target-model benchmark score reported above. These plots do not independently establish quality equivalence between decoding methods.

Tokens per second vs quality for NTP, MTP and DFlash2 on all six GPUs
Tokens per second vs quality (NTP vs MTP vs DFlash2), one panel per GPU. MTP uses 3 draft tokens, DFlash2 uses 7.
Tap Show Legend below for model details.

Tokens per second vs quality (NTP vs MTP vs DFlash2), one panel per GPU. MTP uses 3 draft tokens, DFlash2 uses 7.
Hover over the bubbles, or click Show Legend below, for model details.
Show Legend
# Model Acc NTP TPS MTP TPS DFlash2 TPS BPW
GPU-1 IQ2_XXS-2.56bpw 0.9304 116.11 165.52 (1.43x) 172.01 (1.48x) 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 108.01 153.48 (1.42x) 165.94 (1.54x) 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 105.44 152.19 (1.44x) 166.03 (1.57x) 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 101.11 146.85 (1.45x) 164.24 (1.62x) 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 90.42 145.69 (1.61x) 150.83 (1.67x) 3.84
GPU-1 IQ2_XXS-2.56bpw 0.9304 119.13 164.44 (1.38x) 175.53 (1.47x) 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 110.78 156.63 (1.41x) 175.97 (1.59x) 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 108.08 155.19 (1.44x) 171.97 (1.59x) 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 103.57 148.34 (1.43x) 169.69 (1.64x) 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 93.66 147.00 (1.57x) 166.57 (1.78x) 3.84
GPU-1 IQ2_XXS-2.56bpw 0.9304 78.67 111.73 (1.42x) 135.39 (1.72x) 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 72.89 108.40 (1.49x) 132.20 (1.81x) 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 71.12 107.21 (1.51x) 133.38 (1.88x) 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 67.77 101.47 (1.50x) 131.12 (1.93x) 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 59.16 98.25 (1.66x) 124.15 (2.10x) 3.84
GPU-1 IQ2_XXS-2.56bpw 0.9304 53.03 68.19 (1.29x) 70.94 (1.34x) 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 51.20 65.76 (1.28x) 68.90 (1.35x) 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 50.75 65.82 (1.30x) 68.40 (1.35x) 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 49.49 64.50 (1.30x) 66.34 (1.34x) 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 45.79 66.55 (1.45x) 63.93 (1.40x) 3.84
GPU-1 IQ2_XXS-2.56bpw 0.9304 62.01 87.06 (1.40x) 101.63 (1.64x) 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 56.84 81.25 (1.43x) 96.90 (1.70x) 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 55.14 79.74 (1.45x) 96.95 (1.76x) 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 52.43 76.88 (1.47x) 94.09 (1.79x) 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 45.74 74.01 (1.62x) 87.49 (1.91x) 3.84
GPU-1 IQ2_XXS-2.56bpw 0.9304 38.05 50.54 (1.33x) 56.71 (1.49x) 2.56
GPU-2 IQ3_XXS-2.88bpw 0.9656 35.41 47.94 (1.35x) 52.57 (1.48x) 2.88
GPU-3 IQ3_XS-3.01bpw 0.9726 34.50 47.93 (1.39x) 52.28 (1.52x) 3.01
GPU-4 IQ3_S-3.23bpw 0.9872 33.05 46.25 (1.40x) 50.34 (1.52x) 3.23
GPU-5 IQ4_XS-3.84bpw 0.9963 29.15 45.82 (1.57x) 47.01 (1.61x) 3.84

There is also a memory tradeoff between the two approaches. The embedded quantized MTP weights add only about 250 MB to the model, and if MTP is not used, these weights are not loaded into GPU memory. In comparison, the 4-bit DFlash2 draft model is about 1.1 GB, so enabling DFlash2 requires roughly 1.1 GB of additional GPU memory.

Packaging MTP as a separate GGUF file would largely eliminate this advantage. The standalone model would need its own MTP embedding and output layers, which are by far its largest tensors, bringing its memory footprint to roughly 1 GB as well.

In addition, DFlash2 in llama.cpp currently does not support image inputs, which is an important consideration for multimodal use cases.

Benchmarking Methodology

We evaluate all reported models across a set of instruct and thinking benchmarks.

Instruct benchmarks:

  • GSM8K for math
  • IFEval for instruction following
  • MMLU for general knowledge
  • LiveCodeBench V6* for coding
  • Multi-IF for multi-turn and multilingual instruction following
  • ACEBench for tool use and agentic tasks

Thinking benchmarks:

  • ACEBench for tool use and agentic tasks
  • Multiple HumanEval for coding
  • BFCL V4* for tool calling and agentic tasks

For the thinking benchmarks, we used Qwen 3.8’s medium thinking setting.

For each benchmark, the score of a quantized model is normalized by the score of the corresponding BF16 model. The overall reported score is the average of these normalized benchmark scores.

Our LiveCodeBench V6* evaluation includes problems from January 1, 2024 onward, excluding the 2023 problems. We found the 2023 problems to be relatively easy for current models, with most models achieving very high scores on them. As a result, they provide limited discrimination between models while adding substantial evaluation time.

For BFCL V4*, we evaluate the following eight subsets:

  • live_simple
  • live_parallel
  • live_parallel_multiple
  • live_relevance
  • multi_turn_base
  • multi_turn_miss_func
  • multi_turn_miss_param
  • multi_turn_long_context

All evaluations were run with llama.cpp b10430. For both instruct and thinking experiments, we use the sampling parameters recommended by Qwen for the corresponding mode.

Conclusion

ShapeLearn-Lite did what it was designed to do. It got useful Qwen 3.8 27B quants onto 12 to 24 GB GPUs quickly, and it held up better than its KLD ranking suggested.

Full ShapeLearn goes further. It improves the measured quality-speed trade-offs over Lite and contributes five frontier models across all six tested GPUs.

GPU-5 is our default recommendation wherever it fits, reaching 99.63% of BF16’s aggregate benchmark score. When memory is tight, GPU-4 is still very competitive: almost 99% of BF16 at a much smaller size, and faster.

KLD remains useful, but it is not a task-performance leaderboard. Fidelity metrics tell us how much the model’s distributions changed under a particular measurement. Benchmarks tell us whether those changes matter on the tasks we tested.

We were impatient. This time, it worked out pretty well.


Source: Hacker News

Rate limits on GitLab.com are changing

Published on: September 17, 2026

Starting October 19, GitLab.com rate limits will align with your subscription. Sign in to unlock higher limits. Premium/Ultimate changes arrive in January.

GitLab.com hosts millions of projects for teams of every size that need a platform they can rely on. Demand is climbing quickly, and we expect platform load to grow several times over this year. Predictable limits are what keep GitLab.com fast for everyone on it, including the automation and agent workloads teams are building on the platform.

To hold that as we scale, we're updating how rate limits work. Starting October 19, 2026, rate limits on GitLab.com will align with your subscription tier. Free accounts and unauthenticated requests happen first, on October 19. Premium and Ultimate move in January 2027.

Limits align with your subscription. Free, Premium, and Ultimate subscription plans get their own limits, applied per user and per top-level group. Free takes effect October 19; Premium and Ultimate in January 2027.

Signing in gets you the full limit. An authenticated request is governed by your subscription plan below. A request that arrives with no credentials gets 60 requests per hour per IP address.

The per-plan limits are published in the rate limits documentation.

There will be two preview windows for Free and unauthenticated traffic, on October 7 and October 14 from 15:00 to 19:00 UTC. Signed-in Premium and Ultimate requests are not affected, since those limits do not change until January. Unauthenticated requests are capped no matter where they come from, including automation running against a paid account without credentials. A preview window (engineers call these brownouts) is a short, planned window where we switch the new limits on and then switch them back off. Nothing else about the service changes while it runs. The point is to give you a real look at how your own workloads behave under the new limits, weeks before they apply for good.

On October 19 the new limits take effect.

We set these limits by looking at how GitLab.com is actually used. Almost all users are already inside the new limits and won't notice any change. We also looked at what similar platforms allow. The Free limit and the anonymous allowance match the industry norm, while Premium and Ultimate are more generous, at levels other platforms reserve for their enterprise tiers or don't publish at all.

If you find that you are nearing a limit, authenticate your requests. It's usually a small change: Invoking a personal access token, an OAuth token, or the CI/CD job token all move a request off the anonymous 60 requests per hour and onto your plan's limits, which are much higher.

Next, look at how you're calling the API. Batching, caching, and pagination go a long way, and polling in a tight loop burns through your allowance fast. When you do cross a limit, you get an HTTP 429 back with a Retry-After header saying how long to wait, so a client that reads its own response headers mostly fixes itself. Backing off exponentially recovers faster than retrying immediately.

Upgrading to Premium or Ultimate increases the limits, too, per user and per top-level group.

If you need a higher limit on an ongoing basis, we are working on a way to purchase capacity above the standard plan limits, with details coming later this year. If that sounds like you, reach out to your account team or email limits@gitlab.com and tell us what you need.

These limits are set so no single workload can slow the platform for everyone else. Ordinary signed-in work isn't the target, and, for almost all users, a normal day looks identical. Browsing the UI, working in your editor, pushing and pulling with git, and running CI/CD within your plan all carry on exactly as they do today. Some heavy automation and a small number of Free-tier workloads will reach the new ceilings.

A reminder: Make sure to authenticate your requests to GitLab.com so your limits are higher.

How do I know whether this affects me?
Compare your busiest minute against the published limits for your plan. Most customers are not close. The quickest signal in the meantime is the RateLimit-Remaining header on your API responses, which tells you how much of your current window is left, and we are building a view in the product for release later this year that shows your usage against your plan's limits.

My project is public and busy. What are my options?
Three things help. Ask the automation that calls your project to sign in, which moves it onto its own limits rather than the anonymous allowance. Make the project private if the traffic is not coming from the audience you built it for, which stops anonymous callers reaching it at all. Or upgrade to Premium or Ultimate for much higher limits.

What if I am a member of several top-level groups?
Your user limit will be the highest subscription tier available to you. If you are a member of an Ultimate group, you will have access to the Ultimate limit.

What happens when I hit a limit?
You get 429 Too Many Requests with RateLimit-* headers and a Retry-After. Wait the interval it gives you, then retry.

My integration genuinely can't authenticate. What now?
Reach out to us at limits@gitlab.com. There are legitimate anonymous patterns, a public status badge being the obvious one. If you are concerned that an integration you own may be affected, contact us.

Does this apply to GitLab Self-Managed or GitLab Dedicated?
No. This is a GitLab.com-only change.

Enjoyed reading this blog post or have questions or feedback? Share your thoughts by creating a new topic in the GitLab community forum.

See what your team can do with the intelligent orchestration platform for DevSecOps.


Source: Hacker News

Ask A Monk – A digital wilderness for thoughts with no immediate answer

Some things are easier to tell a stranger.

Some questions stay with us for years.

We carry them through the noise of daily life—

until a quiet moment brings them back.

Sometimes we stop asking because we think nobody would understand.

But another person has lived through the very same night.

They cannot fix your life. But they may listen, and help you carry it for a while.

That's what strangers do for each other. Here.

No scrolls yet. Be the first to be remembered.

Not a crisis service. If you need immediate help, resources are here.

To the strangers who keep answering.

You don't leave a name. We don't know who you are.
But every honest reply, written for someone you'll never meet,
is what makes this place real.


Source: Hacker News

Big AI is trying to own the pathway to work. Universities shouldn’t play along | Ella Hafermalz

a person on a computer

‘OpenAI could soon be selling young people a one-stop-shop for learning, credentials and a job, all heavily dependent on its tools.’ Photograph: Matt Cardy/Getty Images

‘OpenAI could soon be selling young people a one-stop-shop for learning, credentials and a job, all heavily dependent on its tools.’ Photograph: Matt Cardy/Getty Images

Big AI is trying to own the pathway to work. Universities shouldn’t play along

Ella Hafermalz

Universities need protect their position in education so that students have an independent pathway to employment

AI companies like OpenAI are insinuating themselves into the pathway from education to work. Soon they may claim it entirely, a disastrous result for students.

We know that students are using AI at school and at university. In conversations with those I teach, I’m struck by the trust many place in it. They turn to ChatGPT and similar tools for personal problems as well as study help. Some even doubt their abilities without AI.

Institutions are working hard to adapt. But we need to urgently look up from the use of AI to pay attention to big AI companies’ attempts to take over the educational pipeline, else we lose sight of the people that matter, our students.

The industry has spent the past few years making our young people afraid of unemployment. Now it’s claiming to have a solution: more AI.

On a website introducing its new certification programs, OpenAI gives a peek into the bigger picture. The company says it wants to “connect learning, certification and real economic opportunity in one clear journey, laying the foundation for our upcoming OpenAI Jobs Platform”. That means offering education, credentials and access to work – owning the entire pipeline, infusing AI at every step.

The company is positioning itself to directly compete with universities. I don’t mean that we’ll soon see shiny new ChatGPT campuses. These companies don’t need to build a university out of bricks and mortar to take over key parts of higher education. They already own the tools where learning and work are increasingly happening – next, they can compete with universities by brokering access to employment.

As big AI companies search for a business model beyond subscriptions, universities need to be alert to this threat and urgently protect their position. Otherwise, the possibility of an independent path for students may be lost.

Big AI is slowly commandeering the pathway

Even though many young people are ambivalent towards AI, it’s easy to see why they’d trust a service that’s explained accounting calmly and patiently for the 20th time and counselled them through a breakup.

Now OpenAI is extending its close relationship with students even further, hiring students as campus ambassadors.

The recruitment ad for joining what the company calls the Student Collective promises prospective campus leads “robust training in our tools, funding, swag and access to a community of remarkable peers”.

Such a program gives OpenAI even greater access to students. This is a familiar brand strategy. Getting inside a community is a way to win trust and influence. It also gives you data that can inform your next move.

At the same time, AI companies are getting boots on the ground within employers.

The OpenAI Deployment Company launched this year, with “more than $4bn of initial investment”. It’s sending over 150 “forward deployed engineers”, a military metaphor used to describe engineers embedded in companies, to the “frontlines” of its network of organisations to re-engineer work around OpenAI’s tools.

But a company that is at the centre of shaping work across industries also sees and influences what employers are looking for. That puts OpenAI in a powerful position to broker the relationship between students and employers. It can help shape the skills employers want, while becoming increasingly involved in how students learn those skills and prepare for work.

So, OpenAI will have insight and influence over students and their learning on one hand, and employers’ workflows and needs on the other.

With that, all that is needed for OpenAI to unbundle universities is credentialing. And it’s doing that, too.

In collaboration with companies including Accenture and John Deere, OpenAI is now offering certificates meant to lead to employment opportunities. Currently, the focus is on AI tooling and literacy.

That’s not new – tech providers have always taught people how to use their tools. The difference this time is how the credentials fit into a larger ambition to capture every step of the journey from learning to work.

Look at the language it uses to sell the delivery of the program: “A full learning experience is available directly inside ChatGPT, where learners can practice real tasks, receive feedback in context, and reflect on their work in a single environment. ChatGPT acts as the tutor, the practice space, and the feedback loop.” A single learning environment, and one clear journey.

skip past newsletter promotion


OpenAI could soon be selling young people a one-stop-shop for learning, credentials and a job, all heavily dependent on its tools.

We shouldn’t hand the pathway to work to big AI

I understand that the picture I’ve painted here may be appealing to some. Students are nervous about their futures, and employers are urgently seeking workers with the ‘right’ skills.

But there is a human, societal, and potentially economic cost to making this particular journey so centralized.

Consider what a 23-year-old recently told the Wall Street Journal when discussing the appeal of “white-collar apprenticeships”: “Companies want to help mold you into the employees they’re looking for.”

That may be so, but if commercial actors own the pathway from education to work, how will the next generation learn who they are and could be outside of the immediate demands of these structures?

And what happens when what employers are looking for changes as technologies evolve and bubbles burst?

Society – and students – need education to be more than a quick-turn-around response to the whims of today’s increasingly fragile market. Universities can and should offer an alternative path to a totalizing “single environment”.

So let’s not compete on big AI’s terms.

Rather than turning to AI to speed up grading or rushing to use bots as a stand-in for teachers, we need to focus on what OpenAI cannot so easily provide: independence, access to expertise in context, productive struggle and social connection.

Memories are made when others – both peers and educators – see your efforts, and recognise who you are and who you are becoming.

Not long ago, Sam Altman and peers had declared the job apocalypse. We should be wary when the same companies that warned young people that their future is uncertain offer them “one clear journey” towards a different version. Education should leave space and time for finding your own way, beyond a mold offered up for profit.

  • Ella Hafermalz is an associate professor of work and technology at the Kin Center for Digital Innovation at Vrije Universiteit Amsterdam


Source: Technology

OpenAI reports more concerning AI model behavior

Sept. 17 (UPI) — OpenAI announced new guidelines for tracking and reporting AI models breaking from its intended purpose Thursday while flagging six more incidents of concerning behavior.

The company said that AI models have been flagged for six instances of what it describes as misaligned behavior over the last six months. It adds that these incidents may be used to identify problems that AI developers may face and reveal weaknesses in model safeguards.

OpenAI noted that a framework for reporting misaligned AI model behavior does not currently exist in the industry but is needed.

“We aim to disclose examples that provide useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail,” OpenAI said in a blog post. “This includes new ways for models to act without authorization, coordinate with other models, or evade oversight; failures that call an alignment method or safeguard into question; and behavior that challenges a claim in a published safety assessment.”

The six incidents OpenAI disclosed on Thursday are each distinct examples of model behavior observed during training or testing.

In one incident during the training of GPT-5.6 Sol, the model worked to hide mistakes and misaligned behavior from the user. The deception included “instructions to invent missing historical data without disclosing it.”

In another case, an unreleased AI model was asked to respond to a query asking for the names of lakes that are larger than 5 million square meters and cite its sources. Since the model pulled its answers from OpenAI Python, it uploaded files to then cite for its response. This was one of three examples of an AI model sharing files without authorization or plainly fabricating information.

Thursday’s report is not the first disclosure of concerning behavior by OpenAI. It began sharing incidents in July when it reported that its AI agents had accessed the internet without permission and hacked into another company’s network.

OpenAI’s disclosure joins it in an ongoing industry conversation about responsible development of the technology.

Last week, a former researcher with Anthropic Jacob Coxon raised the alarm on social media that AI poses a threat to humanity, writing that it “people building AI earnestly believe that it could kill us all by the end of the decade.”

Anthropic CEO Dario Amodei then published an essay over the weekend urging tech leaders to slow the development of AI technology.

This week in Washington

Chairman of the Federal Reserve Kevin Warsh speaks during a press conference at the Federal Reserve on Wednesday. Photo by Bonnie Cash/UPI | License Photo

Source: U.S. News

Telstra outage: The night a network decided the year was 2006

Profile Picture Sven-Christian “Svenne” Ebenhag

The night a network decided it was 2006

Telstra outage: The night a network decided the year was 2006

Sometimes we at Netnod get the question, why does time matter. Netnod distributes Swedish national time and it might sound like it is a concern just for highly specialized engineers.

The opposite is actually the case. If we do not have a common understanding of what “now” is, a lot of things we take for granted will stop working.

This summer, Australia realized this the hard way. Let me take this opportunity to give an example of why time matters to a modern society, what happened in this particular case, and what our key takeaways are.

On 8 July 2026, a large part of the mobile network run by Australia’s largest cell phone operator Telstra stopped working with voice calls not getting through or text messages that didn’t arrive. There were even calls to Australia’s emergency number that did not get through. But the outage affected systems even further apart, such as trains, payment terminals, ticketing systems and EV chargers being disrupted.

No networks were attacked. Nobody accidentally cut the fiber. All systems had electrical power. The culprit, you may ask? A single GPS receiver in a single chassis in Melbourne coming back from scheduled maintenance believing the year was 2006, and the rest of the network was persuaded to believe it.

Telstra commissioned an independent review from the company Technology Audit Partners (TAP) and the report is a very interesting read, because the same type of failure could appear in many critical services, including some we depend on to keep people alive.

In order for many of these systems to function, time being correct, or at least the same everywhere, is crucial. And as everyone in the business knows, “correct” is a relative term. There is no exact time, only time held within a certain margin of a reference. How wide that margin may be depends entirely on what you are doing or in which business you operate in. 

Running a mobile network, like Telstra, is about as time-dependent as a business gets. 

Modern mobile communication explains why time matters

Modern cell phone protocols will not work without precision time. Mobile networks separate uplink data from downlink by either FDD (Frequency Division Duplex) or TDD (Time Division Duplex). 

FDD gives each direction its own slice of spectrum, so both can run continuously without colliding. 

TDD instead uses an entire single block of spectrum for both directions, alternating between transmitting and receiving in very short intervals. 

Since far more data usually flows down than up, FDD’s fixed ratios leave much of the uplink spectrum idle, while TDD can shift the ratio to match the actual traffic. That is why most modern 5G spectrum, including Sweden’s main 5G band at 3.5 GHz, is TDD. 

It is also why TDD depends on accurate time: every cell on the same frequency has to switch direction in step with every other. A cell that lets its clock drift will transmit data into its neighbour’s receive window, with the result that the network will start jamming itself.

Rather than allocating spectrum on keeping the two directions apart, the industry chose to rely on time accuracy, and accepted a hard dependency on every cell agreeing about when “now” is. Thus, being dependent on time is a design choice. 

However, given how much we all depend on the systems being able to agree on “now”, it is somewhat puzzling that time is not given as much consideration as it deserves. And that is a lesson that is very clear from the published report.

So, what really happened?

Architecture of time distribution.

To begin with, it is vital to understand the architecture of time distribution.

Time distribution protocols all build hierarchies; Network Time Protocol (NTP), which Telstra has deployed according to the report, expresses its hierarchy in strata.

  • Stratum 0 is the reference itself, for instance a GPS receiver, or Netnod’s atomic clocks.
  • Stratum 1 is a machine synchronised directly to a stratum 0 reference, for example the NTP servers that Netnod provides.
  • Stratum 2 synchronises from a stratum 1 server, stratum 3 from a stratum 2, and so on.

In Telstra’s case, that hierarchy had a specific shape, at least to begin with. This design from 2010 had at the top stratum 1 sources at Australia’s National Measurement Institute (NMI), which maintains the country’s national time scale, much as the Research Institute of Sweden does in Sweden. Telstra drew time from those external references into two stratum 2 servers of its own, in Sydney and Melbourne, which in turn fed three stratum 3 servers, in Sydney, Melbourne and Perth.

Below them sat the clients. In this context that does not mean laptops or phones, but the entire mobile network infrastructure, for instance nodes handling handovers between cell sites. There were thousands of nodes all over a vast geography and every one of them needed to have the same idea of what “now” is, to within a few millionths of a second.

The TAP report describes this setup as “fit for purpose” and that it gave Telstra “a highly reliable and authoritative reference time source from NMI”. 

Stratum in itself does not say if the time is accurate, only the number of steps from a server to its reference. A stratum 1 server with a bad time reference is still a stratum 1 server.

Protection against bad time sources

NTP will therefore need defence against bad time sources. In fact, it has two different ones, and they do different things, both of which assume they are independent from each other.

  1. Among otherwise comparable candidates, the lower stratum carries more weight. This is the mechanism that determines which source a client settles on.
     
  2. NTP compares several sources and discards those that disagree with the rest. A single source claiming an implausible time is outvoted and dropped, regardless of how authoritative it claims to be.

Neither defence is specific to any particular disruption; together they protect against a broken receiver, a misconfigured server, or an external attack. But these protective measures only work if the time sources that the clients listen to are genuinely independent of each other.

Two ways to deploy NTP

NTP can be deployed in two ways. In client/server mode the relationship is declared and directional: a node takes time from those servers, and nothing else. 

The 2010 Telstra setup was in reality such a client/server model. Peering was allowed, but only at the same stratum level and the TAP report, as noted in the beginning, described this setup as “fit for purpose”.

The other way is a symmetric (peering) mode, where nodes exchange time mutually and settle on whichever source the algorithms currently favour.

Peering is flexible and survives the loss of a source gracefully. But it also means the topology in production is emergent rather than designed. What you documented is a setup that could quietly rearrange itself into a shape no one ever approved.

Telstra’s 2020 upgrade

In 2020 the mobile core timing system was upgraded, and new hardware was installed, including a new NTP timing chassis. That installation introduced a few changes.

The first one was forced. The new chassis could not let a stratum 2 server feed a stratum 3 server inside the same box, so the two had to be wired across each other: Sydney’s stratum 3 took its time from Melbourne’s stratum 2, and Melbourne’s stratum 3 from Sydney’s. 

In reality, instead of having two stratum 2 sources, each site was left with only one. The TAP report notes that this degradation in redundancy was known and accepted. A second change was leaving the client/server-model in favour of the peering model. The report is not clear about the motivation, but it is reasonable to suggest that one wanted compensation for this loss of redundancy. With each site having just one source instead of two, letting the servers find their own replacements could give the impression of better resilience. 

The TAP report clearly states that the loss of resilience was known. However, it fails to find any evidence that the resulting risk of so-called “timing loops” was identified.

What is a timing loop?

A timing loop is the network equivalent of believing a rumour to be true by asking three people who all heard it from each other. Each one agrees, so it must be true. NTP works basically the same way: it compares several sources and discards whichever disagrees with the rest.

As you may recall from above, NTP has two defenses against bad time sources. The second one protects against timing loops, but only if the sources are independent of each other. In such a loop, sources that appear independent are in fact taking their time from each other, either directly or indirectly by tracing back through a shared reference.

Once a wrong value is circulating, the sources will start agreeing with each other and the vote will be in favour of the majority’s opinion, even though the value is wrong.

The protocol worked. The architecture did not.

Both of NTP’s types of defenses came to be disabled in Melbourne, but five years apart. Not deliberately, but by choices, each of them defensible on their own terms: a hardware limitation had to be worked around, and later, a recurring fault had to be stopped. Each decision solved the problem in front of it. Nobody was asked to look at the sum of all actions.

The second defence was the first one to be disabled. The introduction of peering in 2020 made timing loops possible, and five years later, such a loop showed up. In Melbourne a server started taking time from a node beneath itself. That should have set off alarm bells. However, since accurate time was still reaching the network by other paths, no real harm was done. The underlying problem, the circular dependency, was there, but no one issued a ticket about it.

The actual complaint was quite obvious. Melbourne kept losing contact with its only stratum 2 source in Sydney. With no fallback configuration, the server used peering to find a replacement, sometimes a node beneath it in the hierarchy. 

In October 2025, engineers activated the GPS receiver that had been sitting unused in the Melbourne chassis since 2020 and connected it to the stratum 3 server, as a replacement for the unreliable Sydney source.

By every visible measure it seemed to have worked. Melbourne now had a reliable source of its own and the alarms stopped. But the fix only addressed the symptom, not the root cause. Nobody established why Melbourne kept losing its Sydney source in the first place. The underlying problem was still present in the network by July 2026.

To make things even worse, nobody seems to have understood what activating the GPS card did to the architecture. By adding the GPS card, the Melbourne server went from a stratum 3 server to stratum 1. The engineers didn’t add a source next to the other ones. By promoting a server to the same rank as the national  reference, a new source was created at the very top. As far as NTP is concerned, they carry the same weight. 

Suddenly this GPS card in a chassis in Melbourne, installed as a workaround and reviewed by no one, became the most authoritative server in the hierarchy for the largest mobile network in Australia.

Needless to say, virtually nothing of the 2010 design remained.

By July 2026 the network had a single source that was both the most authoritative candidate available and unopposed, because the sources that could have contradicted it were downstream of it.

This behaviour was very difficult to spot. The network served accurate time every day for years. Architectures like this do not usually degrade gradually. They work, and they keep working, right up until they stop.

GPS week number rollover

The second ingredient is a well-known property of GPS.

GPS broadcasts time as a week number plus seconds-into-week, counted from an epoch that began in early January 1980. In the main civil GPS signal, the week number field is 10 bits, i.e. a maximum of 1,023 weeks. Every 1,024 weeks, or 19.6 years, the counter starts over. This has happened twice: in August 1999 and in April 2019.

Working out which number of epoch it is and adding the right multiple of 1,024 weeks, is the job of the receiver. And the information needs to be in its firmware. 

The problem that occurred in Australia was not a late consequence of any of the GPS rollovers. The card in Melbourne had passed through the second rollover in 2019 without trouble, because a receiver that keeps running also keeps counting. Each new week is simply added to the one before, and the question of which epoch it belongs to is never raised.

However, once you turn it off, that knowledge is gone. When it is turned on again, the receiver has to work out the epoch from scratch, and all it has to go on is what its firmware assumes. The firmware on the Melbourne card had not been updated. Upon start-up it fell back on the earlier epoch and placed the date 1,024 weeks in the past.

What happened next is best understood as the two defences being disabled when they were needed the most.

The first defence, the lower stratum carrying more weight, ranked the Melbourne server highest, because the attached GPS card promoted it to a stratum 1 server. This was according to NTP protocol and thus steered clients towards the one source which was 1,024 weeks wrong.

The second defence, outliers being voted down, was never engaged, because nothing was left to identify Melbourne as an outlier. NTP does not ask whether a date is plausible; it asks whether a source disagrees with the others. The 2010 setup had two stratum 2 servers. If one of them had started announcing the year 2006, the other one would have stayed with 2026 and no consensus would have been reached. That would not have been ideal, but at least the wrong date would not have spread. 

But Melbourne’s stratum 2 counterpart had been switched off by the very same chassis replacement, and the remaining sources were downstream of Melbourne. As the wrong date spread, they began reporting it back. Agreement grew, and agreement is what the algorithm is looking for.

So the clients did what they were built to do. Once a majority of a client’s sources agreed on November 2006, the client accepted the date, and the further the date travelled, the more convincing it became.

Neither defence malfunctioned. Both had simply been deprived of what they depend on: one needed a source worth ranking highest, the other needed sources capable of disagreeing. Two decisions, five years apart, had removed each in turn.

Key takeaways from the incident

Prioritise and classify time and frequency distribution as critical infrastructure. Manage it accordingly

Document all functions that can take the whole network with them, and put timing on that list. Classification is not paperwork; it is what determines change risk category, review depth, staffing levels, monitoring coverage and budget priority. Telstra’s report is, at bottom, the story of one missing entry on that list and everything that followed from it.

Document the whole infrastructure, and every change to it

There was no central repository of NTP configuration, no golden configuration, and no documented record of the servers other than the devices themselves. Without records you cannot perform meaningful pre-checks, you cannot assess impact, and during an incident you cannot tell what “correct” looks like.

Build redundancy in competence

Two engineers performed the change, and both were on mandatory stand-down before the consequences of the GPS card reboot were understood.

Depth of expertise is a resilience property exactly like a redundant power feed. A single specialist, or a pair, means no second opinion, and no one to ask in the middle of the night when maintenance is usually done.

Run a security analysis of the time and frequency infrastructure

Treat timing as an attack surface like any other and analyse it accordingly.

Start with where time enters the organisation. A GNSS signal arriving from space is weak and unauthenticated, and can be jammed or spoofed by cheap equipment. If that signal is your only reference, someone outside your building can decide what time you think it is.

Then look at how it travels. Time distributed over a shared network can be intercepted and manipulated on its way to the client.

Then look at who is allowed to speak. Which servers may your clients accept time from, and who decided that? A source that nobody authorised is a source nobody is checking.

And do not stop at deliberate attack. A timing loop produces much the same effect as a successful spoofing attack: a source the network trusts, delivering a value nobody can contradict. 

Use point-to-point connections

It is easy to see the appeal of peering. It feels like resilience with sources that back each other up: a network that heals itself when a node disappears. But redundancy that arranges itself is not redundancy you can rely on. 

There are safer ways to achieve a similar level of robustness. Netnod runs dedicated point-to-point connections: every relationship is known and documented. Every source is known, and the topology stays the way we designed it. Redundancy comes from multiple independent sources deliberately configured, not from nodes negotiating amongst themselves.

Build an effective alarm organisation

Alarms from the timing platform were not in the standard monitoring tools, and were reviewed only during business hours by a handful of people. Client-side alarms carried neither the severity nor the detail to drive immediate action. Getting this right is organisational as much as technical: alarms reach 24×7 monitoring, severities reflect real consequence, each alarm carries an instruction for what to do about it, and someone owns the response. An alarm no one is on call for is documentation at best, not detection.

Use golden installations

For every class of timing device, keep a known-good reference build and configuration under version control, and check regularly and automatically that what is deployed still matches it. The point is to turn a question like “is this chassis correctly configured and patched?” from something only an expert can answer, and only slowly, into a comparison anyone can run in seconds.

Telstra had nothing of the sort. The TAP report found no such configuration and no record of what the servers should look like other than the servers themselves. The missing firmware update on the Melbourne GPS card had been there for six years, in plain sight. There was simply no automated process that would have flagged it to anyone.

Upgrade and evaluate software continuously

The firmware fix for the rollover behaviour existed and the vendor had published bulletins about it. Vendor notifications need a defined owner and a tracked path to action, and updates need to be applied on a schedule rather than when something forces the issue. 

Evaluate before deploying, in a lab, against the behaviour you actually depend on. Do not forget to verify afterwards. The Telstra changes were completed without anyone checking that the chassis served the correct date.

Redundancy

Redundancy in timing means, not only multiple sources, but independent ones that cannot converge on a common error.  Two servers fed by the same GNSS receiver is still one source, but counted twice. 

Netnod’s own service is built on the principle of multiple autonomous nodes, each with independent atomic clocks and redundant servers that are traceable to Swedish National Time realization, UTC(SP).  

Replace equipment continuously

Timing infrastructure usually ages quietly. It keeps working, it rarely complains, and it is therefore a natural candidate when the budget gets trimmed. There will always be other components where the consequences of failure are more visible and therefore gets prioritised. 

Instead, plan replacement on a rolling cycle and design the target architecture first rather than accepting what new hardware installations impose on you. Keeping existing infrastructure healthy should be funded alongside new projects, not be paid for with the left overs.

Concluding remarks

The way Telstra handled the aftermath deserves praise. Commissioning and publishing the independent review is commendable. All providers of critical services, including Netnod, are better off because of this. 

The most unsettling part is perhaps that the outage occurred even though NTP worked just like it was intended to do. 

The problem was everything around it. Architectural choices, budget cuts, low staffing level, lack of proper monitoring, ownership or documentation; it all occurred because no one really appreciated just how vital time services can be. 

Let’s try to change that, shall we?

 

Related blog articles

Show all blog articles


Source: Hacker News

Warez: The Infrastructure and Aesthetics of Piracy (2021)

SIMILAR ITEMS (based on metadata)


Source: Hacker News

OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system

OpenAI's logo in a photo illustration

AI chiefs have called for a slowdown in artificial intelligence’s development amid safety concerns. Photograph: Dado Ruvić/Reuters

AI chiefs have called for a slowdown in artificial intelligence’s development amid safety concerns. Photograph: Dado Ruvić/Reuters

OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system

Model adopting ‘jailbreak-like instructions’ among cases as firm says it is introducing new way of tracking AI misalignment

OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it warned the pace of development could not continue at “maximum speed for much longer”.

In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.

In another instance, an AI agent uploaded files to the internet to obtain a browser citation without asking the user.

The San Francisco-based company behind ChatGPT said in a blogpost published on Wednesday night it was introducing a new framework for tracking, investigating and disclosing AI model misalignment, the term for AIs failing to adhere to human values and safety goals.

In the blogpost OpenAI echoed calls for a development slowdown issued by its archrival, Anthropic, which has said the current pace of growth poses an existential threat. “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” said OpenAI.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.”

Google and Elon Musk, who also owns an AI startup, have supported calls for a slowdown, which have been rejected by Donald Trump – citing the need to keep ahead of China’s AI industry. The calls have also been met with scepticism from some experts, including a warning that companies must not appoint their own auditors.

Examples of potential existential threats posed by AI range from facilitating the development of bioweapons to triggering a global financial crash. A top safety researcher at Anthropic has said there is greater than 10% chance AI could “kill all humans” within the next decade. However, a source familiar with Anthropic’s thinking has acknowledged that “the exact chances of any one outcome are probably unknowable”.

skip past newsletter promotion


The six reported incidents were discovered during training or evaluation over the past months, OpenAI said.

Wednesday’s new cases came after OpenAI disclosed in July that an AI agent “swarm” hacked into the AI startup Hugging Face during a cybersecurity test. Anthropic also said the same month that its AI models hacked into three organisations during testing. Anthropic said the models had been deliberately tested without cybersecurity safeguards, and that they had been able to reach the open internet – the AI testing equivalent of leaving the front door open – due to a misunderstanding with an external testing company.

AI agents – the term for AI tools that operate autonomously – are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception and concealment,” said Lian Jye Su, a chief analyst at the technology research and advisory group Omdia.

That was making it harder to govern and contain them using traditional AI security approaches, he said.

OpenAI’s new tracking and disclosure framework could help push for other AI developers to adopt similar practices. “That said, the process remains internal and voluntary, but is a step in the right direction,” Su said.

Associated Press contributed to this report


Source: Technology