Train a small model that chooses among a changing list of text options.
A Jev-like model takes a piece of text and a list of N text options. It returns one probability for each option. It does this in one pass instead of writing an answer word by word. Jev is TypeSafe’s commercial model for this kind of task. TypeSafe has not published its design. This repository is an independent starter model with the same input and output shape.
Demo
The same option-attention head can score controller buttons from image patches. This ten-second film joins two selected five-second windows: live deadly_corridor combat on the seven Doom buttons, then a chess controller walking to and playing moves with five keys. The diagram shows the tensors used for each decision. The Doom window came from the supplied joint checkpoint, which averaged 0.60 kills and -97.50 reward across its ten recorded episodes. The chess window came from the stronger chess-only checkpoint, which scored 4 wins, 46 draws and 0 losses in 50 sampled games against a random mover, but 0 wins, 2 draws and 48 losses against Stockfish level 0. The windows were selected for activity and are not typical-play or competence claims.
Install the game extras and record a fresh 640 by 480 Doom trace from the released joint checkpoint:
The release includes the Doom example, the chess example, the single-game checkpoints and the shared 12-option checkpoint. Both games import the visual scorer from jevlike.vision; there is no second model copy in either example.
Architecture
Each option becomes a query vector, which is a short list of numbers representing its text. The query assigns attention weights to the context tokens. Those weights make one context vector for that option. A shared dot product turns each option and context pair into one score. A softmax, which converts scores into probabilities that sum to one, runs across the options.
The default encoder learns byte embeddings from scratch. An encoder is the part that turns text into vectors. The optional Hugging Face path uses a frozen pretrained encoder, whose existing weights stay fixed while the small scorer learns.
Data format
Use one JSON object per line:
{"context":"The customer needs a refund.","options":["refund","sales","technical support"],"label":0}
label is the zero-based index of the correct option. Each row may have a different number of options, with a minimum of two.
Quickstart
Run these commands from the repository root. They create local synthetic data, train on it, evaluate the saved model and score one new menu.
The evaluation prints top-1 accuracy, which is the fraction of correct first choices. Top-3 accuracy is the fraction with the right answer among the three highest scores. Expected calibration error compares confidence with observed accuracy. The command also prints a shuffled-context control, which pairs each menu with the wrong context. A useful model should beat that control.
Use your own data
Export train, validation and test JSONL files in the format above.
Keep all options that the model will see at prediction time in each row.
Split related records together. For example, keep all records for one customer or one target page in one split. This prevents near-duplicates from leaking into the test set.
Run jevlike-train with your train and validation files.
Run jevlike-eval once on the held-out test file. Held-out means the file was never used for training or model selection.
The default byte encoder truncates context to 192 bytes and each option to 32 bytes. Raise --context-tokens or --option-tokens when your text needs more room. Training supports CPU, Apple MPS for a Mac GPU, and CUDA for an NVIDIA GPU through --device.
Use a frozen pretrained encoder
Install the optional dependency and name any compatible encoder from Hugging Face:
The checkpoint stores the trained scorer head and the encoder name. It does not copy the frozen encoder weights. Loading the checkpoint therefore needs access to the same Hugging Face model.
--rank sets the width of the small scorer head. A wider head has more trainable weights and uses more memory.
Wikispeedia example
scripts/get_wikispeedia.sh downloads the public SNAP archives and builds next-click JSONL files. The data stay outside this repository.
Cite Robert West and Jure Leskovec, Human Wayfinding in Information Networks, WWW 2012. Review the source data terms on the SNAP dataset page.
What to expect
In the experiments that led to this starter, the one-pass scorer reached about 98% accuracy on synthetic menus. On target-disjoint Wikispeedia next-click data, a frozen Qwen2.5-0.5B encoder plus the scorer reached 26%, against about 8% for shuffled and random-encoder controls. A small model trained from scratch on 40,000 clicks reached 29%. At eight options, one pass was about 100 times faster than a small decoder forced to write 400 tokens.
These numbers describe local experiments, not this quickstart run. We did not show equal quality with Jev or reproduce TypeSafe’s private training method.
Limitations
This is a research starter, not a copy of Jev.
Accuracy depends on data quality, split quality and the encoder.
The byte encoder is cheap but weak on language meaning.
The pretrained path may download a large model and needs more memory.
One-pass scoring requires the complete option list before prediction.
The speed comparison used a small local decoder rather than a large commercial model.
Licence
Code is released under the MIT License. Downloaded datasets and pretrained models keep their own terms.
Twelve years ago, I wrote 174 lines of PHP as a stopgap for AOL’s content management system. I put it on Packagist in case anyone else needed the same patch, and somehow it’s been installed nearly 20 million times since. Today I marked it deprecated.
A temporary shim
In 2014, we were in the middle of upgrading AOL’s CMS from PHP 5.2 to 5.3. Part of that upgrade was dropping version 1 of the pecl_http extension, which gave us a function called http_build_url(). A CMS deals with a lot of URLs, and ours called that function in dozens of places. I wasn’t touching those. The function seemed straightforward enough to reproduce, so I wrote my own http_build_url(), defined only if the real one didn’t already exist. The old code never knew anything had changed.
Composer was just taking off at the time, which made sharing it easy. I figured it would earn its keep for a year or two, until the PHP community moved on to something better.
That’s a lot of installs
Well, it wasn’t temporary. It’s been installed from Packagist nearly 20 million times, and it still picks up over 400,000 installs a month.
And it turns out Composer is only part of the picture. WPML, the market-leading multilingual plugin for WordPress, bundles the polyfill directly in its codebase, and WPML says it’s installed on over 1.5 million sites. The domain-name library idna-convert depends on it too, which is how it ships inside the source of SPIP, a French content management system, and how it ended up packaged in Debian and Ubuntu. Between all of them, there’s a pretty good chance you’ve visited a website that is still running my code.
I never imagined it would go this far.
Coming back to it
I didn’t grasp how far it had spread until a few months ago, when I looked at the package for the first time in years. I knew it had users. By 2021 I’d been out of PHP for a while, and the downloads were surprising enough that I asked for a new maintainer. Three people offered. Shortly after I asked, we lost a family member unexpectedly, and it turned our world upside down for a while. I never followed up, and that’s on me. By the time things settled, other goals had taken over, and I forgot about the package for years.
Along with the numbers, there were a handful of GitHub issues, including one where joining a path onto a URL with a trailing slash strips every letter “a” out of the path. So much for straightforward. Under a comment that reads // Workaround for trailing slashes, my code tacks an “a” onto the path so there’s always a last segment to cut off, then cuts it off with a find-and-replace. When the path ends in a slash, that last segment is just the “a”, and the find-and-replace takes every other “a” in the path with it. I can’t believe the bug went unnoticed for as long as it did.
So I had a decision to make. I could dive back into PHP after almost a decade away, hand the package to one of the people who’d offered, or let it keep sitting there.
None of the above
It was always meant to be temporary, so I’m retiring it. The PHP League’s URI library has been the community’s answer for years, and PHP 8.5 now ships a standards-compliant URI API in the language itself (thanks to jawira for pointing me at it). Both are better than a 174-line shim from 2014. Maintaining the package would only delay the move everyone should be making, and handing it over would add a risk on top of that. I don’t doubt anyone who offered, and ozh has kept a fork going for YOURLS. But a widely installed package with a new maintainer nobody downstream has vetted is exactly what attackers look for. Veritasium’s video on the xz Utils backdoor is the best telling I’ve seen of how that plays out.
The package will keep installing, but it won’t get new fixes, including for the missing-”a” bug. After this long without a change, even a one-line fix could have unintended consequences for someone, with no one around to support it. The README shows how to switch.
I wrote this code to ease a painful migration, for myself and anyone else going through the same one. Thank you to everyone who sent a pull request or offered to take it over, and to the people who kept filing issues long after I’d stopped reading them. It was a good run for a temporary fix.
P.S. We never migrated AOL’s CMS off the “temporary” polyfill. It ran there until the whole platform was shut down around 2020.
tl;dr: As of version 0.1.4, the gearhash crate has gained a NEON backend
which makes it roughly 2× faster on ARM64 at typical chunk sizes. It is
selected automatically on aarch64 and backwards compatible, so consumers of the
crate don’t need to do more than just update. Read on if you’re interested in
the details of how this was achieved, or skip straight to the final
results.
At the end of 2019, I was building a personal backup system, and as part of
this, became interested in a technique called content-defined chunking. The key
idea behind it is that instead of chunking files on fixed chunk boundaries, you
run a sliding window hash function across the file and trigger a chunk boundary
whenever the hash has a particular value. The downside is that this gives you
variable length chunks over a distribution, but the upside is that your chunking
is now much more resilient to byte sequences being inserted or removed from the
middle of files.
Anyway, as part of this I came across the FastCDC
paper. Its building block is the GEAR
rolling hash. Because I like fast things, I spent quite some time trying to work
out how to convert the serial algorithm published in the paper into a SIMD
algorithm. I ended up publishing the result of this as gearhash, a small Rust crate with
optimizations for SSE4.2 and AVX2.
When I wrote the crate, ARM64 was not really a target worth optimizing for. AWS
had offered ARM64 instances for a year, but only the first-generation Graviton
A1 family, built on Cortex-A72 cores and marketed for scale-out workloads rather
than general compute. Graviton2, the first generation with a competitive core,
was announced at re:Invent the same month as my first commit and did not reach
general availability until May 2020. Apple announced the M1 in November 2020.
Fast forward to today, a lot has changed. Apple has pushed ARM64 into the
mainstream of consumer hardware. AWS has shipped several further Graviton
generations and says that for three years running more than half of the new CPU
capacity it added has been Graviton. GitHub Actions added free ARM64 runners for
public repositories in 2025. On all of those machines, the gearhash crate was
falling back to the scalar loop.
On top of this, while gearhash initially had virtually no production users
aside from myself, it has since become a core part of the Xet
client, Hugging Face’s storage
protocol for large files on the Hub, which has replaced Git LFS as the default.
For gearhash this means we’re now doing between 10k and 20k downloads per day.
This renewed interest in the crate helped me find the motivation to see where I
can push things further.
The gear hash kernel is defined as a serial function over 64-bit unsigned
integers:
hash = (hash << 1).wrapping_add(table[byte as usize]);
Two properties make this difficult to vectorize:
It is a serial dependency chain. Every byte’s hash depends on the
previous byte’s. There is no data parallelism to extract from a single
stream.
The table lookup is a gather. 256 × 8 bytes is 2 KB, far too large for
any in-register permute. Every byte costs a real load.
After banging my head against this for a bit, I ended up making an observation
about the first property: the hash is 64 bits wide and shifts left by one bit
per byte, so after 64 bytes the starting value has been shifted out completely.
That means you can start hashing at any offset in a buffer with hash = 0, warm
up over 64 bytes, and from then on the hash is bit-identical to a pass from the
start.
What this enables is that a chunk can be split into strips: seed lane 0 with
the real incoming hash, seed every other lane by hashing the 64 bytes that
precede its strip, and run all strips in lockstep. When a lane reports a match,
you just need to work out which match is earliest, which is where most of the
complexity in the implementation ended up being.
I started out by doing a straight port from the SSE4.2 implementation. aarch64::uint64x2_t is two 64-bit lanes, the same as x86_64::__m128i, so the SSE4.2 structure
maps over almost mechanically.
The one thing that did not map over is the mask extraction. NEON has no
equivalent of pmovmskb, so getting the lane comparison results into a scalar
register takes a narrowing shift and a move, which I wrapped in a small movemask helper.
The result was disappointing: 0.92×, slower than the scalar code.
To understand why, we need to take a look at the loop-carried latency on ARM64.
Per iteration the NEON version would do this:
add.2d v1, v1, v1 ; h << 1, which LLVM emits as an add to itselfadd.2d v1, v1, v_g
On Apple cores each of these are ~2 cycles each (per Dougall Johnson’s M1
tables), so ~4 cycles
per iteration, and an iteration covers 2 bytes (one per lane), which comes out
to ~2 cycles per byte.
The scalar version, hash = (hash << 1) + table[b], compiles to a single
shifted-register add, add x0, x1, x0, lsl #1, with ~2 cycles of latency. That
is also ~2 cycles per byte.
Which means that the vector version does the same amount of work per unit of
critical path as the scalar one, but on top of that has to pay for the loads and
the mask extraction. It cannot come out ahead.
To win on NEON, the dependency chain itself has to get shorter.
If the chain is 2 ops per 2 bytes, why not make it 2 ops per 4 bytes by writing
out two steps of the per-byte update and multiplying through:
h₁ = (h << 1) + g₀
h₂ = (h << 2) + (g₀ << 1) + g₁
With this, h₂ depends on h through a single shift and a single add,
provided you precompute G = (g₀ << 1) + g₁. G depends only on table
lookups, not on h, so it is off the critical path.
Result: 0.92× → 1.13×, better, but still well short of the expected 2×.
add.2d v2, v1, v1 ; h << 1add.2d v2, v3, v2 ; h₁ = (h<<1) + g₀shl.2d v1, v1, #2 ; h << 2add.2d v3, v3, v3 ; g₀ << 1add.2d v1, v1, v4 ; (h<<2) + g₁ <-- on the h chainadd.2d v1, v3, v1 ; ... + (g₀<<1) <-- also on the h chain
Turns out, LLVM had just gone and reassociated it! I wrote (h << 2) + (G₀ + G₁) and it emitted ((h << 2) + G₁) + G₀. This is a legal transformation of
course, but it puts a second add back on the dependency chain.
You cannot stop the compiler reassociating a sum, but you can (try to) stop it
seeing one. The combined term is built from two table lookups, and those arrive
in general-purpose registers anyway, so the combining can just happen there:
let (t00, t01) = (table[b00 as usize], table[b01 as usize]);let (t10, t11) = (table[b10 as usize], table[b11 as usize]);// Combining the two table entries in scalar registers keeps the vector operand// opaque, which stops the compiler from reassociating the addition below into two// dependent vector adds on the loop-carried `h` chain.let g1 = vcombine_u64( vcreate_u64((t00 << 1).wrapping_add(t01)), vcreate_u64((t10 << 1).wrapping_add(t11)),);let h2 = vaddq_u64(vshlq_n_u64::<2>(h), g1);
Checking disassembly now showed only shl.2d → add.2d on the chain.
Result: 1.13× → 1.46×, finally starting to be meaningfully faster, but not quite
fast enough!
Encouraged by the result of unrolling to two steps at once, I tried the same
with four intermediate states, each still computed directly from h:
let h1 = vaddq_u64(vshlq_n_u64::<1>(h), g[0]); // hk == (h << k) + g[k-1]// ...let h4 = vaddq_u64(vshlq_n_u64::<4>(h), g[3]);
This halves the length of the dependency chain per byte again, so I expected
another large step. Measuring it however, there was no difference at all.
It was at this point that I suspected the limit may no longer be the latency
between iterations, but instead simply the CPU throughput.
Following that hunch, my focus shifted to try and reduce instruction counts
instead. The loop was now doing two loads per byte: one for the byte itself and
one for its table entry. While the latter is unavoidable, we can now can
replace those four consecutive byte loads with one unaligned 32-bit load, then
peel one byte off each word per step with a shift.
Result: 1.46× → 1.63×, another large step towards the 2× goal.
With the loads optimized, the next largest block of instructions per iteration
was the boundary test: at every step, the code checks for a chunk boundary by
masking the hash and comparing it to zero.
While I had unrolled to four intermediate states per iteration, it was still
probing them one by one. That is four separate moves out of the vector unit,
each with a branch waiting on it.
Because the common case is no match, what we can actually do is combine the four
tests inside the vector unit and make one trip out. If none of the four
positions in either strip is a boundary, then the iteration can move on after one
move to a general-purpose register and one branch:
let t = vandq_u64( vandq_u64(vtstq_u64(h1, maskv), vtstq_u64(h2, maskv)), vandq_u64(vtstq_u64(h3, maskv), vtstq_u64(h4, maskv)),);if movemask(t) == u64::MAX { i += UNROLL; continue;}
Only in the uncommon case when that check fails does the code look at the four
states one by one. With the 16-bit mask the Xet client uses for its 64 KiB
chunks, that happens about once every 8192 iterations.
Result: 1.63× → 1.81×. Quite happy with this, and here is where I stopped for
now.
Everything combined, on the crate’s 11-bit benchmark mask, that takes the NEON
path from 0.92× for the direct port to 1.81× in the version that I published as 0.1.4.
However, one thing I realized while working on this which is quite obvious in
retrospect, is how dependent the benchmark is on the mask density. That’s
because the optimized path has a fixed cost per call that the scalar path does
not, and sparser mask means more boundaries and more calls. To see how much that
matters, I ran benchmarks across a range of masks with 4 to 20 bits set.
bits set
mean chunk
scalar MB/s
NEON MB/s
ratio
4
16 B
996
250
0.25×
8
256 B
1864
1641
0.88×
11
2 KiB
1981
3565
1.80×
16
64 KiB
1999
4277
2.14×
20
1 MiB
2000
4347
2.17×
The crate’s own benchmark, at 1.8×, sits on the steep part of the curve, which
keeps rising until it flattens out at about 2.17× between 64 KiB and 1 MiB. The
16-bit row is the mask the Xet client uses, at 2.14×.
Sidenote: Below roughly 350-byte average chunks the per-call cost outweighs the
gain and the NEON path becomes progressively slower than scalar. I wouldn’t
expect anyone to use this type of configuration, but falling back to the scalar
path for masks with few bits set seems like a cheap way to close that gap.
With NEON now roughly twice as fast as scalar, it feels like it’s time to take
another look at the x86 backends. Who knows, some of the tricks I learned along
the way on the NEON implementation might carry over. Stay tuned!
Linum v2 was bottlenecked by the enormous size of its attention context window. A 720p, 5 second clip cost a whopping 110K tokens. To put that in perspective, LLMs see samples with fewer than 8K tokens for 97% of their pretraining. Attention is quadratic in cost, so the biggest lever we have to accelerate model training is pruning the context window down.
Most generative image and video systems are Latent Diffusion Models (LDMs). They split compression and generation into independently trained modules: the Variational Autoencoder (VAE) and the DiT (Diffusion Transformer). Recently, pixel-space models like the JiT have shown to be a promising alternative. It reduces two models into one and allows the diffusion model to construct a latent space specifically for generation, rather than rely on one built for reconstruction.
When trained on our (image, caption) dataset, the JiT seems to struggle to produce finegrained details. We propose a novel encoder-decoder architecture (JiT-DDT) that recovers this detail and trains much more efficiently than its LDM counterpart. Against our Linum v2 baseline, the JiT-DDT trains a text-to-image model with 3.6× fewer GPU-hours, even though it generates images with 4× the pixels.
JiT-DDT code and model weights are available under the Apache 2.0 license. We hope that by sharing our findings with the broader community, we can encourage others to also explore more efficient training methods. This should be treated as a research artifact, not a full model release. Stay tuned for more research checkpoints like this, en route to Linum v3.
Hitting the VAE compression wall
Almost all generative image and video models are Latent Diffusion Models (LDMs). These have two key components, a Variational Auto Encoder (VAE) for compression and a Diffusion Transformer (DiT) for generation.
Operating in raw pixels is too expensive (especially for video), so we first need to find a way to reduce RGB pixels into a smaller amount of tokens for the DiT. This is where the VAE comes in. It’s trained for compression and reconstruction. Specifically, it pushes our pixel-space samples through a probabilistic encoder, spits out -dimensional tokens, and then pushes these latent tokens through a probabilistic decoder to land back in pixel-space.
The VAE is trained to compress and reconstruct
input x
Encoder
encoder
μ = [?, ?]
σ = [?, ?]
μ, σ
z ∈ ℝᵏ
sample z
Decoder
decoder
output x̂
‖x − x̂‖²
+ β · KL(q‖N)
loss
ready
gradient (purple) reaches every weight* simplified: in practice the KL term is ≈ 0, previously we trained a σ-VAE with an L1 reconstruction loss plus LPIPS and GAN losses; see our VAE post.
When building a LDM, you train the VAE separately and then freeze it (i.e. no gradient flow from the DiT into the VAE). This way the latent space stays static throughout the course of DiT training. You run the VAE’s encoder to embed your data, train the DiT to traverse the VAE’s latent space, and then transform the DiT-generated latent tokens into pixel space using the VAE’s decoder.
The VAE is trained once and frozen; the DiT learns to move through its latent space
input x
❄Encoder
encoder
= E(x)
z
(1−t)·z
+ t·ε
INTERPOLATE
zₜ
DiT
trainable
dit
v̂
pred
‖v̂ − v‖²
v = ε − z
loss
ε ~ N(0, I)
gaussian
ε
SAMPLE ε
t ~ LogitNormal
sample t
ready
VAE frozen (dashed) · DiT trainable (purple) · gradient stops at the DiT · t = 0 clean image, t = 1 pure Gaussian noise
We want to eke out as much token compression as possible from the VAE, so that we can curb the cost of attention in our DiT. But if you take a survey of the popular open source text-to-image models like FLUX, Ideogram, and Z-Image, you’ll notice that they all cap out at 16×16 token reduction. This aligns with our experiments on Image-Video VAEs from a few years ago. Unfortunately, it seems like there is an empirical ceiling on the amount of compression we can get out of a standard CNN VAE without degrading the reconstructions.
Unlocking aggressive compression with a unified model
Last fall, Tianhong Li and Kaiming He published a paper (JiT) that achieves 32×32 token reduction by throwing away the VAE altogether and pushing the compression task into the DiT itself.
Patchify: 4×4-pixel patches → 48-dim tokens → linear bottleneck to 12
16×16 pixels, 3 channels (RGB) each
cut into 4×4-pixel patches (16 patches)
each patch is tokenized independently:
16 pixels × 3 channels become one 48-dim token
a linear layer W ∈ ℝ12×48 projects 48 dims down to 12
ready
Illustrative. In JiT at 512px we use 32×32 patches, so a 512×512 image becomes 256 tokens, each starting at 32·32·3 = 3,072 dims; the bottleneck maps that to 256.
This approach to reducing token counts isn’t particularly new. It was invented for vision transformers (ViT) half a decade ago, and it’s pretty commonly paired with a VAE to further condense token sequences before they enter the DiT.In Linum v2, our VAE gave us 8×8 (h×w) compression and 16-dimensional latents. At the base of the DiT, we applied 2×2 patchification to get 16×16 token compression and 64-dimensional latents. We used it in Linum v2 and so do models like FLUX.
So, why hasn’t anyone tried this before? This feels like a free lunch. You get a (potentially) lossless way to cut down attention cost, and it’s bone-dead simple.
In early 2025, papers like VA-VAE demonstrated that DiTs struggle to learn from high dimensional inputs.There are small hacks like using an external model as a regularizer during VAE training (e.g. DINOv3) that (likely) enabled models like FLUX-2 to make the leap from 64 latent dimensions to 128 latent dimensions for their DiT. But, these strategies just kick the can down the road on a clear learnability problem within the DiT. Aggressive patchification explicitly pushes information into the channel dimension, so it triggers this instability. But as it turns out, this is not intrinsic to the architecture. Rather, it’s downstream of the v-prediction, v-loss flow matching objective that everyone’s been using to train diffusion models these past few years.
A quick refresher on flow matching
In old school 2022-era denoising diffusion (DDPM), we iteratively noise a sample and train a neural network to remove the noise. This way at inference time we can use our neural network to transform Gaussian noise into a sample from our data distribution over a sequence of steps. This formulation has a host of issues (e.g. oversaturation in generation, unstable learning, distillation collapse), so in the intervening years the field has shifted away from it towards flow matching.
In flow matching, we construct a straight line path between every sample in our data distribution and a sample of Gaussian noise:The path between noise and samples does not have to be straight. But in practice, we all do it.
At , we recover . At , we get , where .We follow the DDPM convention throughout this post: is data, is noise. Some flow matching papers run the other way, with as noise and as data. The two formulations are equivalent. Then we train a network to approximate the velocity along that path:
We call this v-prediction, v-loss because the neural network is explicitly predicting velocity and it’s trained on the MSE between its velocity prediction and the ground-truth, conditional velocity field.
V-prediction and the curse of dimensionality
If you’re training a flow matching model you don’t necessarily need to train your neural network to predict and regress velocity. The three terms are linearly re-arrangeable; so you can mix and match , , and across prediction and regression targets:
Nine ways to write one objective
Three targets, each linear in the other two
Rearrange one identity to fill each off-diagonal cell
Pick what the network predicts (columns) and what the loss measures (rows). Each off-diagonal cell is one of the three identities above, rearranged to turn the prediction into the loss target. x-prediction with v-loss is the cell we use. Table after Li & He (2025).
In JiT, Li and He revisited the v-prediction, v-loss decision that the field’s been making since the inception of flow matching. They took a toy distribution (points on a spiral) and then projected these points from 2D to different high dimensional spaces of increasing size. For each of these spaces, they trained flow matching models with x-prediction, -prediction, and velocity-prediction; and found that the x-prediction was the only model type to accurately generate samples from the spiral distribution at large dimensions. DiTs have been struggling to learn from high-dimensional inputs because of the curse of dimensionality.
A 2D spiral buried in a D-dimensional space by a random projection. As D grows, epsilon- and v-prediction collapse while x-prediction keeps recovering the spiral. Figure 2 from Li & He (2025).
Velocity is . When we do v-prediction, our neural network has to implicitly learn the signal () and noise (). Noise is a random Gaussian that will cover the entire -dimensional space. So, as we scale the problem of fitting noise (within the velocity term) becomes exponentially harder. This is why aggressive patchification failed in the past and why LDMs have been struggling to learn from high-dimensional VAE latents. As we grow the channel dimension, we end up in the degenerate case where our DiT is struggling to learn high dimensional Gaussian noise.
By switching to x-prediction, we can try to side-step the curse of dimensionality. If we believe that images and videos naturally lie on a low dimensional manifold, we should be able to have our models predict effectively even with high .
Noise fills the whole D-dimensional ball; images sit on a thin sliver of it
ε ~ N(0, ID)
n = 0 · intrinsic dim = 512
x₀ ~ image manifold
n = 0 · intrinsic dim = 3
1 dot = 1 sample · D = 512 · each dot is plotted at coordinates 1, 2 and 3 of its 512same sphere, same n on both sides · the manifold is illustrative, not a real image set
In high-dimensional space (D = 512), noise is truly random. It spreads across the entire space, is incompressible, and cannot be described by any smaller number of dimensions (left). Images are intrinsically low-dimensional, so even in a high-dimensional space they cluster in a small subspace (right).
By predicting v, the model has to learn both the noise ε and the structure. Noise is the harder of the two, and the bigger D gets, the more of the model’s capacity goes to fitting it. By switching to x-prediction, the model can spend its full capacity on the low-dimensional signal, even when D becomes large.
Empirically this works, if you do x-prediction, v-loss.The network predicts . We convert that prediction to a velocity and take the MSE against the true velocity. Since , this is just x-loss scaled by : the same objective, weighted toward small (low noise, nearly clean images). In turn, this unlocks our ability to apply aggressive patchification, blow up the channel dimension, and push the compression problem into the DiT.
Extending JiT for text-to-image models
When we read about JiT, we were really excited to give it a go, since it was explicitly able to achieve 32×32 token reduction. But, we’d be remiss to say this is the only way to achieve this level of compression. Or, that everyone agrees that this is the best way to achieve this amount of compression.
LTX has been able to do this in their video models by altering their VAE’s decoder to make it an explicit denoiser (i.e. they finetune the VAE decoder with a flow matching objective). More recently, Minimax H3 has achieved 32×32 compression in their VAE by swapping out the standard ~80-150M parameter CNN Decoder with a 2B parameter transformer (roughly the size of our entire Linum v2 model). And on toy benchmarks like ImageNet, LDMs still out-perform pixel space models by a smidge. Nevertheless, we think we can overcome some of the limitations present in the original JiT paper and match LDMs’ performance in generative image and video.
Ideologically, we believe simple tends to beat complex when it comes to training neural networks at scale. Papers like E2E-VAE from last fall have shown that allowing your DiT to backpropagate (smartly) into the VAE can improve generation results and accelerate convergence dramatically.
To us, it makes logical sense that if we can specifically tailor the “latent space” for generation rather than rely on one built for compression, we can get better results. And we get the added benefit of having one cohesive model, rather than two disjoint ones.
We also think that ImageNet benchmarks on JiT understate its potential. The JiT might be able to achieve better compression than an equivalent VAE, by leaning on the scaffolding provided by the text prompts.
Text-to-image baselines
When we pretrained Linum v2, we relied on a VAE + patchification stack that afforded 16×16 token reduction. So, we trained on ~600M samples at 256px resolution before introducing 180p video and scaling up to 512px resolution.
For our JiT baseline, we wanted to get a sense of the output quality with the same image-latent-token budget. That meant we trained on 512px images with 32×32 token reduction.
JiT baseline
Linum v2 (ours, previous)* · 256×256
2.0B latent-space DiT + VAE
256 latent tokens
* image-only checkpoint
JiT (wide) · 512×512
2.0B active pixel-space DiT
256 pixel tokens
GPU-hours
0
8.3× fewer
0
samples seen
0M
5.8× fewer
0M
The JiT is able to learn 512×512 images at an identical token count (256 tokens). Convergence seems to happen much faster, but faces look airbrushed and oversaturated.
Our JiT setup
By moving from LDM to pixel-space, we transitioned from v-prediction, v-loss to x-prediction, v-loss. But, we also made a slew of other tweaks to the network:
One single-stream DiT does the compressing and the generating
Instead of alternating blocks of self-attention (image/video) and cross-attention (text-to-image/video), we concatenate visual tokens and text tokens into a single stream that goes through the DiT. This increases the attention sequence in every block and increases the FLOPs per token, but should allow for significantly more expressive relationships between text and image tokens.
v2 block: self-attention, then cross-attention v3 block: one self-attention over image and text
Wider instead of deeper
Our old model was a 40-layer transformer with 2048 hidden size. Here, we switch to a 23-layer transformer with a wider 2944 hidden size. Wider networks have become standard in recent DiT architectures (e.g. Z-Image), so we adopted the same.
Perceptual losses
When you train a VAE, you use perceptual losses like LPIPS and adversarial loss via a GAN to push the reconstructions towards what humans like. MSE on its own gives you a blurry mess. Now that we don’t have a VAE decoder, we need the JiT itself to leverage these losses to generate stuff humans like. We still use LPIPS, but instead of a GAN we use a P-DINO loss. Both are only applied when .
Moonshot’s Kimi models proved that the Muon optimizer works really well at scale. As they recommend, all the 2D matrices in our network (e.g. q/k/v matrices for attention, FFN weights) are optimized with Muon, while layers at the input/output of the network (e.g. patchification, output head) and scales/biases (e.g. AdaLN) are still optimized by AdamW.
PixelREPA auxiliary loss
It’s become pretty common to accelerate the convergence of your DiT by having an earlier layer in the network (e.g. layer 8 of a 23-layer transformer) align to the embedding of in an auxiliary model (e.g. DINOv3). This technique is referred to as REPresentation Alignment (REPA). We’ll dig into this (and the limitations) later in the blog, so hold on for that. But for now, plain REPA did not work well for the JiT. Instead, we adopted PixelREPA which masks out x% of tokens in our visual token hidden state, pushes it through a shallow transformer, and then applies the typical cosine-distance loss between all visual tokens (including the masked tokens) and the auxiliary representation from DINOv3.
DINOv3 uses 16×16 patches. We need the token count between the DINO representation and our hidden state to match, so we downsample the images before they go through DINO. For example, if we’re doing a 32×32 patchification on 512×512 images, we will have 256 tokens. We downsample the image to 256×256 before passing it through DINO’s 16×16 patchification to also get 256 tokens.
h₈mask20%transformer2 layersprojcosDINOv3-L(x₀)
PixelREPA · tapped after block 8 · x₀ downsized to 256px so DINOv3 gives 256 tokens
Sigmoid attention gating
Now that we’re moving from a cross-attention to a single-stream DiT architecture, we may be at a higher risk of attention sinks. We adopt sigmoid attention gating to neutralize this issue.
hAttentionσ(W_g h + b)×to residual
attention gate · 23 gates per block (1 for each attention head) · × elementwise multiply
Qwen text embeddings
Instead of T5-XXL text embeddings, we use hidden states from a more modern decoder-only LLM, Qwen3.5-4B. One downside to using a LLM is that it’s unclear what hidden state to take as your embedding. Most modern LLMs use some sort of alternating sequence of sparse/linear attention and full-attention. We take the hidden states calculated after full-attention blocks, concatenate them together, and have the DiT learn a transform to combine these representations into a single text condition. Recent Ideogram and FLUX models are more aggressive here, using larger LLMs and aggregating information across all hidden states. Given the size of our DiT, it seemed like overkill to go down that path.
text conditioning · three hidden states, one learned projection
The JiT recipe
Noise schedule
At train time, we get to pick the distribution from which t is sampled. Empirically, there’s a small band of values at high t where the structure of the image is determined. This is the hardest part of the trajectory for the model to learn, so we skew timesteps accordingly.
Note that the σ term is tied to pixel count: σ = 1 is for 256×256, σ = 2 is for 512×512. The intuition is that we need to shift more aggressively at higher pixel counts. There is more redundant information within the image, so we need to noise more. This is the exact schedule from JiT.
Loss
The MSE is x-prediction with the velocity weighting (x-prediction, v-loss), clamped at t = 0.1 so the weight caps out at 100×. LPIPS and P-DINO are perceptual losses on the predicted image. PixelREPA aligns the block-8 hidden state with DINOv3 features of the clean image (downsampled 2× to match the token grid of h₈).
Recovering finegrained details in pixel-space
One of the biggest limitations that folks have observed about JiTs is that they struggle to generate the finegrained details. Our baselines corroborate this. If we want to really get our pixel space models to sing, we need to fix this.
DDT: Decoupled Diffusion Transformer
LDMs face the same issue, but to a much lesser extent.
In early 2025, Shuai Wang and team tackled this problem directly with their DDT (Decoupled Diffusion Transformer), scoring SOTA on ImageNet gFID at the time. They observed that —
In each denoising step, diffusion transformers encode the noisy inputs to extract the lower-frequency semantic component and then decode the higher frequency with identical modules. This scheme creates an inherent optimization dilemma: encoding low-frequency semantics necessitates reducing high-frequency components, creating tension between semantic encoding and high-frequency decoding.
DDT splits the model into two components, a “conditional encoder” for low-frequency structure and a “velocity decoder” for high-frequency detail. They give the encoder most of the layers, since the most difficult portion of the probability path to master is the transition from random noise to basic structure.When you train a diffusion model, you have to sample timesteps between and to turn your clean samples into noise-interpolated . Since Stable Diffusion 3, it’s been widely known that there seems to be a small band of high- values where the overall structure of the image is determined. We oversample from this part of the distribution, to accelerate convergence. This finding was the inspiration for the DDT to allocate the bulk of its parameters to the encoder.
DDT splits one DiT into a large structure encoder and a shallow detail decoder
DDTv̂velocity decoder≈ 1/4 of the layersz_tself-conditionconditionencoder≈ 3/4 of the layersREPADINOv2(x₀)tclass yx_tdecoder:high-frequencydetail from x_tencoder:low-frequencystructureclass y reaches the encoder onlyconventional DiTv̂DiT blocksN×REPADINOv2(x₀)x_ttclass y
blue = encoder · orange = decoder · dashed box = auxiliary loss
Within the encoder they apply REPresentation Alignment (REPA), an auxiliary loss that accelerates training by aligning the hidden states of an early layer of the model to the DINO representation of the clean image (). If you keep REPA loss active throughout all of training in standard DiTs, it actually hurts overall FID.
The DDT avoids this problem by giving the decoder the noised image () so it can extract the finegrained details that REPA might otherwise destroy. Moreover, the DDT frees up the decoder to focus solely on details by creating an information bottleneck. The decoder doesn’t get the class label. Given its limited capacity, it’s forced to rely on the encoder’s hidden states to ascertain structure. And in turn, it allocates its parameters to focus on detail recovery.
Naturally, we tried to port over the core ideas from the DDT so that we could recover detail in our pixel space model. We call this new architecture JiT-DDT (creative, we know).
Ours is trained with x-prediction, v-loss, unlike the DDT which was trained with the classic v-prediction, v-loss formulation.
This is crucial. If you downsample an image, you strip it of most of its high frequency detail, leaving behind low-frequency structure. So in an x-prediction-world, we can get our encoder to learn structure explicitly by predicting a low-resolution version of our input image, .
Concretely, we split our DiT in half. We give the encoder and decoder their own input patchification and output heads, so they can specialize. We have the encoder predict a 64×64 version of the input 512×512 image (8× downsampled) and pass its hidden states to the decoder. This way the decoder gets a structural sketch of the output, its own view of the noised image, and the text prompt to create the full resolution image.
JiT-DDT: the encoder plans a 64×64 image; the decoder uses that plan, full-resolution x_t and text to generate the 512×512 image
blue = encoder · orange = decoder · teal = rep tokens · gray = frozen · dashed = loss‖ concat · MSE gradient from decoder backpropagates all the way through the encoder
One added benefit of this design is that we can have asymmetric patch sizes between the encoder and the decoder. If we’re training on 512×512 images, the encoder will be tasked with regressing the 64×64 version of the image. It needs less information to do this task, so we can use 64×64 patches and operate on a 64-token sequence. Meanwhile, the decoder can use smaller patches to help recover detail (e.g. 32×32 patches or even 16×16 patches).
We make two additional deviations from the original DDT’s architecture:
Our encoder and decoder are equally sized. In the original DDT, the encoder predicted at the same resolution as the decoder. That’s not the case for us. We’ve simplified the problem dramatically for the encoder by having it predict an 8× downsampled image, so it doesn’t make sense to have the encoder be way bigger than the decoder. In the future, we’ll have to run ablations to find the ideal encoder/decoder block ratio.
Decoder gets the text condition. DDT was a class-conditional model on ImageNet. That’s a relatively tiny domain compared to open world image and video generation. A lot of the detail that we want to recover will be annotated in the text, so we thought it’d be better to give the decoder access to this information. We tried removing the text condition in one of our ablations, and it was a wash. So, for the rest of our DDT experiments, we retain the text condition in the decoder.
The JiT-DDT loss: the encoder is scored on the 64×64 plan, the decoder on the 512×512 image
Encoder · the 64×64 plan
Decoder · the 512×512 image
We compute the MSE loss twice, once for the encoder against an 8× downsampled version of x₀ and once for the decoder against the original x₀. We keep the alignment loss in the encoder, moving it earlier in the network (tapped after block 6 of 12, against DINOv3 features of x₀ downsampled 4× to match the token grid of the encoder’s h₆). And, we keep the perceptual losses on the full resolution outputs of the decoder. The gradient runs end to end, so the decoder’s loss trains the encoder too.
JiT-DDT 64/32 (baseline) reduces the oversaturation issue
JiT (wide) · 512×512
2.0B active pixel-space DiT
256 pixel tokens
JiT-DDT 64/32 (baseline) · 512×512
2.2B active pixel-space DiT
320 pixel tokens = 64 encoder + 256 decoder
Both models are trained on the same 100M samples. The JiT-DDT is ~28% slower to train than the JiT (still much faster to converge than v2). Details are learned earlier in training. Images are a lot less oversaturated, but they only look 10–20% better.
Adjusting the noise schedule
We were honestly surprised that the images from the JiT-DDT weren’t that much better than the JiT. So, we ablated a bunch of different training and architecture decisions (e.g. warm-start the encoder before adding the decoder, dropping text from the decoder, etc.).
Nothing worked, until we started tweaking the noise schedule. It’s the most obvious knob to tune, but somehow we haven’t found any research dialing this in for pixel-space models.
Partway through training, widen the noise schedule toward clean images
share of training timesteps within 0.08 to 0.13 · hover the plot to move
phase 1 · LogitNormal(0.8, 0.8)
<0.05%
phase 2 · LogitNormal(−0.2, 1.0)
3.0%
phase 2 / phase 1
107×
Widening the noise schedule helps JiT-DDT recover details like freckles and hair texture
JiT-DDT 64/32 (baseline) · 512×512
2.2B active pixel-space DiT
320 pixel tokens = 64 encoder + 256 decoder
JiT-DDT 64/32 + noise shifting · 512×512
2.2B active pixel-space DiT
320 pixel tokens = 64 encoder + 256 decoder
Architecture is identical but noising schedules are different. In the baseline, we did 100M samples in Phase A. In the second noise shifting version, we did 109M samples in Phase A and 33M samples in Phase B.
Architecture refinements
Last fall, Alibaba’s Z-Image became the best small, open-weight model on the market. Their technical report contains a lot of juicy details, but we were most interested in the tweaks they made to the architecture:
Four Z-Image changes to the DiT block
beforetext tokensimage tokens‖N×shared DiT blocksaftertext tokensimage tokenstext refiner2 DiT blocksno AdaLNimage refiner2 DiT blocksAdaLN on t‖N×shared DiT blocks
Refiners clearly helped our model. They’re cheaper versions of the MM-DiT blocks invented by BFL in FLUX. Both help the model massage the modalities before combining them in a shared DiT trunk. The other knobs (AdaLN truncation, post-norm gate + tanh, RMSNorm) didn’t move the needle for us, so we omit them from our experiments.
Refiners fix the excessive freckling and the color grading
The encoder and the decoder each get four modality-specific blocks before the shared DiT trunk: two for image, two for text (8 new blocks in total, 477M additional parameters).
We trimmed one shared block from each stack (2 total) which brought the refiner model within ~5% of the no-refiner model’s FLOPs. The modality-specific blocks are much cheaper to run, since they see about half the sequence length of the shared blocks.
Note that the two runs shift the schedule at different points: the refiner model switches at 80M samples and trains 58M more under the wider schedule, while the previous model switches at 109M and trains 33M more.
From Linum v2 to JiT-DDT
We’ve covered a lot of ground, so let’s recap real quick.
Our goal is to reduce tokens in the DiT context window. That way we can accelerate training and inference. Traditionally, DiTs have struggled to learn from high-dimensional inputs because of the curse of dimensionality implicit to v-prediction.
If we swap in x-prediction, we can get DiTs to successfully learn from high dimensional samples. We can use this fact to apply linear patchification, throw away the VAE, and push the compression problem into the DiT. This way we can develop the latent space specifically for generation and at the same time get the token savings we’re looking for.
The one downside to this approach is that the JiT struggles to learn finegrained details out of the box. Humans perceive these details quite easily, so we need these if we want to generate good images and videos. Our JiT-DDT is one way we can get pixel-space models to learn structure and detail.
JiT-DDT trains in 3.6× fewer GPU-hours at 4× pixels
Linum v2 (ours, previous)* · 256×256
2.0B latent-space DiT + VAE
256 latent tokens
* image-only checkpoint
JiT (wide) · 512×512
2.0B active pixel-space DiT
256 pixel tokens
JiT-DDT (ours, new) · 512×512
2.5B active pixel-space DiT
320 pixel tokens = 64 encoder + 256 decoder
GPU-hours
0
0
3.6× fewer
0
samples seen
0M
0M
4.2× fewer
0M
Why does the JiT-DDT work?
We think that it’s useful to look at the JiT-DDT in the context of three papers (iREPA, Self-Flow, RAE v2), to try to unpack why our architecture works in the first place.
DiTs struggle to learn structure on their own
As we mentioned earlier, REPA has become a standard way to accelerate DiT convergence. The original authors tried a few different vision encoders and found that DINOv2 worked the best. But, it wasn’t until iREPA late last year that anyone took a serious look into why DINO seems to work so well.
iREPA trained a bunch of generative image models on ImageNet with REPA, using a larger test bed of vision encoders.DINOv3, modern JEPA variants and Masked Autoencoders. They looked at the models’ gFID scores and tried to determine whether generation quality could be attributed to either the vision encoder’s understanding of the holistic image or its understanding of spatial structure.
For holistic understanding, they relied on linear ImageNet probes. For spatial structure, they constructed a suite of self-similarity metrics. These quantify how much more correlated patches from an object are to each other than patches from other objects in the same image (e.g. patches of a lion’s mane should be more correlated with other parts of the lion’s mane than patches of the background skyline).
They found that higher ImageNet probe accuracy predicted worse gFID, while higher spatial self-similarity predicted much better gFID. Accordingly, it seems like REPA accelerates training by getting early layers of the network to see local structure, not holistic visual concepts.
Here, we have two distinct vision encoders, WebSSL-1B and SpatialPE-B. WebSSL-1B scores higher on the ImageNet probe (76.0% vs. 53.1%) but lower on spatial self-similarity (0.18 vs. 0.34).
Pay attention to the red box in the middle column, highlighting the dead grass. Yellow is most correlated, green is somewhat correlated, and blue is least correlated. In WebSSL-1B, the grass is correlated with everything but the lion (e.g. correlated with the sky). Meanwhile in SpatialPE-B, it’s only correlated with the other blades of grass.
The iREPA authors find that DiT aligned to WebSSL-1B generates worse images than those aligned to SpatialPE-B (gFID 26.1 vs. 21.0). Figure from iREPA (2025).
We think that this finding rhymes with the encoder-prediction task in our JiT-DDT. Downsampling images (e.g. 512×512 to 64×64) strips images of all detail, leaving us only with structure. By predicting the low-resolution image early in the JiT-DDT, we are providing a similar signal.
Learning structure earlier in the DiT unlocks better image generation
Taking a step back, it feels really weird that we’re aligning a multibillion parameter DiT to the hidden space of a ~100M unsupervised vision encoder. Bigger models should have more capacity, so it’s sus that we’re relying so much on the representation space of a tiny model.
Black Forest Labs (the authors of Stable Diffusion and FLUX) seem to agree with our premise. In Self-Flow, they throw away DINO and achieve better FID results by aligning to the hidden states later in the network.
Self-Flow: the student sees patches noised at two levels (); an EMA teacher sees the cleaner version, and its deep hidden states replace DINO as the alignment target. Figure from Black Forest Labs (2026).
Another paper from last year found that the later layers of the DiT learn structure quite quickly, while early layers lag significantly. If the deeper layers already learn this structure without external intervention, we can simply align to them. This way you accelerate learning, without the representational ceiling imposed by traditional REPA.
We see Self-Flow and our JiT-DDT as cousins of sorts, tackling 3 core problems with different solutions:
Slow Structure Learning in Early DiT Layers: Self-Flow aligns to later layers that have learned structure. We make structure learning explicit by regressing the low-resolution images with our encoder.
REPA’s Loss of High Frequency Details: We view Self-Flow as a form of self-distillation. It allows the model to make better use of its billions of parameters, freeing up later layers to generate detail once the early layers learn structure. We achieve the same effect by having two patchifications: one coarse and the other fine. The encoder learns structure explicitly, propagates its representation, and frees up the decoder to explicitly learn detail.
Insufficient Exposure to Low-Noise Timesteps: Self-Flow relies on dual-timestep noising, which provides additional exposure to low-noise timesteps. We explicitly widen the noise distribution, after structure is learned.
PixelREPA still wins at ~100M samples
Early on, we tried JiT + Self-Flow and it performed worse than JiT + PixelREPA.
Our gut is that this discrepancy just comes down to the amount of samples seen during the training. We use 100-150M samples per experiment. We can’t tell from BFL’s primary figure how many images they used in ImageNet training.
When we dropped PixelREPA from our JiT-DDT, the images were 5-10% worse. So, it looks like distillation from the auxiliary vision model remains helpful in low sample regimes.
Since we view JiT-DDT as a cousin of Self-Flow, we’d like to eventually train our architecture on 10x more samples with/without PixelREPA and see if we can get better generations without the auxiliary vision encoder.
One more thing to call out is that there is a clear discrepancy in the effect the REPA has in pixel space versus VAE latent space. Plain REPA actively hurt our JiT. That’s why we switched to PixelREPA in the first place. We ablated whether to keep PixelREPA on for the entirety of training or switch it off midway (as is conventional wisdom). The results were a wash; PixelREPA’s masking op might be a regularizer helping us avoid overfitting to DINO space.
Boosting gradients early in the DiT accelerates learning
While Self-Flow finds a path forward without DINO alignment, others have gone the other way. In RAE v2, the authors achieve SOTA on ImageNet FID by training a DiT in DINO space.
Instead of using patches like our pixel space models or VAEs like BFL, they run DINOv3 on all of their images, summing together the hidden states across many layers of DINO to come up with a representation. They then train two independent models, the flow matching generative model and a decoder from DINO space back to pixel space.
If you’re training in DINO space already, it’d be logical to axe out REPA. But turns out, it still unlocks better generations in RAE v2. Let’s pause for a second. That’s really weird.
The authors find that REPA reduces to x-prediction within RAE v2, because the DiT’s latent space and the alignment loss are both derived from DINO. Obviously, this rhymes with our JiT-DDT; we’re also doing x-prediction early in our DiT via our encoder. But, we think this points at a deeper point — the DiT has a gradient propagation problem.
Self-Flow in latent space, RAE v2 in DINO space, and JiT-DDT in pixel space all improve model performance by introducing a loss term earlier in the network. It seems like all these models need additional gradient highways to learn more effectively. We’re actively digging into this and will report back on this soon.
Appendix
Below are side-by-side comparisons of Linum v2 and JiT-DDT on 26 different prompts. All images generated by Linum v2 are 256×256 and all JiT-DDT images are 512×512. A few things stand out:
Three golden-brown croissants rest in a row on a rustic wooden plate, the plate itself sitting atop a folded beige cloth napkin on a weathered wooden table, viewed from a slight high angle. Each flaky crescent is dusted with a sprinkle of vibrant red pepper flakes. Natural side light rakes across the laminated layers to emphasize their crisp, buttery texture. A shallow depth of field softens the table edge and background.
02 / 26
Linum v2
JiT-DDT
A golden retriever jumps out of fresh powder snow in a sunlit forest clearing, centered in the frame with its ears perked up. Fine snow dusts its paws and catches the low morning light. Tall evergreen and bare deciduous trees ring the clearing in the background, softened by a shallow depth of field. The crisp, high-key winter light leaves the brightest snow overexposed.
03 / 26
Linum v2
JiT-DDT
A close-up portrait of a young white woman with vibrant, fiery red hair cascading over her shoulders in soft waves, framed from the shoulders up and centered against a softly blurred warm-toned background. Her fair, lightly freckled complexion sets off piercing green eyes and a subtle, closed-lipped smile. Soft natural light enters from the left of the frame, highlighting the texture of her hair and the curve of her cheek while leaving the right side in gentle shadow. A shallow depth of field renders the background into smooth, neutral bokeh. Lights dangle out of focus on the left side of the frame.
04 / 26
Linum v2
JiT-DDT
A 3D animated image of a small rounded robot with big expressive blue eyes standing near a potted sunflower on a windowsill, centered in the frame and rendered with soft subsurface lighting. Its dented metal body has a cheerful yellow paint job with scuffs. Warm morning light streams through the window behind it, casting a gentle glow on the leaves. The background kitchen is softly blurred
05 / 26
Linum v2
JiT-DDT
An antique wooden globe on a brass stand sits on a leather-topped desk in a dim study, positioned slightly left of center, its aged map showing faded oceans and hand-drawn continents. A green banker’s lamp on the right casts warm light across the globe and a stack of leather-bound books. A window behind shows a rainy gray evening. The rich browns and greens give the scene a scholarly, quiet mood.
06 / 26
Linum v2
JiT-DDT
A Bengal tiger wades through a shallow jungle river, its orange and black striped body centered in the frame and water splashing around its chest as it moves toward the viewer. Dense green foliage and hanging vines line the riverbanks. Dappled sunlight through the canopy throws bright spots across the water and the tiger’s wet fur. Its amber eyes are fixed directly ahead in sharp focus.
07 / 26
Linum v2
JiT-DDT
A three-tier chocolate birthday cake covered in glossy ganache and topped with a ring of lit rainbow candles sits centered on a white marble table in a dark room. The candle flames cast a warm, flickering glow across the cake and the scattered confetti below. Fresh raspberries and mint leaves decorate the edges. The background is nearly black, making the flames and dripping ganache the focal point.
08 / 26
Linum v2
JiT-DDT
A close-up portrait of a sweat-drenched boxer with wrapped hands raised in a guard, framed from the shoulders up and centered against the dark ropes of a gym ring. A single hard light from the upper right carves sharp highlights along his brow and cheekbones and leaves the left side of his face in deep shadow. Beads of sweat catch the light. The background of hanging heavy bags is nearly black and completely out of focus.
09 / 26
Linum v2
JiT-DDT
An expressive charcoal drawing of a galloping horse in profile moving from right to left, its mane and tail rendered in loose, energetic strokes on textured white paper. Heavy blacks define the body and legs while smudged gray tones suggest dust and motion around the hooves. The background is left mostly untouched with a few sweeping gestural marks. The drawing feels raw and immediate, with visible finger smudges.
10 / 26
Linum v2
JiT-DDT
A middle-aged Asian chef in a white double-breasted jacket tosses vegetables in a flaming wok, positioned center-left in a busy stainless-steel restaurant kitchen. Orange flames leap from the pan and light his focused face from below, while cool blue fluorescent light fills the background. Steam and small sparks scatter across the frame. Shot at a slight low angle with a fast shutter that freezes the tumbling vegetables mid-air.
11 / 26
Linum v2
JiT-DDT
Sweeping orange sand dunes stretch to the horizon under a clear sky at sunrise, with a sharp crest running diagonally from the lower left to the upper right of the frame. Low sunlight from the right carves the dunes into bright ridges and deep purple shadows. Fine wind-blown ripples texture the sand in the foreground. A lone line of camel tracks curves over the nearest dune toward the distance.
12 / 26
Linum v2
JiT-DDT
A weathered elderly fisherman with a white beard and deep sun-creased wrinkles sits on the edge of a wooden dock, framed from the chest up and slightly right of center. He wears a faded navy knit cap and a yellow oilskin jacket and looks off to the left with pale blue eyes. Golden late-afternoon light rakes across his face from the left, catching the texture of his skin and beard. A calm harbor with moored boats blurs softly into the background.
13 / 26
Linum v2
JiT-DDT
A close-up of dark espresso pouring from a chrome portafilter into a small white ceramic cup, centered in the frame, with a thick golden crema swirling on the surface. The stainless-steel espresso machine fills the background in soft focus. Warm cafe lighting from the upper left reflects in the chrome and the liquid stream. Small droplets are frozen mid-splash near the rim of the cup.
14 / 26
Linum v2
JiT-DDT
A smiling middle-aged woman in a green canvas apron arranges a bouquet of peonies and eucalyptus behind the counter of a small flower shop, framed from the waist up and centered. Buckets of tulips, roses, and sunflowers crowd the foreground in the lower third, and shelves of potted plants fill the background. Soft daylight from a storefront window on the right falls across her face and the pale pink petals. The depth of field is shallow, keeping her and the bouquet sharp.
15 / 26
Linum v2
JiT-DDT
A shaggy Highland cow with long ginger hair covering its eyes and wide curved horns stands in a misty green Scottish field, framed from the chest up and centered, looking toward the viewer. Soft overcast light gives the fur a warm, tactile quality. Rolling hills and a low stone wall fade into fog behind it. Dew glistens on the grass in the foreground and the cow’s wet nose catches a small highlight.
16 / 26
Linum v2
JiT-DDT
Dozens of colorful hot-air balloons in stripes of red, yellow, blue, and green drift over a misty valley at dawn, with the largest balloon filling the upper left of the frame and the others scattered toward the horizon. Low sunlight from the right rims the balloons in warm gold. Rounded rock formations and green fields sit below, partly hidden by mist. The sky is a pale wash of peach and lavender.
17 / 26
Linum v2
JiT-DDT
A Roman-style mosaic of a large fish swimming to the left, composed of thousands of small tesserae tiles in blues, greens, gold, and terracotta, filling the frame against a background of pale stone tiles. The fish’s scales are picked out with alternating light and dark tiles, and a wavy band of blue tiles runs along the bottom. Grout lines and slight irregularities in the tile edges are visible. Even, diffuse light shows the surface texture.
18 / 26
Linum v2
JiT-DDT
A close-up of a mottled orange octopus draped over a coral outcrop, its curling arms and suckers filling the lower half of the frame as it looks toward the viewer with a golden eye. Deep blue water and small drifting particles fill the background. Dappled sunlight from the surface above casts shifting light patterns across its textured skin. Small purple sea fans and yellow coral polyps frame the edges.
19 / 26
Linum v2
JiT-DDT
A dramatic oil painting in the style of nineteenth-century marine art depicting a three-masted sailing ship heeling in a violent storm, positioned center-left as enormous green-gray waves crest around it. Torn sails and rigging strain in the wind, and a break in the dark clouds on the upper right lets a shaft of pale light fall on the foam. Thick impasto brushstrokes render the spray and cloud. The palette is deep blues, grays, and bone whites.
20 / 26
Linum v2
JiT-DDT
A black vintage typewriter with round chrome-rimmed keys sits on a wooden desk with a sheet of white paper rolled into its carriage, centered and photographed straight on from a slight high angle. A single line of typed text is visible on the page. Soft window light from the right highlights the keys and casts gentle shadows between them. A small brass lamp and a stack of books blur in the background.
21 / 26
Linum v2
JiT-DDT
A detailed graphite pencil sketch of an old man’s face in three-quarter view, framed from the shoulders up and centered on cream-colored paper. Fine cross-hatching builds up the deep wrinkles around his eyes and the texture of his short beard, while lighter strokes suggest a flat cap. The left side of the face is shaded in darker tones and the right fades into untouched paper. A few faint construction lines remain visible near the edges.
22 / 26
Linum v2
JiT-DDT
A steaming bowl of tonkotsu ramen sits centered on a dark wooden counter, viewed from a slight high angle, with slices of pork belly, a soft-boiled egg halved to show its orange yolk, green onions, and a sheet of nori arranged on top of the noodles. Wooden chopsticks rest across the rim on the right. Steam rises and catches a warm overhead light. The cloudy broth and glistening pork are rendered in sharp, appetizing detail.
23 / 26
Linum v2
JiT-DDT
A vintage red bicycle with a wicker basket of fresh baguettes leans against a sun-bleached yellow plaster wall, positioned center-right in the frame. A green wooden window shutter and a small pot of red geraniums sit on the sill above the bike on the left. Bright afternoon sunlight from the right casts a crisp shadow of the bicycle across the wall and the cobblestones. The colors are warm and saturated.
24 / 26
Linum v2
JiT-DDT
A bearded street musician in a brown corduroy jacket plays a saxophone beneath a dripping awning on a rainy city street at dusk, positioned right of center. Neon signs in pink and teal reflect in the wet pavement in the lower half of the frame, and blurred pedestrians with umbrellas pass on the left. Cool ambient light mixes with a warm glow from a shop window behind him. Raindrops streak across the foreground.
25 / 26
Linum v2
JiT-DDT
A macro photograph of a bright green tree frog with orange toes and red eyes perched on a floating lily pad, centered in the frame and viewed at water level. Dew drops bead on its glossy skin and the leaf’s surface. Soft, diffused morning light gives an even glow, and the pond behind dissolves into a smooth green and blue bokeh. A single pink water lily blooms softly out of focus in the upper right.
26 / 26
Linum v2
JiT-DDT
A vintage silver and black film camera rests on a scuffed oak desk beside a stack of faded photographs and a coiled leather strap, centered in the frame and viewed from a slight high angle. Warm afternoon light from a window on the left glints along the chrome dials and lens ring. A half-empty cup of black coffee sits out of focus in the upper right. The shallow depth of field keeps the lens crisp while the desk edge softens.
Authorship statement
We wrote all the words on this page. We used Claude Fable 5.1 to help us build the diagrams.
The Huggingface Model Card and Github Repo were written automatically by Claude Fable 5.1. We pointed Claude to our internal, experiment repo and had it pull out (and clean up) the necessary code.
Who are we?
We’re two brothers training text-to-video models from scratch, trying to make animation accessible to everyone.
Get Field Notes
Technical deep dives on building generative video models from the ground up, plus updates on new releases from Linum.
Starting today, Claude Cowork and chat are merging into one Claude. Bring a quick question, or hand over a report due at noon, and Claude takes it from there, even after you’ve closed your laptop. It’s rolling out on Pro and Max plans over the next few weeks, with more plans to follow.
What Claude makes doesn’t need its own place either. Claude Docs and Claude Slides are new today, and Claude Design now works inside your conversations too. Ask for a document, and you and Claude write it together. Ask for a presentation, and Claude drafts the slides. You can edit directly, present straight from Claude, or download as PowerPoint or PDF. All three are in beta on paid plans, and Enterprise admins choose when to turn them on. If you use Claude Design on its own, it keeps working as before.
We built Cowork as a separate place for bigger work, and Design for visual work. People used both, and told us the frustrating part was deciding where a task belonged. What they’d started in one also didn’t carry into the other. So we stopped making you choose. Claude can now figure out what a task needs, so what Cowork and Design can do is available from any conversation, with the context, skills, and connectors you already have.
“I could have Claude pull up [my legal research database], and it would pull all the cases, read them, figure out which other cases I might need, download them, and store them in a folder for my personal review.” – Andrew Keller, Senior Economist
What it looks like
A weekly report is due at noon. Before you head out, you ask Claude what moved in the pipeline last week, then add: “Write it the way we always do, flag anything that slipped, and put the highlights in five slides for the leadership meeting.” If something’s unclear, Claude asks. You can check progress from your phone on the way to the office. By the time you’re at your desk, the report is waiting as a doc with notes from your teammate, and the slides are ready to open, adjust, and download as PowerPoint. Both came out of the same conversation, so the slides already match the report. You fix a line yourself, leave a comment for Claude on a slide, and share it. Schedule it for every Monday, and Claude starts on the report without being asked.
Anything you make with Claude Design, Slides, or Docs lives at one shareable link you can open on your phone. You can select an element and move it, or tell Claude what you want changed.
You can choose how Claude checks in with you. By default, Claude asks before taking an action. If you’d rather let it keep working and check in only when something needs a closer look, you can turn that on. You keep the final say.
Getting started
If you mostly use chat, you don’t have to do anything different. When you want to hand over something bigger, try the next report or deck you’d normally build yourself. If you’ve been working in Cowork, everything is where you left it: your chats, projects, artifacts, connectors, and skills. When you open the app, pick up where you left off.
This is rolling out to Pro and Max plans first, in the Claude app on web, desktop, and mobile over the coming weeks to existing and new users on these plans. There’s nothing to turn on. Team and Free plans will follow soon, and Enterprise admins will hear from us at least 30 days before anything changes for their organizations.
The full list of capabilities is in the Help Center. If you’ve been saving up a big messy project, now’s the time.
[1] button Change ticket type · Round trip
[2] combobox Where from? · San Francisco
[3] combobox Where to? · empty
[4] textbox Departure · empty
...
The operations are CLICK, TYPE_TEXT, SELECT, SCROLL_UP, SCROLL_DOWN, WAIT, DONE, and BLOCKED. Only supported operations and targets are offered.
one TypeSafe request
┌───────────────────────────┐
page → element table → operation │
│ click_target │
│ type_text_target │
│ select_target, if present │
└─────────────┬─────────────┘
use the matching target
│
CLICK [7] ─────┤──→ browser
TYPE_TEXT [3] ─────┘
↓
small LLM → text → browser
Target questions are speculative. If the operation is CLICK, only click_target can execute. Two decisions, one network round trip. Each target head contains only compatible elements. Native dropdown choices carry an observed element/option index.
There are no site-specific action scripts or prepared field strings in the policy. The Flights example supplies a goal and independently verifies the outcome. The screenshot renderer adds labels afterward; it does not drive the browser.
Try it
git clone https://github.com/browser-use/jev-ultrafast.git
cd jev-ultrafast
uv sync
cp .env.example .env
# Add TYPESAFE_API_KEY and TEXT_MODEL_API_KEY.
uv run jev
Open http://127.0.0.1:8766 and click Start demo → Run automatically. The inspector shows numbered elements, operation probabilities, target probabilities, and executed actions. Choose next pauses before execution.
Chrome connects through Browser Harness, installed by uv sync. Run uv run browser-harness --doctor if it needs connecting. Allow remote debugging in Chrome when prompted.
TEXT_MODEL_API_KEY is an OpenRouter key in the example configuration. The current demo uses inception/mercury-2.5 with reasoning disabled. Gemini, GLM, and DeepSeek can also use the OpenAI-compatible text helper; configure the appropriate model, endpoint, and reasoning setting.
Use the library
fromjev_ultrafastimportAgentwithAgent(
"https://www.google.com/travel/flights?hl=en",
"Find one-way flights from Zurich to London on September 20, 2026, ""for one adult in economy. Stop when matching flight options are visible.",
) asagent:
forstateinagent.run():
print(state["elapsed_ms"], state["status"])
Run with uv run --env-file .env python your_script.py. The same policy can run a different task:
uv run --env-file .env python examples/run.py
--url https://en.wikipedia.org/wiki/Main_Page
--goal 'Find and open the Wikipedia article about Gödel’s incompleteness theorems.'
uv run --env-file .env python examples/flights.py --keep-open performs the flight search, checks the actual route/date/results, and saves its trace. It does not select or book a flight.
Why it moves
One request per decision cycle. Operation and target heads share the same observed state.
No screenshots in the default agent loop. Jev consumes structured state. The inspector opts into screenshots; the video uses a separate continuous screencast.
One browser call per snapshot. Read visible controls, their names, values, and text atomically. Keep references to the actual DOM nodes.
Validate the selected target. Clicks check the document, form values, target, and nearby context. Animation alone does not force another prediction. Resolve current geometry and reject covered controls before input.
Wait for useful state. After typing into a combobox, wait for visible suggestions, capped at 200 ms. Other interactions get at most two animation frames or 50 ms. These reads happen after execution is logged.
Send visible text. Offscreen article bodies and footers do not fill the model context.
Reuse an interrupted text request. A generated value survives a stale-page retry only if the entire text-helper input is unchanged.
Every executed target is resolved from an observed node. The executor rechecks page freshness and click occlusion. Model output never becomes selectors, coordinates, shell commands, or executable JavaScript. Text-helper output must parse as a small JSON object before typing.
The current video is a 7,073 ms Google Flights run. Timing starts after initial page observation and includes model calls, generated text, browser work, stale decisions, and loading waits. A fresh independent check verifies the one-way setting, Zürich, London, September 20, 2026, and visible flight options. The video plays at 1×, with no opening hold and a 0.5-second final hold.
In six alternating runs with identical models and settings, both versions passed 3/3. Median task time went from 9.450 s → 7.092 s, a 25% reduction; median browser protocol calls went from 1,092 → 101. This is three repeats of one task on one browser profile, not a general reliability benchmark.
The same policy opened the requested Wikipedia article in 2.798 s and passed a local hotel search/filter task in 1.896 s. Runs, failures, source hashes, and measurement boundaries are in performance.md.
A DONE choice still requires independent outcome verification. The DOM reader handles common HTML and ARIA controls, not the full accessible-name specification. Shadow roots, frames, canvas, uploads, pop-up tabs, nested scrolling, and arbitrary keyboard widgets remain outside this MVP. Owned tabs share the existing Chrome profile.
Development
uv run ruff check .
uv run pytest
node --check jev_ultrafast/static/app.js
node --check jev_ultrafast/snapshot.js
uv build
Tests are offline. uv run python scripts/check_guards.py checks real controls in a local browser without model calls. Live examples and recording scripts make paid API calls. scripts/record_flights.py <new-folder> captures original browser timestamps; scripts/render_demo.py <recording-folder> renders that verified run at 1× and crops out the Google account strip. Credentials and raw traces stay ignored.
The spec framework for building the right thing and building it right
Synopsis
OpenSpec is a lightweight and configurable framework for creating and managing software
specifications.
With OpenSpec, you capture what you want to build in a spec and keep your team and coding
agents aligned as the work evolves. We help you refine the requirements, validate that they
describe the right thing, and verify that the implementation matches.
Abstract:Ternary Large Language Models (LLM) store every weight as one of three symbols ${-1,0,+1}$, so the cost of a ternary model is conventionally referenced to the information-theoretic $log_2 3 approx 1.585$ bits per weight. The prevailing deployment format packs five ternary weights into one byte (five-trit packing), and due to the power-of-two group sizes used in practice this rounds up to $1.625$ bits per weight. This effective storage bit-width treats the three symbols ${-1,0,+1}$ as equiprobable. We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to $51.5%$ of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layout comprised of a dense presence bitmap plus a compacted sign vector, and costs $2 – z$ bits per weight element given a zero density $z$ in the model’s weights. BITCOS stores weights more compactly than the five-trit packing in 26 of the 29 tested models, and reaches $1.485$ bits per weight on the sparsest of them. BITCOS is amenable to efficient unpacking on modern processors and GPUs, and we present optimized unpacking sequences for AVX-512, AVX2 and Intel Xe2 GPUs. Measured against production state-of-the-art ternary matrix-vector multiplication kernels, at the zero densities real-world ternary models exhibit, the realized gain with our proposed layout is up to $1.28times$. Finally, we illustrate end-to-end LLM inference results on 5 different platforms (client and server CPUs, integrated and discrete Xe2 GPUs) where decode throughput improves by up to $1.18times$ on CPUs and $1.27times$ on GPUs.
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Cyclomatic Complexity (or CC) in C# is a code metric that counts the number of linearly independent execution paths through a method. Concretely, it is computed as 1 plus the number of branching constructs in the method body (such as if, while, for, case, &&, ||, ?: and ??). The higher the score, the harder the method is to read, test and safely change. A score of 1 means a single straight path, around 10 is the traditional upper bound recommended by Thomas McCabe, and anything above 25 is flagged as excessive by Microsoft’s CA1502 analyzer.
This guide explains, with C# examples, how Cyclomatic Complexity is calculated, what thresholds matter in practice, how to measure and visualize it in real .NET codebases, and how to go beyond the raw score by pairing it with test coverage and IL-level analysis.
Cyclomatic Complexity was introduced by Thomas J. McCabe in 1976 as a way to quantify the structural complexity of a piece of code. The idea comes from graph theory: every method can be represented as a control flow graph where nodes are blocks of statements and edges are jumps between them. On that graph, the Cyclomatic Complexity is given by the classic formula:
1
M=E–N+2P
where E is the number of edges, N the number of nodes, and P the number of connected components. For a regular method with a single entry and a single exit, this collapses to 1 + the number of decision points, which is the form most tools actually compute.
What this number really tells you is the minimum number of test cases you need to exercise every independent path through the method. That is why Cyclomatic Complexity has stuck around for almost half a century: it is a structural metric, but it has a very concrete operational meaning for everyone who has to maintain or test the code.
The following expressions are not counted for CC computation:
1
2
3
4
5
elsedoswitchtryusing
throwfinallyreturn
objectcreation
method call
field access
Two details that trip people up: else does not increment the score because the alternative path was already created by its matching if; and a switch contributes one unit per case (and one for default), not one for the switch keyword itself. C# pattern-matching constructs (and, or, the modern switch expression with patterns) also add to the score, the same way their classic counterparts do.
Example of Cyclomatic Complexity Impact in C#
Exhibiting a Complex Method
Here is a complex method with entangled if and else scopes. The keyword if is used six times and && is used once. Hence its Cyclomatic Complexity score is 8:
C
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
publicstaticclassOrderLogic{
publicstaticvoidProcessOrder(
intorderId,
boolisPriority,
boolisInternational,
boolisGift,
boolisCouponApplied,
decimal orderTotal){
if(orderId<=0){
Console.WriteLine(“Invalid order ID.”);
return;
}
if(isPriority){
Console.WriteLine(“Processing priority order.”);
if(isInternational){
Console.WriteLine(“Processing international priority order.”);
if(isGift){
Console.WriteLine(“This is a gift order.”);
}
}
}else{
Console.WriteLine(“Processing standard order.”);
if(isInternational){
Console.WriteLine(“Processing international standard order.”);
}
}
if(isCouponApplied&&orderTotal>100){
Console.WriteLine(“Applying discount for orders over $100.”);
}else{
Console.WriteLine(“No discount applicable.”);
}
}
}
Eight independent paths means at least eight tests to fully cover this single method, plus a non-trivial amount of head-scratching every time someone has to add a new business rule. This is exactly the kind of method where a regression slips in unnoticed.
Refactoring the Complex Method in Several Simpler Methods
The method above can be refactored into several less complex methods. In the code we use CC to refer to each method Cyclomatic Complexity score:
Console.WriteLine(“Processing international priority order.”);
}
}else{
Console.WriteLine(“Processing standard order.”);
if(isInternational){
Console.WriteLine(“Processing international standard order.”);
}
}
if(isGift){
Console.WriteLine(“This is a gift order.”);
}
}
privatestaticvoidApplyDiscountIfEligible(// CC 3
boolisCouponApplied,
decimal orderTotal){
if(isCouponApplied&&orderTotal>100){
Console.WriteLine(“Applying discount for orders over $100.”);
}else{
Console.WriteLine(“No discount applicable.”);
}
}
}
Benefits of Refactoring
Simpler Control Flow: The main method now delegates specific tasks to smaller, more focused methods.
Easier to Test: You can test each smaller method independently.
Lower Cyclomatic Complexity: The complexity is spread across multiple methods, making each method easier to understand and maintain independently.
Better Naming: Method names like IsValidOrder or ApplyDiscountIfEligible document intent, so a reader does not have to mentally simulate the body to understand the high-level flow.
Note that the total Cyclomatic Complexity summed across the four methods is actually slightly higher than the original 8. That is fine and even expected. What matters for maintainability is the complexity per method, because that is the unit a developer has to reason about at a time.
Cyclomatic Complexity Thresholds: What Score Is Too High?
There is no single sacred number, but the literature converges around the same ranges. The table below summarises what most teams and tools use as a guideline:
Cyclomatic Complexity
Risk profile
Practical interpretation
1 – 10
Simple, low risk
McCabe’s original recommendation. Easy to test, easy to read.
11 – 20
Moderately complex
Still manageable, but worth a second pair of eyes during review.
21 – 50
Complex, high risk
Hard to test exhaustively. Strong refactoring candidate.
> 50
Untestable
Bug magnets. Often legacy hotspots that need to be broken down.
Two reference points are worth keeping in mind. McCabe himself recommended splitting modules that exceed a Cyclomatic Complexity of 10. Microsoft’s CA1502 analyzer defines “excessive complexity” as a score greater than 25 by default. Mark Seemann argues for an even tighter ceiling of around 7, mirroring Miller’s “magical number seven, plus or minus two” for human short-term memory.
In practice the right threshold depends on the codebase. A parser, a serializer or a state machine will routinely live in the 15-25 range without being objectively bad. A piece of business logic that scores 25 almost always is.
Measuring Cyclomatic Complexity in C#
You can’t improve what you don’t measure, so using a tool to evaluate code complexity is essential. Calculating this metric helps developers identify areas that might need refactoring to improve code quality.
NDepend is a great option for this, as it measures the cyclomatic complexity of methods in C# code. For instance, it includes the Search Methods by Complexity feature, which helps identify complex methods for further analysis.
Visual Studio itself ships a “Calculate Code Metrics” command (Analyze > Calculate Code Metrics) that reports Cyclomatic Complexity per method, type and assembly. The CA1502 analyzer can be wired into your build to actually fail on methods above a configured threshold, which is useful for new code. Roslyn-based analyzers like SonarAnalyzer.CSharp and third-party tools such as ReSharper or CodeRush also surface the same metric inside the editor.
Ruling C# Cyclomatic Complexity
NDepend offers several rules like Avoid methods too big, too complex that flag methods with excessively high Cyclomatic Complexity scores, highlighting potential issues in the code.
You are probably working with a large legacy codebase, making it impractical to refactor every complex method. This is why it’s essential to measure Cyclomatic Complexity against a baseline, allowing you to focus on new or refactored methods that are too complex. There are two rules for that:
This baseline-driven approach matters more than any absolute threshold. When you start tracking complexity on a well-established codebase, you will inevitably find complex methods that have been stable and well-tested for years. The real risk is not the static score, it is what happens when those methods start growing. Using NDepend’s CQLinq, you can express that idea directly:
The query warns whenever an already-complex method (CC > 10) becomes even more complex between two analysis snapshots. In other words, it surfaces the modifications that are actually risky, while leaving stable legacy untouched.
Visualizing C# Cyclomatic Complexity
A colored treemap can be used to visualize the cyclomatic complexity of your C# methods. In this visualization, each rectangle represents a method:
The size of the rectangle corresponds to the number of statements in the method.
The color of the rectangle reflects the method’s cyclomatic complexity.
The advantage of a treemap over a flat list is that complexity hotspots literally jump out of the picture: a large, dark red rectangle inside an otherwise calm area is exactly the kind of method that deserves an architecture conversation.
C# Cyclomatic Complexity and Tests
Writing tests for your code is nowadays an essential practice for every professional C# developer. Typically, a test covers a single execution path, while the Cyclomatic Complexity score of a method represents the number of independent execution paths. Therefore, Cyclomatic Complexity provides a rough estimate of how many tests are required to fully test a method.
By running tests, you can determine the code coverage for each method. A method partially covered means that not all its independent execution paths are challenged by tests. The rule Methods should have a low C.R.A.P score spots methods that both have high Cyclomatic Complexity scores and are poorly tested (C.R.A.P stands for Change Risk Analyzer and Predictor). The matched methods clearly indicate pain points in your code and should be tested and refactored.
The CRAP score defines a specific mathematical formula to combine complexity and coverage. Expressed in CQLinq, that formula is:
C
1
2
3
4
5
6
7
8
// <Name>CRAP</Name>
fromminJustMyCode.Methods
let CC=m.CyclomaticComplexity
let uncov=(100–m.PercentageCoverage)/100f
let CRAP=(CC*CC*uncov*uncov*uncov)+CC
selectnew{m,CRAP}
The CRAP score scales with the square of complexity and the cube of uncovered percentage, then adds the raw complexity as a floor. The practical takeaway: a method with CC = 30 and 100% coverage is far less of a liability than a method with CC = 12 and 0% coverage. The second data point of coverage shows that not all complexity is created equal.
Going Beyond Cyclomatic Complexity
Cyclomatic Complexity is a great start for reasoning about your code’s complexity, not the end of the conversation. Two extensions are worth knowing about.
Pair Complexity with Branch Coverage
If you have two methods with the same Cyclomatic Complexity, and one is fully covered by branch tests while the other has none, the risk profile is wildly different. A CQLinq query that captures this idea:
It raises a warning for any complex method without total branch coverage, which is a much more nuanced view of risk than “this method has CC > 10”.
IL Cyclomatic Complexity for Third-Party Code
Any non-trivial C# project drags in third-party libraries. Most of the time you treat them as black boxes; sometimes you regret it. NDepend’s CQLinq offers a property called IL Cyclomatic Complexity that applies the same metric to the .NET Intermediate Language inside the DLLs you depend on.
This lets you measure how complex the methods of a candidate library actually are, not just how nicely their public API is documented. If you analyze a dependency and find that its internal methods are routinely above CC = 30 in IL, that is a strong signal it will be hard to debug when something goes wrong inside it.
How to Reduce Cyclomatic Complexity in C#
Refactoring for lower Cyclomatic Complexity is mostly about extracting decisions out of the hot path. The techniques below come up over and over again on real codebases.
1. Extract Method
The most boring and the most effective. Pull a contiguous block of decisions into a well-named method, exactly like in the ProcessOrder example earlier. Each call site loses one chunk of complexity; each new method is small enough to test on its own.
2. Early Return / Guard Clauses
Instead of deeply nesting validation in if blocks, return early on invalid input. The cyclomatic count is the same, but the cognitive load drops sharply because every following line can assume the input is valid.
C
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
// Nested
publicdecimal CalculateFee(Order order){
if(order!=null){
if(order.IsValid){
if(order.Customer!=null){
returnorder.Total*0.05m;
}
}
}
return0;
}
// Guarded
publicdecimal CalculateFee(Order order){
if(order isnull)return0;
if(!order.IsValid)return0;
if(order.Customer isnull)return0;
returnorder.Total*0.05m;
}
3. Replace Conditionals with Polymorphism
A long switch on a type discriminator is a classic smell. Each case adds 1 to the score and the method becomes the only place that needs to change whenever a new variant appears. Moving the per-variant behaviour into derived classes (Strategy, Template Method, or simple subclass overrides) typically collapses the central method to CC = 1.
4. Use Modern C# Pattern Matching and Switch Expressions
Switch expressions are still counted by analyzers, but they encourage flatter, side-effect-free code than chains of nested ifs. Combined with pattern matching, they often replace a CC of 8-10 with a single, declarative expression:
If a method is mostly a list of “input X maps to output Y”, a Dictionary<TKey, TValue> or a static readonly array is almost always preferable to a chain of ifs. The dictionary lookup is CC = 1, regardless of how many entries it holds.
6. Boolean Parameters Are a Code Smell
A method that takes several bool flags almost always hides several methods in a trench coat. Splitting SendEmail(bool html, bool urgent, bool dryRun) into more specific methods reduces both the per-method complexity and the chance that a caller passes the wrong combination of flags.
Frequently Asked Questions
What is a good Cyclomatic Complexity score in C#?
Below 10 is considered safe, between 10 and 20 needs attention, above 25 is what Microsoft’s CA1502 rule flags as excessive. McCabe’s original recommendation, still widely quoted, is to refactor any method that exceeds 10.
Does the else keyword increase Cyclomatic Complexity?
No. The else branch is already implied by its matching if, so it does not add a new independent path. Only the if itself counts.
Does a switch statement count once or per case?
Per case (and per default). The switch keyword itself does not increment the score. A switch with 6 cases plus a default contributes 7 to the method’s Cyclomatic Complexity.
Does Cyclomatic Complexity equal the number of unit tests I need?
It is a useful lower bound, not a contract. Cyclomatic Complexity gives you the number of linearly independent paths, which is the minimum number of tests required for full path coverage. In practice you may need fewer (some paths are infeasible) or more (data combinations within a single path can still misbehave).
What is the difference between Cyclomatic Complexity and Cognitive Complexity?
Cyclomatic Complexity measures the number of paths. Cognitive Complexity, popularised by SonarSource, also weighs nesting depth and boolean operator combinations because they make code harder to read even when they do not add new paths. The two metrics are complementary: low cyclomatic, high cognitive is rare; low cognitive, high cyclomatic is also rare; methods that are bad on one are usually bad on the other.
How do I measure Cyclomatic Complexity in Visual Studio?
In Visual Studio, go to Analyze > Calculate Code Metrics > For Solution. The resulting window lists Cyclomatic Complexity per method, type and project. For continuous enforcement, enable the CA1502 analyzer and configure its threshold via a CodeMetricsConfig.txt additional file.
Does Cyclomatic Complexity work on async or LINQ code?
Yes. The C# compiler rewrites async/await into a state machine, but most tools (NDepend, the Roslyn analyzers, Visual Studio’s metrics) compute Cyclomatic Complexity on the source-level method. await itself does not count, but the if, while and catch constructs surrounding it do. LINQ query operators are method calls, which do not count either, but lambdas passed to them are analyzed as their own methods.
Conclusion
Cyclomatic Complexity is one of the oldest code metrics still in active use, and the reason is simple: it captures something that maps directly to real-world pain. A high score means more paths to reason about, more tests to write, more places where a bug can hide. Keeping it under control, especially on the methods that change the most, pays for itself within weeks.
That said, the score on its own is half the picture. A complex but heavily tested method is rarely the one that wakes you up at night; a moderately complex method with zero coverage that gets touched every sprint will. Pair Cyclomatic Complexity with coverage (or with the CRAP score), enforce a delta-based rule rather than a flat threshold on legacy code, and use IL Cyclomatic Complexity to sanity-check the libraries you depend on. Combined, these techniques turn a 1970s metric into a surprisingly modern early-warning system for technical debt.
If you want to try this on your own codebase, download a free trial of NDepend and run an analysis: the complexity hotspots usually become obvious within the first few minutes.
This article is brought to you by the team behind NDepend — a proven .NET static analysis tool for improving code maintainability, security, and overall quality. Whether you’re modernizing a legacy .NET application or starting fresh in C#, get started with your free full-featured trial today!