Will DeepSeek V4.1 Flash fit on your machine?

DeepSeek V4.1 Flash ships as 510.3 GB across 48 files. Most of that never has to be in memory at once. Set your memory below and see which build fits, and what ends up on disk. Only the byte count matters here, so two machines with the same memory get the same answer.

You will see this called a 749B model and a 510 GB download, and both are right. The experts ship packed two 4-bit weights to a byte, so the 384 experts per layer occupy 296 GB of the 307 GB backbone, instead of the terabyte they would need at full precision. Parameter count is the wrong unit for this question. Bytes is the right one.

expert modulesattention and the restlookup tables (can live on disk)

There will be another one of these in a fortnight, and the sizing question starts over.

How much memory DeepSeek V4.1 Flash needs

510.3 GB on disk. 307.2 GB has to be resident as shipped, 150.8 GB at 2-bit. The remaining 203.1 GB is lookup tables and can stay on an SSD.

The lookup tables do not need to be in memory

Of the 510.3 GB, 203.1 GB is a pair of hashed n-gram lookup tables. The number looks fatal until you check how much of it gets read. From the released inference code: 48 rows per token at 264 bytes each, so 12.4 KiB per token. Next to the 4.51 GB of expert weights each token pulls in, the tables are rounding error.

So they are a latency cost, not a bandwidth cost, and they can sit on an SSD. Qwen3.8 has a similarly large static n-gram table, and on an M5 Max the difference is measurable: on an M5 Max under llama.cpp, leaving that table on SSD instead of in memory costs 0.9% of decode speed, 40.09 against 40.47 tok/s.

What the 510.3 GB is made of

PartAs shippedMust be in memory?
Expert modules384 per layer, 6 used per token, already 4-bit296.0 GBThe 6 per token do. The rest can page, but paging them is what makes it slow.
Attention, shared expert, embeddings, vision11.2 GBYes. Read on every token.
Lookup tablestwo hashed n-gram tables203.1 GBNo. 12.4 KiB read per token; SSD is fine.

Context memory is not the constraint

ContextMemory

890 bytes per token, so a full million-token context stays under a gigabyte. Plan around the file size.

How fast does it actually run?

Three published local measurements as of 18 September 2026. Other figures in circulation are either hosted API throughput or the previous model, V4 Flash, which has no lookup tables and is a third of the size.

It does not have to fit. The rows below include a 128 GB machine running the 4-bit build with the experts streamed from SSD, faster than the 256 GB machine that holds the whole backbone in memory. Fitting decides whether you can load it without a disk in the path; it does not decide whether it runs.

MachineBuildDecode

There will be another one of these in a fortnight, and the sizing question starts over. I do this for every open-weight release, from the shard headers, usually within hours of the weights landing, and they all land here.