Learn/Memoryrusty_allocrusty_zstd

How rusty_alloc optimized large memory allocation, and what that means for rusty_zstd at L1 and L3

rusty_alloc made a 2 MiB alloc and free 57 percent cheaper. rusty_zstd 0.3.0 is faster at L1 and L3, and that release measured the allocator change as neutral.

Signed by Tim Almond·
Painted dusk press room where one wide oak crate rests on a long shelf lined with soft mesh lights, and two smaller hand presses wait beside folded paper
Did rusty_alloc make rusty_zstd 43 percent faster at L1 and L3?

No. rusty_zstd 0.3.0, released 2026-10-08, encodes in 0.61 times the time of 0.2.5 at L1 and 0.72 times at L3. That release moved the CLI to rusty_alloc 2.2.5 and measured the allocator swap as neutral, 0.98 to 1.00 times. The large allocation win is inside rusty_alloc: a 2 MiB alloc plus free is 57 percent cheaper than in 2.2.0.

A large allocation and a compressor level are easy to fold into one sentence. They are two clocks. rusty_alloc made a large memory allocation cheaper inside the heap. rusty_zstd 0.3.0, cut on 2026-10-08, encodes faster at L1 and L3 because the encoder changed. The changelog for that compressor release says the move onto a newer rusty_alloc measured neutral.

The source for the allocator is the rusty_alloc repository. The source for the compressor is the rusty_zstd repository. The crate page is rusty_zstd on crates.io. Read the two notes as a pair, and keep the percents attached to the bench that produced them.

The large allocation optimization inside rusty_alloc

rusty_alloc is the house allocator: a pure Rust global heap, the pure Rust alternative to mimalloc when a process wants that shape. Since 2.2.0, four rounds of instruction-count work landed in the ledger. The row that matters for a large allocation is the huge op: one 2 MiB alloc plus the free that returns it.

A 2 MiB alloc and free

The microbench is bench/opscan.c. Lower is better. Against 2.2.0 at commit c9631f4, a huge 2 MiB alloc plus free moved from 689.68 to 293.68, a 57 percent drop. A big 4 KiB alloc plus free moved from 121.00 to 92.00, a 24 percent drop. Mixed sizes fell 16 percent. A small 32 byte alloc plus free fell 4 percent. Cross-thread free fell 32 percent. These are allocator operations. They are not zstd levels, and they are not whole-program compressor time.

The same README records a Rust GlobalAlloc path the C instruments never exercise. Boxed trees, buffers, and Rc fell 21.1 percent in whole-program instructions versus 2.2.0. Cache-line-aligned buffers on realloc fell 27.3 percent. Hash maps fell 5.8 percent. Short-lived threads barely moved, 0.8 percent. A large allocation is one of those paths. A hash map of short strings is another. Quoting one percent for both is how the story goes wrong.

What the rusty_zstd check on that allocator proved

On 2026-09-25 the house ran real consumers, not only the counter. SpaceDB's workspace tests and rusty_zstd's tests ran on the newer allocator on Windows and Linux. The Silesia corpus went through the rusty_zstd CLI at levels 1, 3, 9, and 19, with 4 threads. Every file compressed byte-identically to the same CLI on rusty_alloc 2.2.0, and round-tripped through C zstd 1.5.7 both ways. That is the downstream claim: a large allocation change did not move the bytes. It is not a claim that L1 or L3 got faster.

If you have seen 43 percent next to this allocator, it is probably the other essay. A game memory allocator for developers reports battle frames, on a 60 fps budget, where the allocator alone took more than 1 ms. That count fell by 41 percent and 43 percent versus glibc in the battle scene. Useful, and a different instrument. It does not transfer to compressor levels.

What rusty_zstd changed at L1 and L3 today

rusty_zstd 0.3.0 is the pure Rust alternative to libzstd at this snapshot. The changelog titles the release in one line: encode is 1.8 times faster. The table under that line is the one to quote.

Encode time at L1 and L3

The standing board is six Silesia files: dickens, samba, x-ray, mozilla, nci, and xml. Levels are 1, 3, 5, 7, 9, and 12. Whole files, one core, best of five alternating rounds. Encode time against rusty_zstd 0.2.5:

| Level | L1 | L3 | L5 | L7 | L9 | L12 | geomean | |---|---|---|---|---|---|---|---| | Time vs 0.2.5 | 0.61x | 0.72x | 0.68x | 0.61x | 0.54x | 0.31x | 0.56x |

L1 at 0.61 times is about 39 percent less encode time on that board. L3 at 0.72 times is about 28 percent less. The geometric mean of 0.56 times is about 44 percent less encode time across all six levels, which is the "1.8 times faster" in the release note. A flat "43 percent at L1 and L3" is not what the table says. L1 and L3 are the two levels, and they are not the same percent.

Output at L1 through L3, and at L13 and above, is byte-identical to 0.2.5. L4 through L12 got smaller and still emit plain zstd frames. libzstd 1.5.7 decodes those frames byte for byte. Decoding in rusty_zstd did not change in this release.

The byte-identical speed sits in the encoder. L1 through L4 run Fast and DFast in their own scan loops. Literals go out in libzstd's Huffman shape. The sequence coder writes one merged extra-bits field. A dictionary is primed once per dictionary and parameter set, then restored per call. The row match finder, which changes output, starts at L4 and up. It is not the L1 and L3 result.

The allocator line in the same release

The same changelog has an alloc bullet. The CLI and the bench binaries moved from rusty_alloc 2.0.5 to rusty_alloc 2.2.5. Measured neutral: 0.98 to 1.00 times. The optional rusty-alloc feature still installs 1.1.6 through rusty_alloc_default 0.1.2. So the large allocation work and the L1 and L3 work shipped near each other, and the compressor release itself says the heap swap did not move its clock.

That is the split to keep. Large memory allocation got cheaper in the allocator's own bench. L1 and L3 got faster in the encoder's own bench. Putting the second percent on the first component invents a cause the measurement refused.

How to use both numbers without mixing them

A reader who wants one number will average these. The honest page keeps three clocks: the allocator op, the codec, and the CLI that wraps the codec with file IO.

CLI speed, codec speed, and the heap

The rusty_zstd README, remeasured for 0.3.0 on 2026-10-08, also compares the CLI with C zstd 1.5.7. Nineteen files, each capped at 8 MiB, 143,865,706 bytes in the staged corpus. Each arm has its own copy. Ratios are C time divided by ours, so above 1 times means rusty_zstd is faster. L1 encode lands at 1.31 to 1.58 times. L3 encode lands at 0.90 to 1.35 times. Size is a little larger: about 2.00 percent at L1 and 1.76 percent at L3 on the uncapped corpus.

Codec only, in process, against zstd -b on one thread, is a different table. There C still leads at L1 by about 6 to 9 percent on five of the six files, and x-ray is the exception at 1.27 times. At L3, C leads by 1 to 26 percent. The README prints that beside the CLI numbers so a slow C CLI, busy writing files, is not mistaken for a faster codec. Neither table is an allocator result. The alloc bullet already said the heap swap was neutral.

Where a large allocation still earns its keep

Neutral on one compressor board does not mean the large allocation work was wasted. A process that actually allocates 2 MiB blocks, or that spends time in GlobalAlloc on wide buffers, is the workload the 57 percent and the 27 percent describe. SpaceDB installs rusty_alloc by default so a memory bug aborts instead of quietly corrupting a replica. A library turns default features off so it does not force that choice on every downstream binary.

The same allocator is already under other products. MATA.NETWORK uses it as the unified allocator, and the field note is MATA.NETWORK in the wild. RAG Converter uses it in wasm, which is RAG Converter in the wild. dldeploy.com uses it against malloc and free on device, with a measured gain above 23 percent on ESP32, and the field note is dldeploy in the wild. Those are their benches. They are not the L1 and L3 table.

If you are choosing a heap for a compressor, start from the neutral result and then measure your own files. If you are choosing a heap for large buffers, start from the 2 MiB row. The map of the rest of the house is Vision.

FAQ

Quick answers for builders evaluating this technology.

What did the large allocation optimization in rusty_alloc measure?

On the allocator microbench, a 2 MiB alloc plus free fell from 689.68 to 293.68, 57 percent, against rusty_alloc 2.2.0. A 4 KiB alloc plus free fell 24 percent. Those are allocator operations, not compressor levels.

How much faster is rusty_zstd at L1 and L3?

Against rusty_zstd 0.2.5, encode time is 0.61 times at L1 and 0.72 times at L3 on the standing board. Across L1, L3, L5, L7, L9, and L12 the geometric mean is 0.56 times, which the changelog calls 1.8 times faster.

Did the rusty_alloc upgrade cause that L1 and L3 speed?

The 0.3.0 changelog says the CLI and bench moved from rusty_alloc 2.0.5 to 2.2.5 and that swap measured neutral, 0.98 to 1.00 times. The L1 and L3 speed is encoder work, and L1 through L3 output stays byte-identical to 0.2.5.

Where does the figure of 43 percent come from?

Not from this pair of releases. A geometric mean of 0.56 times is about 44 percent less encode time across six levels. A separate game-allocator essay uses 43 percent for battle frame spikes versus glibc. Neither number is an L1 and L3 allocator result.

What did the rusty_zstd check on the new allocator prove?

On 2026-09-25 the Silesia corpus through the rusty_zstd CLI at levels 1, 3, 9, and 19 matched the same CLI on rusty_alloc 2.2.0 byte for byte, and round-tripped through C zstd 1.5.7. That check is identity, not a speed delta.