Learn/Memorygame developersallocatorrusty_alloc

A Game Memory Allocator Swap, Measured on a Real Title

Game developers can measure a memory allocator swap on a real title: fewer frame spikes, a cheaper start-up, and a memory cost you can switch.

Signed by Tim Almond·
Painted hillside workshop at dusk where glass jars of different sizes wait on a wooden bench beside a quiet lake
What does a game memory allocator swap change on a real title?

On recorded Endless Sky traces, rusty_alloc reduced start-up allocator time and allocator-caused frame spikes versus glibc, showed no clear change in total battle cost, and used more peak memory unless purge delay was enabled.

A game memory allocator decides how often a frame waits on malloc and how much RAM a session keeps after the fight ends. We replayed the recorded allocation stream of Endless Sky, an open-source space game, through four allocators under repeated runs. The result is real and specific: fewer allocator-caused frame spikes and a cheaper start-up, at a memory cost. Game developers can run the same kit on their own client and dedicated server.

The crate under test is rusty_alloc. The program around it is The Importance of Remade With Rust. Allocators on the bench: glibc as the baseline, rusty_alloc 2.2, mimalloc 2.3, and jemalloc 5. Five repetitions each, rotated order, a null arm, and 95 percent bootstrap intervals. That set is the game memory allocator comparison a studio can defend. A row counts as a difference only when its whole interval clears the baseline and the null arm (glibc run twice) did not move.

What a game memory allocator swap changed

Two recordings came from the unmodified Endless Sky binary. Start-up made 2.4 million allocator calls on 25 threads. The first 60 seconds of a large battle made 12 million calls on 28 threads, with 816 MB live at peak. Each claim below showed up in at least two separate full runs on fresh virtual machines. A single-run result is not listed as a finding.

Fewer frame spikes at start-up and in battle

Start-up allocator time was 26 percent and 39 percent lower than glibc across those runs, 104 ms against 171 ms. On 60 fps frames (16.7 ms) where the allocator alone took more than 1 ms, the count fell by 30 percent and 49 percent at start-up, and by 41 percent and 43 percent in the battle (76 frames against 133). The worst 1 percent of battle frames, measured as allocator cost, was 24 percent and 31 percent lower.

Total battle cost did not move in a way the noise floor would sign. Allocator time plus the page faults of first writing new memory landed at 0.90x and 1.07x, and the interval includes 1.0. rusty_alloc spends less time inside the allocator. More of the page-fault cost of fresh memory lands on the game's first write. Counted together, the battle total matches glibc within noise.

A synthetic 20,000-frame session across level loads did show a lower worst-1-percent frame cost: 60 percent, 46 percent, and 59 percent lower across three runs. A synthetic job-system workload showed a win in one run that did not reproduce, so it is not a claim. On a synthetic 30 Hz server tick, rusty_alloc stayed between 0.5 and 1.4 seconds in all 15 runs while glibc ranged from 0.9 to 6.8 seconds. That spread is steadier. The median gap varied too much to quote as one percentage.

Memory is the price, and the switch trades it back

By default rusty_alloc keeps freed memory for reuse. On the battle it held about 403 MB more at peak than glibc (1,589 MB against 1,186 MB) and kept 1.6 GB after everything was freed. That is a policy, not a leak. With memory return on (RUSTY_ALLOC_PURGE_DELAY=10), battle peak memory fell to 283 MB below glibc. The cost was about four times as many allocator calls over 100 microseconds, and a 30 percent higher total. Game developers choose that switch per deployment, not as a hidden default.

How game developers check a game memory allocator on their own build

These figures cannot tell a studio what a much larger title would do on its own hardware. Endless Sky is a smaller game, the battle is a 60-second cut of one recording, and a replay measures the allocator's share of the work, not whole frame time. The original ask for this kit was a Star Citizen dedicated server and client. The method is the same for any studio that can record a build.

Record the live stream, then replay it

Start the unmodified binary under a small recorder (LD_PRELOAD, under 200 lines of readable C). It logs each allocation call and passes it straight through. One command then replays that stream through every allocator, checks each one for correctness, and times it. An overnight run on an idle machine is enough. Recorded traces hold block ids, sizes, threads, and timestamps. They contain no addresses and no memory contents. What leaves the building is the report, and only if you choose to send it.

Windows clients cannot be recorded yet. A converter from Windows heap tracing (ETW) is the planned next piece. Linux recordings can replay on Windows meanwhile. The kit and protocol live in bench/alloc-eval in the rusty_alloc repository.

glibc, rusty_alloc, mimalloc, and jemalloc

jemalloc had the cheapest start-up (48 percent less allocator time) and the lowest battle memory (160 MB under glibc), but its combined per-frame cost in the battle was 69 percent worse at the 99th percentile. mimalloc faults new pages in during the allocation call, so its battle allocator time read twice glibc's while its combined total was level. It also had six to eleven times glibc's count of calls over 100 microseconds on both recordings. If a studio already ships mimalloc or jemalloc, that build is the baseline the next run should use. The kit takes another allocator in a few lines.

The replay refuses a verdict it cannot defend. Every allocator is checked for alignment, overlapping blocks, contents preserved across frees and resizes, and zeroed memory before anything is timed. glibc runs twice under two labels, and when those two disagree the row gets no verdict. Replays follow the recording's own timeline so threads overlap as they did in the game, and a run more than 10 percent behind pace is dropped. rusty_alloc and mimalloc default to opposite memory policies, so each is also run with the other's. Allocator time alone misranks allocators that fault pages at different moments, so the total including first writes stays a headline row.

What game developers should decide before they ship

Dedicated server cost and the hitch a player feels

The game memory allocator decision is three questions. Allocator CPU per server tick, under your entity counts: a lower number with a steadier spread is headroom per instance. How many frames lose more than a millisecond to allocation, and how bad the worst 1 percent are. Peak and retained memory for each allocator and each policy, so a CPU gain is priced before anyone commits. If those numbers favor a swap, the integration test is a one-line global allocator in a build you control.

The machine behind these figures was one 24-thread laptop running Linux under WSL2, on a host shared with other work. Treat the direction as evidence. Remeasure the magnitude on the hardware that will ship.

Where the allocator sits beside the rest of the house

Shipping this game memory allocator is a Memory-band change. rusty_alloc is the crate this measurement is about, next to SpaceDB when a session's durable copy should stay on hardware you hold. The map for that band is Vision: Memory. Studios that keep design docs and patch notes in the browser, with nothing uploaded, already have that habit at ragconverter.com (the field note). The mesh a client can join when the venue is not a single landlord is mata.network (in the wild). When the same local-first rule is a sensor in a house rather than a game process, the surface is dldeploy.com (sensors that stay home). Review the lockfile before the swap ships with Deputy.

FAQ

Quick answers for builders evaluating this technology.

Which allocators were in the game memory allocator comparison?

glibc was the baseline. The others were rusty_alloc 2.2, mimalloc 2.3, and jemalloc 5, five repetitions each, in rotated order, with a null arm and 95 percent bootstrap intervals.

Did the battle get faster overall?

Not in a way the noise floor would sign. Allocator time plus the page faults of first writing new memory landed at 0.90x and 1.07x versus glibc, and that interval includes 1.0.

How can game developers run this without changing the game?

Record the unmodified binary with LD_PRELOAD, then replay that stream through each allocator. Traces store block ids, sizes, threads, and timestamps. They do not store addresses or memory contents.

Why did peak memory go up?

rusty_alloc keeps freed blocks for reuse by default. On the recorded battle that was about 403 MB more than glibc. Setting RUSTY_ALLOC_PURGE_DELAY=10 trades that memory back and raises the count of slow calls.

Does this prove a frame-rate win on every game?

No. The replay measures the allocator's share on Endless Sky, on one laptop under WSL2. Whole frame time, and any larger dedicated server, still has to be measured on that build.