Introduction to CRAM

A new compression model, called CRAM, offers a different path to compression that avoids swap entirely by keeping the compressed data in memory, and it offers up to 452x the performance of ZRAM. Memory compression has been a feature of operating systems for a long time. In Linux, the most popular options are zswap and ZRAM, but both of these options are fundamentally swap-layer features. CRAM is a new take that aims to boost performance tremendously.

CRAM was conceptualized by Gregory Price and his team at Meta. The central realization that spurred the development of CRAM is that the largest portion of the performance hit from compressed memory isn't the compression itself, but rather the fault and swap behavior. The core idea behind CRAM is to perform memory compression completely in memory instead of utilizing swap space.

Architecture and Implementation

CRAM makes use of existing Linux mechanisms to enable radically higher-performance compressed memory, particularly on reads. It uses a private NUMA node instead of pretending to be a block device, which lets Linux continue using all of its memory semantics, including migration and ballooning, to manage CRAM.

A critical part of the architecture is the "Chicken Bit," which tells Linux to stop trying to use CRAM while it is busy managing allocations. Compressibility of data varies tremendously depending on the workload, ranging from easily compressible repeating data to already-compressed non-compressible files. This creates challenges in determining the available logical RAM.

While logical allocation tracking remains an area of ongoing research, the Chicken Bit helps prevent cascading failures, referred to as a "poison storm," when write operations outpace CRAM's ability to allocate resources.

Performance and Speed Comparisons

Because CRAM is stored in RAM and treated as RAM with full cacheline and byte access, read-only data can be accessed with minimal delay beyond hardware-offloaded compression costs. CRAM runs at near DRAM speed. In worst-case read scenarios, CRAM achieves 489 million operations per second compared to ZRAM's 1.1 million.

When write operations are enabled, CRAM remains significantly faster than ZRAM, delivering a 5.4x speedup in the worst tested case of 20% writes. The performance drop during writes is due to the necessity of page faulting and migrating folios back to the original NUMA domain, as direct writes to compressed data would cause corruption.

My explanation of CRAM might not be completely correct; I wasn't at the Linux Plumbers' Conference in Prague to hear the presentation directly, so I'm working off of the slides available from the session information site (h/t to Phoronix for the spot) I think I've managed to grasp the gist of what Price is getting at, though.

Future Integration

While originating at Meta Platforms with large Linux servers as the primary target, ZRAM and Zswap are deployed across the broader Linux ecosystem on devices such as the Steam Deck. CRAM could offer substantial performance improvements for resource-constrained machines once remaining implementation questions are resolved and it is integrated into the mainline kernel.