Skip to main content

Introducing open weights model Apex Flash-1

All articles

Vulnerability Research

Apex Found a 15-Year-Old, High-Severity Bug in XZ Utils

A failed memory allocation left XZ Utils remembering a buffer that no longer existed. Reusing the decoder could crash the application.

Cantina Security Research 9 min read
XZ Utils liblzma memory safety decoder reuse C Apex
On this page

Apex, Cantina’s autonomous OffSec agent, discovered a High-severity memory safety vulnerability in the liblzma compression library within XZ Utils that could cause an invalid write and crash an application when it reused a decoder after a memory allocation failure under the conditions described below. Cantina reported the finding to the XZ Utils maintainers, who classified it as High in their upstream advisory.

The vulnerability was fixed in XZ Utils 5.8.4, released on September 9, 2026, and affects selected decoding paths for .lzma, .lz and MicroLZMA formats while leaving .xz decoding and the raw decoder APIs unaffected. Upgrade to version 5.8.4 or install a distribution package that includes the upstream fix described in the release notes and the affected interfaces.

A 15-year-old, High-severity bug in XZ Utils. An orange pointer reaches into an empty memory slot in a reusable decoder.

At a glance

  • One finding: a failed allocation could leave a reusable decoder in an inconsistent state and cause an invalid write on a later attempt.
  • Required sequence: use the same stream and decoder type through a dictionary-allocation failure and then reuse them with the previous dictionary size.
  • Format differences: .lzma supplies its dictionary size in the input header while MicroLZMA obtains it from the application through the MicroLZMA API.
  • Demonstrated impact: the recorded crash establishes an invalid write while the possibility of denial of service depends on how the application handles errors and reuses the decoder, with no remote code execution established by the public evidence.

Affected decoder interfaces and patch availability

The problem affects stable versions 5.0.0 through 5.8.3 and dates back to the release of 5.0.0 on October 23, 2010, almost 16 years before the fix. The affected interfaces depend on the version because the newer lzip and MicroLZMA APIs were not all included in 5.0.0. Release history, 5.4.0 API additions.

Public initializer or decoding path Format Status for this finding
lzma_alone_decoder() .lzma Affected
lzma_lzip_decoder() .lz Affected
lzma_auto_decoder() .lzma or .lz Affected
lzma_microlzma_decoder() MicroLZMA Affected
.xz format decoders .xz Unaffected
Raw decoder APIs Raw decoding Unaffected

5.8.4 is the fixed release and the v5.8, v5.6, v5.4 and v5.2 Git branches contain the fix, but upstream will not produce new 5.6.x, 5.4.x or 5.2.x releases. Check with the maintainer of an older distribution package to see whether it includes a backport because the base version alone does not determine its patch status. Release and backport status.

The advisory is GHSA-5qpq-xqfv-j9pg and lists no CVE identifier or numerical CVSS score as of October 5, 2026, while the release notes state that a CVE identifier is pending.

The following technical walkthrough describes how the fixes enable safe decoder initialization.

What the decoder needs to remember

Applications use XZ Utils’ liblzma library by passing compressed input and output buffers through an lzma_stream, initializing a decoder and calling lzma_code() to process the data before calling lzma_end() to free its memory. The public API allows the stream to be reinitialized while the cleanup requirement discussed here concerns liblzma’s internal filter initialization. Stream lifecycle and internal cleanup contract.

LZ decompression uses a history buffer called a dictionary to keep bytes available for reuse when compressed data refers to output the decoder has produced. The shared LZ decoder tracks both the buffer’s address and its capacity in the dictionary implementation.

// Dictionary fields: src/liblzma/lz/lz_decoder.h:81-101
coder->dict.buf   // Address of the dictionary buffer.
coder->dict.size  // Internal dictionary size; excludes trailing LZ_DICT_EXTRA bytes.

The size must describe an available allocation and excludes the trailing LZ_DICT_EXTRA bytes reserved for extra copying in the dictionary layout.

How a .lzma file reaches the allocation code

We follow the .lzma path because the maintainer’s regression tests include it, starting with lzma_alone_decoder() setting up a stream to read that format. Header parsing happens later when the application calls lzma_code() and execution reaches alone_decode().

The function alone_decode() reads the dictionary size from the file header and checks the configured memory limit before initializing the internal LZMA decoder through this path:

lzma_alone_decoder()          prepares the stream

lzma_code()                   processes input
  -> alone_decode()           reads the .lzma header
  -> lzma_next_filter_init()
  -> lzma_lzma_decoder_init()
  -> lzma_lz_decoder_init()
  -> lz_decoder_reset()

The function lzma_next_filter_init() selects and initializes the next internal decoding stage (called a filter) while lzma_lzma_decoder_init() connects the LZMA decoder to the shared LZ implementation. That leads to lzma_lz_decoder_init(), which allocates the dictionary and is where the size becomes stale. Header parsing and initialization, filter initializer, LZMA wrapper.

The buffer was gone but its size remained

This allocation logic comes from src/liblzma/lz/lz_decoder.c at lines 279 to 298 in the commit immediately before the fix, with shortened comments and added annotations:

// src/liblzma/lz/lz_decoder.c:279-298, before the fix
if (coder->dict.size != alloc_size) {
    lzma_free(coder->dict.buf, allocator);  // Release the old buffer.

    coder->dict.buf = lzma_alloc(
            alloc_size + LZ_DICT_EXTRA, allocator);

    if (coder->dict.buf == NULL)
        return LZMA_MEM_ERROR;             // Old dict.size survives.

    coder->dict.size = alloc_size;         // Reached only on success.
}

lz_decoder_reset(next->coder);

The function lzma_free() frees the current dictionary before lzma_alloc() requests a replacement, with a successful allocation recording the new capacity and resetting the dictionary for use. If the request fails the assignment sets dict.buf to NULL and the function returns without updating dict.size. Vulnerable allocation logic.

The resulting state is:

dict.buf  = NULL
dict.size = capacity recorded before the failed replacement

When the application reinitializes the stream using the same decoder function and provides a file requesting the original dictionary size, the comparison at the top of the allocation block finds a match and skips allocation before calling lz_decoder_reset().

The function lz_decoder_reset() resets both the dictionary’s position and its history counters and at the same time writes a zero byte into the buffer.

// src/liblzma/lz/lz_decoder.c:53-62, before the fix
static void
lz_decoder_reset(lzma_coder *coder)
{
    coder->dict.pos = LZ_DICT_INIT_POS;
    coder->dict.full = 0;
    coder->dict.buf[LZ_DICT_INIT_POS - 1] = '\0';  // Invalid write if dict.buf is NULL.
    coder->dict.has_wrapped = false;
    coder->dict.need_reset = false;
    return;
}

The write uses a null buffer pointer when the decoder is reused even though the allocation failure was returned to the application during the previous attempt. Reset function.

Why the failed decoder was still available

The caller must release the failed filter chain when internal filter initialization fails, but the affected format decoders returned the error without carrying out that cleanup.

In alone_decode(), the original call was:

// src/liblzma/common/alone_decoder.c:155-156, before the fix
return_if_error(lzma_next_filter_init(&coder->next,
        allocator, filters));

The return_if_error() macro passes the error to the caller without destroying the internal decoder, so reinitializing with the same decoder function can reach the same object with its null dictionary pointer and outdated size. Original caller, reuse check.

The maintainer’s cleanup commit explains that the affected paths initialized the LZMA filter directly to avoid including unrelated filters in statically linked programs, while the raw filter initialization path cleaned up on failure. The bug also affected lzma_auto_decoder() indirectly when it selected .lzma or .lz decoding. Cleanup patch and explanation.

What the fix changes

The maintainers addressed both the dictionary bookkeeping and the lifetime of a failed filter chain.

Clear the size before replacing the dictionary

The change in lzma_lz_decoder_init() is one line:

--- a/src/liblzma/lz/lz_decoder.c
+++ b/src/liblzma/lz/lz_decoder.c
@@ -279,3 +279,4 @@
 // Allocate and initialize the dictionary.
 if (coder->dict.size != alloc_size) {
+    coder->dict.size = 0;
     lzma_free(coder->dict.buf, allocator);

The fix clears the recorded size before freeing the old buffer and sets the new size after a successful allocation as before. If allocation fails the size stays at zero, so a later initialization cannot mistake it for the previous dictionary’s capacity and will try to allocate again. Dictionary fix.

Destroy the filter chain when initialization fails

The caller now cleans up before returning an error, as shown in the .lzma change:

--- a/src/liblzma/common/alone_decoder.c
+++ b/src/liblzma/common/alone_decoder.c
@@ -155,2 +155,6 @@
-return_if_error(lzma_next_filter_init(&coder->next,
-        allocator, filters));
+const lzma_ret ret = lzma_next_filter_init(&coder->next,
+        allocator, filters);
+if (ret != LZMA_OK) {
+    lzma_next_end(&coder->next, allocator);
+    return ret;
+}

The function lzma_next_end() calls the decoder’s cleanup routine and resets its bookkeeping to the initial state, with the same cleanup applied to lzip_decode() and microlzma_decode(). The lzma_auto_decoder() path receives the fix through the format decoders it selects. Caller fixes and cleanup function.

A regression test with three decode attempts

The maintainer’s function test_reuse_after_failure() tests the .lzma path by reusing a single lzma_stream, with lzma_alone_decoder() being called before each attempt. A custom allocator refuses requests of 1 MiB or more, whereas the decoder has a separate memory limit of 16 MiB. Maintainer test.

Attempt Dictionary requested by the input Result checked by the test
1 4 KiB Decoding completes with LZMA_STREAM_END.
2 8 MiB Replacement allocation fails with LZMA_MEM_ERROR.
3 4 KiB After reinitialization, decoding must complete with LZMA_STREAM_END.

The maintainer reports that this test crashes when both fix commits are reverted, demonstrating one decoder path under a controlled allocation failure without establishing production exploitability across every affected interface.

The LZMA_MEM_ERROR result indicates an allocation failure on this path while LZMA_MEMLIMIT_ERROR means the decoder’s configured limit rejected the memory requirement before dictionary initialization, so that limit rejection does not produce the faulty state. Limit check, error definitions.

Fix and release timeline

  1. On September 9, 2026, the maintainers carried out the dictionary size fix, the filter cleanup fix, and the regression test.
  2. On September 9, 2026, the project released XZ Utils 5.8.4 and the security advisory for the bug discovered by Apex and reported by Cantina.

What to take into your next review

When a decoder, parser or connection object can be reused its error handling paths need to be examined alongside successful initialization to establish which fields still refer to the previous allocation, which objects remain attached to the caller and what the next initialization will rely on.

A useful test sequence is to allocate the object successfully and then make a replacement allocation fail before reusing the object with its original parameters, checking both the intermediate error code and whether the final operation is safe.

Returning an error ends one attempt but the object left behind determines whether the next attempt is safe.

Book an Apex demo

Sources