<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alex Nazimov</title>
    <description>The latest articles on DEV Community by Alex Nazimov (@__f57a448).</description>
    <link>https://dev.to/__f57a448</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4091151%2F6956cc6d-0f40-4718-b301-0d195933a11d.jpeg</url>
      <title>DEV Community: Alex Nazimov</title>
      <link>https://dev.to/__f57a448</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/__f57a448"/>
    <language>en</language>
    <item>
      <title>Optimizing Disk I/O in NumPy: Implementing a Fast LZ4 Compression Algorithm via C-Extensions</title>
      <dc:creator>Alex Nazimov</dc:creator>
      <pubDate>Sun, 23 Aug 2026 18:56:41 +0000</pubDate>
      <link>https://dev.to/__f57a448/optimizing-disk-io-in-numpy-implementing-a-fast-lz4-compression-algorithm-via-c-extensions-ilm</link>
      <guid>https://dev.to/__f57a448/optimizing-disk-io-in-numpy-implementing-a-fast-lz4-compression-algorithm-via-c-extensions-ilm</guid>
      <description>&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
        &lt;div class="c-embed__cover"&gt;
          &lt;a href="https://pypi.org/project/numpy-cache/" class="c-link align-middle" rel="noopener noreferrer"&gt;
            &lt;img alt="" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpypi-camo.freetls.fastly.net%2F67af7117035e2345bacb5a82e9aa8b5b3e70701d%2F68747470733a2f2f73746f726167652e676f6f676c65617069732e636f6d2f707970692d6173736574732f73706f6e736f726c6f676f732f73656e7472792d77686974652d6c6f676f2d4a2d6b64742d706e2e706e67" height="55" class="m-0" width="248"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="c-embed__body"&gt;
        &lt;h2 class="fs-xl lh-tight"&gt;
          &lt;a href="https://pypi.org/project/numpy-cache/" rel="noopener noreferrer" class="c-link"&gt;
            numpy-cache · PyPI
          &lt;/a&gt;
        &lt;/h2&gt;
          &lt;p class="truncate-at-3"&gt;
            High-performance LZ4 cache for NumPy arrays
          &lt;/p&gt;
        &lt;div class="color-secondary fs-s flex items-center"&gt;
            &lt;img alt="favicon" class="c-embed__favicon m-0 mr-2 radius-0" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fpypi.org%2Fstatic%2Fimages%2Ffavicon.35549fe8.ico" width="32" height="30"&gt;
          pypi.org
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;
&lt;br&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/macht1212" rel="noopener noreferrer"&gt;
        macht1212
      &lt;/a&gt; / &lt;a href="https://github.com/macht1212/numpy-cache" rel="noopener noreferrer"&gt;
        numpy-cache
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      High-performance LZ4 cache for NumPy arrays
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;NumPy Cache – Fast LZ4-Based Caching for NumPy Arrays&lt;/h1&gt;
&lt;/div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h3 class="heading-element"&gt;High‑performance, lightweight disk cache for NumPy arrays with LZ4 compression – now with configurable compression speed.&lt;/h3&gt;
&lt;/div&gt;
&lt;p&gt;&lt;a href="https://github.com/macht1212/numpy-cache/actions/workflows/ci.yml" rel="noopener noreferrer"&gt;&lt;img src="https://github.com/macht1212/numpy-cache/actions/workflows/ci.yml/badge.svg" alt="CI"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://www.python.org/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/aeb851aab490066e9d7aa5b6af558384aaea59ea8da4a34bf827d961ad57d060/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f507974686f6e2d332e3132253230253743253230332e3133253230253743253230332e31342d626c7565" alt="Python"&gt;&lt;/a&gt; &lt;a rel="noopener noreferrer" href="https://github.com/icons/NumPy-2.5.2-blue.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2Ficons%2FNumPy-2.5.2-blue.svg" alt=""&gt;&lt;/a&gt; &lt;a rel="noopener noreferrer" href="https://github.com/icons/License-Apache%202.0-blue.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2Ficons%2FLicense-Apache%25202.0-blue.svg" alt=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;📌 The Problem&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;When dealing with large NumPy arrays, developers face a classic trade‑off:&lt;/p&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Speed (100 MB)&lt;/th&gt;
&lt;th&gt;Issue&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;np.save()&lt;/code&gt; / &lt;code&gt;np.savez()&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;~45 ms&lt;/td&gt;
&lt;td&gt;Huge storage, slow network transfer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;np.savez_compressed()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;~3.6 s&lt;/td&gt;
&lt;td&gt;Single‑threaded DEFLATE (zlib) is too slow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;numpy_cache&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~100 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Best of both worlds&lt;/strong&gt; ✅&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;There is a clear gap: no lightweight, specialised solution combines &lt;code&gt;np.save()&lt;/code&gt; speed with good compression – until now.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;🚀 Features&lt;/h2&gt;

&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;✅ Blazing fast – 20 × faster than np.savez_compressed()&lt;/li&gt;
&lt;li&gt;✅ Good compression – 2 × smaller than np.save()&lt;/li&gt;
&lt;li&gt;✅ Pure C extension – minimal overhead, maximum performance&lt;/li&gt;
&lt;li&gt;✅ Configurable speed – acceleration parameter (1–16) lets you trade compression ratio for speed&lt;/li&gt;
&lt;li&gt;✅ NumPy integration – works with all numeric dtypes (int, uint, float, bool)&lt;/li&gt;
&lt;li&gt;✅ Multi‑dimensional – supports up…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/macht1212/numpy-cache" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;When processing large volumes of numerical data in Python, developers regularly face a trade-off between disk I/O speed and storage space. Standard tools in the NumPy ecosystem cover the extremes of this spectrum but leave a gap for scenarios requiring simultaneous efficiency in both parameters.&lt;/p&gt;

&lt;p&gt;In this article, we will analyze the limitations of standard approaches and explore the architecture of a lightweight solution based on C-extensions and the LZ4 algorithm. This solution allows caching arrays with latency comparable to an uncompressed memory dump, while achieving a compression ratio that matches or exceeds standard zlib.&lt;/p&gt;

&lt;h3&gt;
  
  
  Analysis of Standard Approaches and Their Limitations
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;np.save()&lt;/code&gt;: Performs a direct binary data dump. It has minimal latency (less than 1 ms per megabyte) but performs no compression. When working with large datasets, this leads to inefficient use of disk space and network bandwidth.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;np.savez_compressed()&lt;/code&gt;: Uses the deflate (zlib) algorithm. It provides good compression, but the compression process is a CPU-bound operation. On a 100 MB array, the write operation can take over 3 seconds, which is unacceptable for model training loops or real-time systems.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;HDF5 (via h5py) or Zarr&lt;/code&gt;: Powerful tools for storing multidimensional arrays. However, they require pulling in heavy dependencies, learning a specific API, and configuring chunking, which is excessive for the simple task of fast caching of intermediate ndarray states.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Solution Architecture
&lt;/h3&gt;

&lt;p&gt;To eliminate this bottleneck, there is a solution that combines the speed of direct memory access with the efficiency of LZ4. Key architectural decisions include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;C-extension (Python C-API)&lt;/strong&gt;: The critical execution path is offloaded to C. This avoids Python interpreter overhead and the Global Interpreter Lock (GIL) during serialization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Direct Memory Access (Buffer Protocol)&lt;/strong&gt;: Using PyArray_DATA allows obtaining a direct pointer to the contiguous memory block of the array, minimizing copy operations (a zero-copy approach where the data structure permits it).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LZ4 Algorithm&lt;/strong&gt;: The &lt;code&gt;LZ4_compress_fast&lt;/code&gt; function is selected, allowing flexible control over the balance between speed and compression ratio via an acceleration parameter.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  API Usage Example
&lt;/h3&gt;

&lt;p&gt;The library interface is intentionally kept minimal to integrate into existing code with minimal changes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;numpy_cache&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;save&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;load&lt;/span&gt;

&lt;span class="c1"&gt;# Initialize an array of ~100 MB
&lt;/span&gt;&lt;span class="n"&gt;arr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;randn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Save with default parameters (acceleration=4)
# The library automatically determines the data type and shape
&lt;/span&gt;&lt;span class="nf"&gt;save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;arr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;cache_data.npc&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Tuning the balance: increasing acceleration to 16 
# reduces the compression ratio but maximizes write speed
&lt;/span&gt;&lt;span class="nf"&gt;save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;arr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;cache_data_fast.npc&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;acceleration&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Load with full metadata restoration (dtype, shape)
&lt;/span&gt;&lt;span class="n"&gt;loaded_arr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;cache_data.npc&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Explanation&lt;/strong&gt;: The &lt;code&gt;save&lt;/code&gt; function accepts an &lt;code&gt;ndarray&lt;/code&gt; object and passes the retrieved data to the C-module, where compression and header writing occur. The file extension can be anything (&lt;code&gt;.bin&lt;/code&gt;, &lt;code&gt;.cache&lt;/code&gt;, &lt;code&gt;.npy&lt;/code&gt;), as parsing relies exclusively on the byte structure of the header.&lt;/p&gt;

&lt;h3&gt;
  
  
  Binary Serialization Structure
&lt;/h3&gt;

&lt;p&gt;To ensure data portability between architectures (e.g., x86_64 and ARM64), a simple and predictable structure is used: a fixed header followed by a compressed data block.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="c1"&gt;// The pragma pack directive ensures no field alignment padding,&lt;/span&gt;
&lt;span class="c1"&gt;// guaranteeing a consistent header size (96 bytes) across all platforms.&lt;/span&gt;
&lt;span class="cp"&gt;#pragma pack(push, 1)
&lt;/span&gt;&lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="n"&gt;CacheHeader&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kt"&gt;uint64_t&lt;/span&gt; &lt;span class="n"&gt;uncompressed_size&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// Original data size in bytes&lt;/span&gt;
    &lt;span class="kt"&gt;uint64_t&lt;/span&gt; &lt;span class="n"&gt;compressed_size&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;// Size of the compressed block&lt;/span&gt;
    &lt;span class="kt"&gt;uint64_t&lt;/span&gt; &lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;          &lt;span class="c1"&gt;// Array dimensions (supports up to 8 dimensions)&lt;/span&gt;
    &lt;span class="kt"&gt;uint32_t&lt;/span&gt; &lt;span class="n"&gt;magic&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;             &lt;span class="c1"&gt;// Magic number 0x4C5A4E43 ("LZNC") for validation&lt;/span&gt;
    &lt;span class="kt"&gt;uint32_t&lt;/span&gt; &lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;           &lt;span class="c1"&gt;// Structure version (current: 1)&lt;/span&gt;
    &lt;span class="kt"&gt;uint32_t&lt;/span&gt; &lt;span class="n"&gt;ndim&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;              &lt;span class="c1"&gt;// Actual number of dimensions&lt;/span&gt;
    &lt;span class="kt"&gt;uint32_t&lt;/span&gt; &lt;span class="n"&gt;dtype&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;             &lt;span class="c1"&gt;// Internal NumPy data type identifier&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="cp"&gt;#pragma pack(pop)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using a magic number allows quickly rejecting files with corrupted structures or incorrect formats during loading, without attempting to decompress invalid data. Support for non-contiguous slices (strided arrays) is implemented by pre-creating a contiguous copy of the data in memory before compression, if the source array is not C-contiguous.&lt;/p&gt;

&lt;h3&gt;
  
  
  Benchmarks and Performance Analysis
&lt;/h3&gt;

&lt;p&gt;Testing was conducted on float32 arrays of ~1 MB. The goal of the tests was to record the overhead of serialization and deserialization compared to standard methods.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write Latency&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Intel i5-1235U (Ubuntu)&lt;/th&gt;
&lt;th&gt;Apple M1 (macOS)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;np.save&lt;/td&gt;
&lt;td&gt;880 μs&lt;/td&gt;
&lt;td&gt;548 μs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;np.savez&lt;/td&gt;
&lt;td&gt;1,083 μs&lt;/td&gt;
&lt;td&gt;554 μs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;np.savez_compressed (zlib)&lt;/td&gt;
&lt;td&gt;28,922 μs&lt;/td&gt;
&lt;td&gt;32,211 μs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NumPy Cache (accel=4)&lt;/td&gt;
&lt;td&gt;756 μs&lt;/td&gt;
&lt;td&gt;635 μs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Observation&lt;/strong&gt;: The write speed of the LZ4-based solution is comparable to uncompressed &lt;code&gt;np.save&lt;/code&gt; and demonstrates an order-of-magnitude reduction in execution time compared to zlib.&lt;/p&gt;

&lt;h3&gt;
  
  
  Read Latency
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Intel i5-1235U (Ubuntu)&lt;/th&gt;
&lt;th&gt;Apple M1 (macOS)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;np.save&lt;/td&gt;
&lt;td&gt;92 μs&lt;/td&gt;
&lt;td&gt;77 μs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;np.savez&lt;/td&gt;
&lt;td&gt;401 μs&lt;/td&gt;
&lt;td&gt;225 μs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;np.savez_compressed (zlib)&lt;/td&gt;
&lt;td&gt;4,859 μs&lt;/td&gt;
&lt;td&gt;2,877 μs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NumPy Cache (accel=4)&lt;/td&gt;
&lt;td&gt;421 μs&lt;/td&gt;
&lt;td&gt;670 μs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Observation&lt;/strong&gt;: During reading, there are expected overheads for LZ4 decompression (~300–500 μs) compared to direct &lt;code&gt;np.save&lt;/code&gt; mapping. However, this time remains an order of magnitude lower than zlib, making this compromise well-justified for most caching tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Compression Ratio
&lt;/h3&gt;

&lt;p&gt;For the test dataset:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;np.save&lt;/code&gt; / &lt;code&gt;np.savez&lt;/code&gt;: ~1.0 MB&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;np.savez_compressed&lt;/code&gt;: ~0.5 MB&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NumPy Cache&lt;/strong&gt;: ~0.4 MB (depending on the &lt;code&gt;acceleration&lt;/code&gt; parameter)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In certain scenarios with numerical data, the LZ4 algorithm shows better compression than zlib due to its dictionary handling specifics and the absence of redundant checks inherent to deflate. However, the challenge of caching purely random data remains.&lt;/p&gt;

&lt;h3&gt;
  
  
  Managing the &lt;code&gt;acceleration&lt;/code&gt; Parameter
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;acceleration&lt;/code&gt; parameter (range 1–16) is passed directly to &lt;code&gt;LZ4_compress_fast&lt;/code&gt;. It determines how aggressively the algorithm searches for matches in the dictionary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1–4&lt;/strong&gt;: Maximum match exploration. Recommended for archival storage where write time is not critical.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;5–10&lt;/strong&gt;: Balance. A value of 4–8 is optimal for most intermediate result caching tasks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;11–16&lt;/strong&gt;: Minimal match search. Used in systems with strict write latency requirements (real-time), where a slight increase in file size is acceptable.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;The presented approach demonstrates that caching NumPy arrays does not always require resorting to heavy formats like HDF5. Using C-extensions combined with LZ4 allows achieving latencies close to uncompressed dumps while significantly saving disk space.&lt;/p&gt;

&lt;p&gt;The project's source code is available under the Apache 2.0 license. Implementation details of the C-module and scripts for reproducing the benchmarks can be found on &lt;a href="https://pypi.org/project/numpy-cache/" rel="noopener noreferrer"&gt;PyPI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>compression</category>
      <category>numpy</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
