DEV Community

Cover image for Streaming Large Files in Node.js: Readable, Writable, and Transform Streams
DEVANSHU PATIL
DEVANSHU PATIL

Posted on AI-assisted

Streaming Large Files in Node.js: Readable, Writable, and Transform Streams

Streaming Large Files in Node.js: Readable, Writable, and Transform Streams

title: "Streaming Large Files in Node.js: Readable, Writable, and Transform Streams"
published: true
published_at: "2026-11-09T09:00:00+05:30"
description: "Master memory-efficient file processing in Node.js using Readable, Writable, and Transform streams. Learn how to handle multi-gigabyte payloads, manage backpressure, and leverage the pipeline API."
tags: [nodejs, backend, javascript, performance]
ai_disclosure_level: some_ai

Handling multi-gigabyte files in a server-side runtime requires careful resource management. A common anti-pattern in Node.js is loading an entire file into memory using fs.readFile() or parsing a massive multipart upload via buffer concatenation. When multiple concurrent requests hit this code path, V8's garbage collector is overwhelmed, and the application quickly runs out of heap memory, resulting in an ERR_STRING_TOO_LONG or JavaScript heap out of memory crash. Node.js Streams solve this problem by processing data piece-by-piece, or chunk-by-chunk, keeping memory consumption constant regardless of the file size. ## Understanding the Core Stream Types Node.js provides four fundamental stream types: 1. Readable: Abstraction for a source from which data can be consumed (e.g., fs.createReadStream(), incoming HTTP requests). 2. Writable: Abstraction for a destination to which data can be written (e.g., fs.createWriteStream(), HTTP responses). 3. Transform: A Duplex stream where the output is computed from the input (e.g., zlib.createGzip(), crypto hashing). 4. Duplex: Streams that are both Readable and Writable (e.g., TCP sockets). ### The Mechanics of Chunks and Buffers Instead of loading a 5GB CSV file into RAM, a Readable stream reads a small segment of the file (typically 64KB by default) into a buffer, emits a data event, and waits for the downstream consumer to process it. ## The Problem with Manual Event Listeners Historically, developers wired streams together using manual event listeners (data, end, error, drain). However, this approach is prone to memory leaks and fails to handle backpressure correctly.

const http = require('http');
const fs = require('fs');

// ANTI-PATTERN: Manual piping without error handling or backpressure management
const server = http.createServer((req, res) => {
  const readable = fs.createReadStream('./massive-video.mp4');

  readable.on('data', (chunk) => {
    const canWrite = res.write(chunk);
    if (!canWrite) {
      // If the destination buffer is full, we must pause the source
      readable.pause();
    }
  });

  res.on('drain', () => {
    // Resume reading when the destination buffer clears
    readable.resume();
  });

  readable.on('end', () => {
    res.end();
  });

  readable.on('error', (err) => {
    console.error(err);
    res.statusCode = 500;
    res.end('Internal Server Error');
  });
});
Enter fullscreen mode Exit fullscreen mode


Manual handling of pause(), resume(), and drain is tedious and error-prone. If an error occurs on the readable stream, the writable stream is left open, causing resource leaks. ## Modern Stream Composition with pipeline The modern, production-grade approach is to use stream/promises pipeline. The pipeline utility module safely pipes between streams, forwards errors cleanly, cleans up all streams when an error occurs or the pipeline completes, and natively manages backpressure.

const http = require('http');
const fs = require('fs');
const path = require('path');
const { pipeline } = require('stream/promises');

const server = http.createServer(async (req, res) => {
  const filePath = path.join(__dirname, 'massive-video.mp4');

  try {
    // pipeline handles backpressure and error propagation automatically
    await pipeline(
      fs.createReadStream(filePath),
      res
    );
  } catch (err) {
    // If headers haven't been sent, we can send a 500 error
    if (!res.headersSent) {
      res.statusCode = 500;
      res.end('Internal Server Error');
    } else {
      // Otherwise, abort the response
      res.destroy();
    }
    console.error('Pipeline failed:', err);
  });
});

server.listen(3000);
Enter fullscreen mode Exit fullscreen mode


## Implementing Custom Transform Streams When building data processing pipelines—such as transforming a massive CSV to JSON, masking sensitive user data, or compressing payloads—Custom Transform streams are necessary. The following example demonstrates a custom transform stream that reads a stream of newline-delimited log strings, parses them, filters out non-error logs, and outputs formatted JSON chunks.

const { Transform } = require('stream');

class LogParserTransform extends Transform {
  constructor(options = {}) {
    super({ ...options, objectMode: true });
    this._buffer = '';
  }

  _transform(chunk, encoding, callback) {
    // Append incoming buffer chunk to internal string buffer
    this._buffer += chunk.toString();
    const lines = this._buffer.split('
');

    // Keep the last partial line in the buffer
    this._buffer = lines.pop();

    for (const line of lines) {
      if (!line.trim()) continue;
      try {
        const logEntry = JSON.parse(line);
        // Filter out non-error levels
        if (logEntry.level === 'ERROR') {
          this.push(JSON.stringify(logEntry) + '
');
        }
      } catch (err) {
        // Handle malformed JSON lines safely
        this.emit('warning', new Error(`Invalid JSON line: ${line}`));
      }
    }

    callback();
  }

  _flush(callback) {
    // Process any remaining data in the buffer
    if (this._buffer.trim()) {
      try {
        const logEntry = JSON.parse(this._buffer);
        if (logEntry.level === 'ERROR') {
          this.push(JSON.stringify(logEntry) + '
');
        }
      } catch (err) {
        // Ignore final malformed chunk or handle accordingly
      }
    }
    callback();
  }
}

module.exports = LogParserTransform;
Enter fullscreen mode Exit fullscreen mode


### Consuming the Custom Transform with pipeline We can now combine our custom transform stream with file system streams safely using stream/promises:

const fs = require('fs');
const { pipeline } = require('stream/promises');
const LogParserTransform = require('./LogParserTransform');

async function processLogs() {
  try {
    await pipeline(
      fs.createReadStream('./application.log', { encoding: 'utf8' }),
      new LogParserTransform(),
      fs.createWriteStream('./errors-only.log')
    );
    console.log('Log processing completed successfully.');
  } catch (error) {
    console.error('Log processing failed:', error);
    process.exit(1);
  }
}

processLogs();
Enter fullscreen mode Exit fullscreen mode


## Backpressure Management Explained Backpressure is the mechanism Node.js uses to prevent a fast producer from overwhelming a slow consumer. When data is pushed into a Writable stream faster than it can be drained (e.g., writing to a slow disk or a congested network socket), the internal buffer of the Writable stream exceeds its highWaterMark. When this threshold is crossed, .write() returns false. The pipeline helper automatically listens to this signal, pauses the upstream Readable stream until the Writable stream emits the drain event, and then resumes data flow. This guarantees that your application's RAM footprint remains stable regardless of file size or I/O bottlenecks. ## Conclusion Node.js Streams are a cornerstone of performant backend engineering. By avoiding buffer accumulation, utilizing stream/promises pipeline for resource management and error handling, and implementing custom Transform classes, you can build systems capable of processing arbitrarily large datasets with minimal resource consumption.

Top comments (0)