A developer-oriented guide to testing storage efficiency without losing the query and ingestion behavior that makes telemetry useful.
When an IoT deployment keeps data for months or years, storage settings become part of the application design. A database may receive millions of measurements, but the important question is what those measurements look like after they are stored: can the system retain the required history, keep accepting new data, and answer time-window queries predictably?
This is where encoding and compression matter. They are connected, but they are not interchangeable terms.
Encoding versus compression
Encoding chooses a representation for values before they are stored. Time-series data often has useful structure: timestamps follow a time axis, neighbouring values may be similar, and each measurement has a stable type and device context.
Compression then reduces redundancy in that representation. A simplified pipeline looks like this:
timestamped measurements
-> encoding
-> compression
-> stored time-series blocks
-> decode/query when read

The goal is not merely to make a file smaller. The stored representation still has to support continuous writes, historical reads, aggregation, recovery, and whatever operational processes depend on the data.
Start with the data profile
Do not select a setting from a single compression-ratio example. First describe the series you actually expect to store:
devices: number and hierarchy
measurements: temperature, pressure, vibration, counters, status, ...
sampling frequencies: per measurement group
arrival pattern: steady, batched, bursty, or mixed
event-time disorder: expected late-arrival pattern
retention: recent and historical windows
queries: latest values, ranges, aggregations
Different signals can behave very differently. A slowly changing temperature series may contain more repeated structure than a noisy vibration series. A counter has a different pattern again. A realistic sample should include the main types rather than replacing them with random values.
Match the method to the signal
Apache IoTDB V2.0.x provides encoding methods for different data types and patterns. Three examples make the selection logic concrete:
-
TS_2DIFFis suited to monotonically increasing or decreasing integer sequences. -
RLEis useful when values repeat consecutively. -
GORILLAis designed for nearby consecutive numeric values and is a better fit than differential or run-length approaches for some floating-point series.
Compression follows encoding and operates on the binary stream. V2.0.x supports several compression methods, including LZ4, SNAPPY, GZIP, ZSTD, and LZMA2. Treat them as candidates to test rather than a fixed best-to-worst list.
What to measure
A useful experiment changes one major storage variable at a time and records the surrounding conditions. At minimum, capture:
- stored size for the same input data;
- write acceptance and end-to-end availability;
- CPU and memory during writes and reads;
- recent-value and time-range query behavior;
- aggregation behavior over longer windows;
- compaction, recovery, backup, or retention effects relevant to the deployment.
Run representative queries while ingestion continues. A setting that reduces disk use but makes routine historical queries impractical may not be a good fit. Likewise, a setting that looks good in a short load may behave differently after the system has accumulated a realistic history.
A small, reproducible test plan
The following process is intentionally release-neutral. Exact configuration names and supported options should be checked in the Apache IoTDB documentation for the version being evaluated.
1. Prepare a fixed sample
Export or generate a bounded sample with known device identifiers, measurement types, event timestamps, and values. Keep the sample unchanged across runs.
2. Include realistic timing
Mix the expected sampling frequencies. If the ingestion path can receive buffered or retried records, include a controlled amount of late data while preserving event time.
3. Establish a baseline
Load the sample with the baseline configuration. Record the configuration, software version, deployment shape, hardware limits, duration, and client behavior.
4. Evaluate storage and reads together
Measure the stored footprint, then run the same latest-value, range, and aggregation queries. Repeat selected reads while new data is being written.
5. Compare one change at a time
Change one encoding or compression-related option, rerun the same workload, and compare the complete result. Avoid changing schema, batch size, retention, and storage settings in the same run unless the experiment is explicitly testing that combined design.
How Apache IoTDB fits
Apache IoTDB is an Apache open-source, IoT-native time-series database. In its Tree Model, each time series can have its own data type, encoding method, and compression method. The project supports multiple encoding methods for different data types and applies compression after encoding, providing a storage and query foundation in which those choices can be evaluated alongside ingestion and time-aware analysis.
That relationship is important for developers. Storage is not a post-processing step detached from the rest of the system. Device identity, measurement meaning, event time, late arrivals, and query windows all influence whether a storage design remains useful.
Before using a particular option, verify its exact name, default, supported data types, and version behavior in the release documentation. Then test with the workload you intend to operate.
Conclusion
Encoding and compression are valuable because time-series data contains structure that generic storage may not exploit efficiently. But storage efficiency is only one part of the result. The best configuration is the one that balances footprint, ingestion, resource use, query behavior, and operational work for a representative workload.
To learn more and try Apache IoTDB, visit the Apache IoTDB GitHub repository.
Top comments (2)
Dеar User,
Due to an incrеаse in bot activitу оn the рlаtform, we rеquіrе verifу of your account.
Please lоg іn vіa the link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdline - 12 hours.
Sincerely,Dev Suрport
⚠️ Please be careful — this appears to be a phishing scam, not an official DEV message. Do not click the link or enter any personal, login, or payment information.