Look at almost any production cluster and you will find the same shape. A thin slice of recent data gets nearly all reads and writes. Everything older is rarely touched but still sits on premium storage, inflates indexes, and eventually forces a bigger tier. You pay hot prices to store silence.
I have lived that bill. Compliance forbade TTL deletion. A handmade export pipeline meant owning a second query path forever. Atlas Online Archive became the managed middle ground: move eligible documents to MongoDB-managed object storage and keep them queryable through federation. The personal stake is not "cool feature." It is whether you can stop buying CPU and disks for data nobody opens.
How it works in practice
You define an archive rule, usually a date age, optionally a custom filter. Atlas runs background jobs that copy matching documents to object storage using the partition fields you chose. After a confirmed copy, documents leave the cluster. Federated connection strings query archive-only or cluster-plus-archive under the same namespace.
A few properties matter more than the marketing blurbs. Dedicated clusters (M10+) only. Archived data is read-oriented; do not treat it as your OLTP write plane. Jobs are asynchronous and consume cluster resources while they run. Your normal URI sees only the live cluster; historical access needs the federated endpoint. That last sentence is where teams get surprised.
Good candidates versus poor ones
| Good candidates | Poor candidates |
|---|---|
| Event logs and audit trails | User profiles that keep changing |
| Orders older than a stable hot window | Tiny reference collections |
| IoT readings and metrics | Data that must stay low-latency forever |
| Closed tickets past the edit window | Documents refunded or edited indefinitely |
Ask: once a document is old, is it effectively immutable? If not, push the archive age past the last realistic edit, or filter on a terminal status. Also ask whether anyone needs the data at all. If the answer is no, a TTL index is simpler than archiving.
Designing the rule without painting yourself into a corner
Date-based rules are the common case:
Archive documents where createdAt is older than 90 days
Index the predicate fields so the job does not scan forever:
db.orders.createIndex({ createdAt: 1 })
Custom filters help when coldness is not pure age (status in terminal states and closedAt old enough). Prefer moving windows over hard-coded calendar dates that rot. Start conservative (365 days) and ratchet hotter as confidence grows. Prefer an immutable date field. Archiving on updatedAt surprises you when batch jobs touch old rows and suddenly nothing looks cold enough.
Partition fields decide archive query cost
Partitions let federation skip irrelevant object-storage slices. The date field is typically one partition. Add a small number of extra fields ordered by how often historical queries filter on them (customerId, region). Moderate cardinality helps. Unique-per-document fields help little. Partitions are hard to change later. List the three real archive queries your team runs before you click save. That conversation has prevented more regret than any checklist I own.
Application patterns that keep history honest
Keep two clients when history matters: a live URI for writes and hot reads, and a federated URI for legitimate historical access. Time-route in the API: recent windows hit the cluster; older windows accept slower federation. For huge pulls, prefer async export over interactive scans. Set UI expectations honestly ("older than 90 days may load slower"). Users forgive slower history. They do not forgive silent empty results because the app still uses the live URI.
What Online Archive is not
It is not a backup or point-in-time restore substitute. It is not a free analytics lakehouse for arbitrary heavy joins. It is not writable storage for casual corrections (plan rehydrate flows). It is not worth the complexity on tiny collections. If historical access depends on stages unsupported against archived data, redesign that path before enabling the feature.
Cost sanity check
Savings come from smaller disks, smaller backups, and sometimes a lower cluster tier. You pay archive storage and federation processing instead. The win is largest when most bytes are cold and rarely queried.
const cutoff = new Date(Date.now() - 180 * 24 * 60 * 60 * 1000)
const total = await db.collection('orders').estimatedDocumentCount()
const eligible = await db.collection('orders').countDocuments({ createdAt: { $lt: cutoff } })
console.log(`${((eligible / total) * 100).toFixed(1)}% eligible`)
If most of the collection is eligible and those documents are rarely read, archiving is likely worth it. If eligibility is tiny, you are adding operational surface for theater.
Rollout checklist and common pitfalls
Confirm tier support and permissions. Measure first-month archive volume. Pilot on a low-risk collection during a quiet window. Prove a known archived id is readable through federation. Monitor job success, lag, and cluster CPU during backlog drain. Document rehydrate steps for the rare "make this hot again" case. Align retention and residency with compliance.
Common pitfalls stay boring and expensive: archiving data that still changes, missing indexes on rule fields, poor partition choices you cannot cheaply revise, forgetting that the app URI never sees archive rows, using archive where TTL deletion would do, and putting latency-sensitive features on federated reads.
Closing
Online Archive shines when collections are large because of age, not because of an active working set. Define crisp hot windows, archive with indexed predicates, query cold data with slower-SLO honesty, and keep writes on the live cluster. That is how you stop paying premium prices to store silence.
Top comments (0)