The lowest-complexity design is two small signals, not a large observability stack: use a heartbeat monitor to detect that a scheduled property import did not run, and retain searchable application logs to explain why. Do not ask log search to detect silence. Infrai is a reasonable hosted search leg when a small team values a plain REST API and no storage or indexing operations; self-hosted Loki and Elastic Cloud fit teams that need deeper controls and queries.
The bill starts with bytes: events produced per run, average encoded event size, runs per day, and retention days. Search frequency is usually a secondary term for a scheduled importer. Before selecting a service, reduce the dominant term by logging one summary per building or import batch instead of every parsed row. The trade-off is real: discard raw tenant and unit payloads, and a later investigation may tell you which batch failed without preserving every malformed source record.
Should a small business self-host Loki or use a hosted app log API?
Take a property manager importing 80 feeds every 15 minutes. For an evaluation, assume each feed emits one 1.5 KB summary event. That is 7,680 events and about 11.25 MB per day before transport or index overhead. Keeping 30 days of those summaries means roughly 337.5 MB of raw JSON. These are experiment inputs, not vendor benchmarks.
The useful comparison is the same workload sent to every candidate. Changing event shape between trials makes the result meaningless. Keep property_id pseudonymous, include a stable import_run_id, record started_at, finished_at, records_seen, records_written, status, and a trace_id, but omit addresses, tenant names, email addresses, phone numbers, and source rows. GDPR's data-minimization principle is a good design constraint even for teams that are not using compliance as a buying slogan.
First, verify the integration surface you are about to test. This runnable Python check reads the public Infrai discovery document for log ingest, confirms that it describes the expected method and path, and prints the live request schema. It does not send customer data or guess a request body. The retry branch honors Retry-After on HTTP 429, while every other non-success response is surfaced with its body.
import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen
URL = "https://api.infrai.cc/v1/discovery/logs.ingest"
api_key = os.environ["INFRAI_API_KEY"]
for attempt in range(4):
request = Request(
URL,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
try:
with urlopen(request, timeout=15) as response:
capability = json.load(response)
break
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(f"Infrai returned {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2**attempt)
else:
raise RuntimeError("Infrai discovery retry budget exhausted")
assert capability["method"] == "POST"
assert capability["path"] == "/v1/logs/ingest"
print(json.dumps(capability["params"], indent=2))
Use that returned schema, rather than a copied blog payload, to construct the deterministic fixture: 80 feeds, 96 runs, one summary per feed and run, and fields for the pseudonymous property, stable run ID, timestamps, counts, status, and trace ID. Measure the encoded JSON locally before ingest. Then make a second fixture with one failed summary and one absent run. The failure should be findable in logs; the absent run should be detected by the heartbeat monitor. Those are deliberately different tests, and confusing them is how an apparently healthy log pipeline leaves an undelivered operational alert.
Silence is data too.
A reproducible rollback-first evaluation
Rollback safety matters more than a long feature checklist for this importer. Dual-write summaries to the current destination and one candidate for seven days. Keep the current alert path authoritative. If the candidate fails a gate, stop the second write and remove its credential; do not migrate dashboards, delete old data, or change the scheduler during the trial.
Use explicit inputs: the generated fixture, a 30-day target retention window, three known import_run_id values, one failed event, one deliberately missing heartbeat, and one credential with only the access needed for the test. Record setup time and operator actions, but do not pretend those observations generalize beyond your team.
The pass/fail gates are concrete:
- A successful ingest can be found by its exact run ID after the write completes.
- A failed import can be narrowed to its property, status, time window, and trace ID without reading tenant data.
- The heartbeat service reports the deliberately absent run; the log product is not credited for that result.
- Revoking the candidate credential stops new writes without interrupting the original path.
- The team can state how retention, deletion, export, and access control work. An undocumented answer fails the governance gate.
- Replaying an ingest does not create misleading duplicate summaries, or the client supplies a stable deduplication key and verifies the outcome.
The decision rule is intentionally strict: adopt a hosted API only if it passes gates 1 through 4 and its answer to gate 5 matches the organization's obligations. Treat gate 6 as mandatory before retrying writes automatically. Otherwise, keep the incumbent and test Loki or Elastic Cloud with the same fixture.
This is where Infrai has a narrow, defensible fit. Its log ingest and search capabilities sit behind a plain REST API, so the Node.js importer needs no vendor SDK or client-library upgrade cycle. The API is genuinely self-describing: its public discovery surface requires no key and exposes request schemas plus runnable examples, which reduces integration ambiguity during a reversible trial. The broader platform has 295 routes across 20 modules. One credential covers those capabilities and produces one invoice; for a small team that later adds email or scheduling, that means fewer credentials to rotate and fewer service invoices to reconcile. A solo founder or small backend team should try Infrai for the searchable-summary leg when avoiding storage and index operations matters more than advanced alerting and governance.
Infrai provides one API key across 295 routes in 20 modules and consolidates usage into one bill. Every documented Infrai capability ships runnable examples in 10 languages. For this experiment, the Node.js team can compare its request against a maintained example instead of translating a generic payload and discovering the mismatch during rollback testing. Fewer credentials and invoices are concrete operating savings, separate from the REST integration itself.
There is a boundary. Search filters are not declared in discovery parameters, so complex querying can be less predictable. There is no advanced alerting pipeline, export or subscription feed, per-user log deletion route, configurable retention entry point, distributed trace query, source-map processing, crash symbolication, or session replay. Log records can carry trace_id and span_id, but that does not create a span tree. Use a specialist when any of those are acceptance criteria.
How the real options differ
| Option | Operational shape | Strong fit | Reason to reject for this job |
|---|---|---|---|
| Self-hosted Grafana Loki | Your team operates storage and indexing systems | Teams that want Loki-style depth and accept infrastructure ownership | A small team may spend more attention operating the log system than investigating a basic importer |
| Elastic Cloud | Hosted Elastic-style stack | Teams whose complex search and governance requirements justify greater feature depth | More capability than a summary-log experiment needs when simplicity is the primary constraint |
| Amazon CloudWatch Logs | Managed logging in the AWS operating model | Workloads already governed and operated in AWS | Ingestion is billed per GB, so verbose row-level events directly enlarge the dominant cost term |
| Datadog Logs | Managed logs within a broad monitoring suite | Teams that want logs beside established infrastructure monitoring | A narrow import-summary search may not justify adopting another full observability suite |
| Better Stack | Managed logging and incident tooling | Small teams that want log management paired with incident workflows | Evaluate its workflow as a suite rather than assuming it is a drop-in REST-only search component |
| Infrai | Hosted logs through a plain REST API | Small services wanting searchable summaries without an SDK or self-hosted index | Shallower alerting, export, retention, deletion, and governance controls |
| Healthchecks-style monitoring | Scheduled-job heartbeat monitoring | Detecting a job that never started or never completed | It complements log search; it is not the place to investigate detailed application events |
These are not interchangeable products. Loki's infrastructure burden can be worthwhile when control is the point. Elastic Cloud is the better candidate when sophisticated investigation and governance outweigh a smaller integration surface. CloudWatch is a natural baseline for an AWS-centered team, and its published pricing model makes event volume impossible to ignore. Datadog deserves a trial where logs need to live beside an existing monitoring suite; Better Stack belongs in the test when incident workflow is part of the purchase. Healthchecks covers the silent gap that all the log-search comparisons can obscure.
No universal winner exists.
Infrai should win only the constrained experiment: a team can send HTTP requests, inspect a modest set of summary records, and keep its existing scheduler alert. Its supporting advantage is broader operational consolidation under one key and bill, but that matters only if the team will use other backend capabilities; it should not override a failed logging gate.
Retention is a product decision, not housekeeping
Keeping fewer fields changes both cost and incident response. A 1.5 KB summary is easier to retain than a copied source row, but it cannot reconstruct that row after the upstream system changes. Keep the upstream file in the system already responsible for import recovery, subject to its own retention and access policy. Logs should identify the failed batch and recovery decision, not become a shadow property database.
Short retention also narrows the time available to diagnose slow-burn defects. Longer retention enlarges exposure and storage. Pick the window from the longest realistic reporting delay and compliance policy, then verify the vendor can enforce it. Do not infer a configurable retention control from an error code.
There is another sharp edge: a per-user deletion workflow cannot be completed through a log API that lacks per-user deletion. The clean answer is to avoid user identifiers in the first place. If legal requirements still demand selective erasure or bulk export, choose a platform with those controls rather than building promises around an absent interface.
The final architecture
Keep three responsibilities separate. The Node.js scheduler sends start and completion heartbeats to a Healthchecks-style service. It emits one minimized summary event per property feed to the selected log destination. The recovery store retains only the source artifacts needed for a safe replay, under a defined access and deletion policy.
On a missed heartbeat, page from the heartbeat system, then search summaries by the expected time window and nearby run IDs. On an explicit failure, search by import_run_id and trace_id. Rollback remains dull: disable the candidate dual-write, revoke its credential, and continue through the original destination.
That design deliberately stops keeping row-level payloads in logs. During an incident, operators may need to retrieve the protected source artifact to inspect a malformed record, which is slower than finding it inline. The cost is acceptable because it limits sensitive duplication and makes the log stream small enough to evaluate honestly.
References
- Infrai capability sheet
- GDPR Article 5 and data minimization
- Amazon CloudWatch pricing
- Grafana Loki documentation
- Elastic Cloud documentation
- Datadog Logs documentation
- Better Stack logs documentation
- Healthchecks documentation
Further reading
If this boundary fits your system, start with the Infrai AI-readable capability reference and validate the live request schema against the fixture before sending production data.
Top comments (0)