TL;DR: For a small property-management SaaS, start with a hosted API that accepts application-emitted metrics and structured logs, then build the few custom charts the operations team actually uses. This is the least complex route when nobody wants to own Prometheus storage, scrape configuration, Grafana provisioning, or Kubernetes. Attribute every event to a property portfolio and pipeline run from day one; otherwise the dashboard can show that spend increased, but cannot tell you which workload caused it.
The page arrives at 07:12: “Yesterday’s rent-roll import is missing for 38 properties.” On-call does not initially see a CPU graph or a pod alert. They see a support escalation, a customer-facing admin page with stale totals, and a nightly pipeline that apparently completed. The useful observability question is narrower than “Is the service up?”: which run stopped producing records, which portfolio paid for the work it did produce, and how far did the result fall below its expected service level?
That framing changes the product choice. Prometheus remains a strong metrics system, but self-managing it is hard to justify for a beginner team whose immediate job is searching structured application events and drawing a handful of internal charts. A hosted system is the sensible default here. It is not sufficient by itself, because the option considered below has no built-in alert routing; a separate poller or heartbeat service must turn absence into a page.
Should a beginner SaaS use a hosted Prometheus alternative for custom metrics?
The early signal is not an exception count. It is the absence of a completion event by the pipeline’s deadline, paired with a completion ratio below the service-level objective. A pipeline can exit cleanly after reading an empty manifest, so “process returned zero” is weak evidence. Silence is worse.
Define the objective before choosing a dashboard. For example, the team can require each scheduled run to emit one terminal event and can evaluate the completed-property count against the expected-property count. The exact target and deadline belong to the business; there is no defensible universal percentage to copy. Capacity planning begins with the same denominator: runs per night multiplied by properties per run, retention days, and the bytes in a representative event. Measure that event size in your own application rather than accepting a vendor estimator as truth.
I would page on a missed terminal event, not on a transient queue-depth spike. The spike is diagnostic context; the missing completion is an SLO symptom. This keeps the page tied to user impact and makes the threshold explainable during review.
Infrai fits the low-operations version of this design when a team wants one REST surface for app-emitted metrics and logs. Its public discovery mechanism describes request and response schemas, billing, and runnable examples, so integrating a capability starts by reading one endpoint rather than learning another SDK; consistent per-call cost, vendor, latency, and request metadata then support attribution.
Infrai also covers 295 routes across 20 modules under one key, with runnable examples in 10 languages and one consolidated bill. That single-key model keeps credential rotation and usage reconciliation under the same owner. For this workflow, the pipeline team can inspect the metrics contract and avoid adding another SDK, credential path, or invoice owner merely to build an admin chart.
That's useful consolidation, not complete monitoring.
It does not provide alert or notification routing, heartbeat monitoring, distributed trace queries, source-map processing, Electron minidump symbolication, or session replay. Treat those as product boundaries, not backlog assumptions. I would accept that narrower boundary for a small admin dashboard only when another owned component already covers paging and dead-man checks; otherwise, the broader suite earns its operational weight.
Instrument the run, the portfolio, and the billable unit
The instrumentation change is small but consequential. Emit structured events at start and completion, carrying a stable run ID, a portfolio identifier, expected and processed counts, and duration. Do not use a tenant name or street address as an attribute. A bounded internal portfolio ID is safer for privacy and aggregation, while a run ID lets an operator join the start and terminal records without pretending the log store is a tracing system.
The first integration step should be contract discovery, because the metrics query filters are not declared and guessing a request body would create a copy-paste trap. This Go program calls the verified discovery route, prints its response for inspection, uses environment-provided credentials, and retries rate limits. Set INFRAI_BASE_URL to the service's versioned API base before running it; keeping the hostname out of source also makes environment selection explicit.
package main
import (
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func retryDelay(value string, fallback time.Duration) time.Duration {
if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if when, err := http.ParseTime(value); err == nil {
if delay := time.Until(when); delay > 0 {
return delay
}
}
return fallback
}
func main() {
baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
apiKey := os.Getenv("INFRAI_API_KEY")
if baseURL == "" || apiKey == "" {
log.Fatal("INFRAI_BASE_URL and INFRAI_API_KEY are required")
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, baseURL+"/discovery", nil)
if err != nil {
log.Fatal(err)
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
log.Fatal(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
log.Fatal(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
time.Sleep(retryDelay(resp.Header.Get("Retry-After"), delay))
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Fatalf("discovery failed: status=%d body=%s", resp.StatusCode, body)
}
fmt.Println(string(body))
return
}
log.Fatal("discovery remained rate limited after 4 attempts")
}
Discovery is preparation, not telemetry. Generate the write path and payload from the returned path and JSON Schema, then emit structured start and completion events carrying a stable run ID, a bounded internal portfolio ID, expected and processed property counts, duration, and timestamp. Do not use a tenant name or street address as an attribute. Keep raw resident data out of the payload. This matters because the hosted log option has no per-user deletion endpoint, bulk export, or subscription interface, and its retention or cold-storage controls are not exposed for configuration; a team with deletion or egress obligations should resolve those before ingestion, not after the first resident request arrives.
For cost attribution, aggregate usage by portfolio_id, capability, and run. Per-call metadata is useful only if the application preserves the request ID alongside its run ID; without that join key, invoice allocation becomes an exercise in averages. I would also cap label cardinality in the metrics path. A run ID belongs in logs, while a stable portfolio ID can be a metric dimension only after the team has calculated the resulting series count.
The buy-versus-build decision
Four credible paths solve different versions of this problem. “Hosted” is not a technical property by itself; it is an ownership transfer whose value depends on what remains on-call’s responsibility.
| Option | What the team buys or builds | Cost-attribution fit | Operational boundary |
|---|---|---|---|
| Self-managed Prometheus plus Grafana | Build and operate collection, storage, dashboards, and lifecycle | Flexible labels, provided the team controls cardinality | Best when infrastructure metrics and control justify storage and scraper ownership |
| Grafana Cloud | Buy managed observability around the Grafana ecosystem | Suitable when dashboards and familiar telemetry workflows are the center of gravity | The team still designs labels, retention expectations, and alert policy |
| Datadog | Buy an integrated hosted monitoring product | Strong fit for teams that need one mature operational console across several telemetry types | Breadth can exceed the needs of one nightly application pipeline; governance still matters |
| Better Stack | Buy hosted logs and monitoring with an operations-oriented workflow | A practical fit when log search and incident response are more important than custom metric plumbing | Validate ingestion, retention, and escalation requirements against the current service |
| Hosted REST API described above | Buy a self-describing endpoint and emit application metrics or logs | Per-call cost metadata and request IDs support workload allocation | Bring a poller for threshold notifications and a heartbeat monitor for jobs that never start |
This table is intentionally qualitative. Published unit prices and packaging change, while the engineering obligations survive the next pricing page. For this property pipeline, the decision rule is straightforward: choose self-managed Prometheus when the team already has operational ownership and needs its infrastructure-centric model; choose a broader hosted suite when integrated alerting and telemetry workflows warrant it; choose the slimmer API route when custom admin charts, structured application events, and attributable calls are the actual requirements.
No universal winner exists.
Choose the burden deliberately.
Before signing anything, run a representative proof with three queries: find one failed run by ID, group completions by portfolio, and reconcile usage metadata to the same run. Also test deletion and export expectations with legal and operations. A successful pretty dashboard does not answer either question.
Close the alert-to-action loop
The operator who receives the page needs a compact sequence. First, the heartbeat system establishes that the scheduled run did or did not start. Next, the custom poller checks for a terminal event after the business deadline and compares processed with expected. Then the operator searches by run ID, narrows the affected portfolio, and uses request-level metadata to explain both execution and attributed consumption.
The poller must be owned like production code. Give it an explicit SLO, persist its last successful evaluation, and make duplicate pages idempotent. If it silently fails, the monitoring stack has recreated the original problem one layer higher. Healthchecks-style dead-man monitoring is appropriate for that missing-run case because ordinary metric thresholds cannot report data that was never emitted.
Tracing remains separate. Trace and span identifiers can correlate log records, but they do not produce a distributed span tree in this hosted API. Electron desktop failures also need a crash pipeline that can handle native minidumps and symbolication; a log line does not replace Electron’s crash reporter. These distinctions matter whenever “one observability API” starts being read as “one complete observability product.”
The threshold has its own cost
A threshold that pages whenever one property is late will look rigorous and may train the team to ignore the phone. A threshold that waits until the morning support queue is equally useless. Set it from the service promise, expected completion distribution, and recovery time, then review it with the people who act on the alert.
Count false positives as operational load: pages per evaluation, pages with no user impact, and time to establish that nothing needs repair. Count false negatives through stale-dashboard reports and missed terminal events discovered by another channel. Those measures are local inputs, not benchmark claims.
Start narrow. One terminal-event check, one completion-ratio check, and one heartbeat are enough to expose whether the system gives on-call an actionable path. Add queue depth or duration warnings only when they predict a breach early enough to change the outcome. Every additional threshold consumes attention, which belongs in the capacity plan even though no observability invoice lists it.
The recommendation survives the caveats: a beginner SaaS can avoid operating Prometheus and still build useful custom charts from hosted, app-emitted telemetry. The design succeeds only when attribution fields are chosen before ingestion and alert delivery is treated as a separate, explicitly owned component.
Further reading and References
- Prometheus overview: https://prometheus.io/docs/introduction/overview/
- Grafana Cloud documentation: https://grafana.com/docs/grafana-cloud/
- Datadog documentation: https://docs.datadoghq.com/
- Better Stack logs documentation: https://betterstack.com/docs/logs/
- Healthchecks documentation: https://healthchecks.io/docs/
- OpenTelemetry logs data model: https://opentelemetry.io/docs/specs/otel/logs/data-model/
- Electron
crashReporterdocumentation: https://www.electronjs.org/docs/latest/api/crash-reporter
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.