DEV Community

Tamiz Uddin
Tamiz Uddin

Posted on Originally published at tamiz.pro

Beyond the PID: Why 'Artifact Age' Beats 'Liveness' in Modern Agent Observability

Originally published on tamiz.pro.

The Illusion of Liveness

In the era of autonomous AI agents, the traditional monitoring paradigm is breaking. For decades, the gold standard of infrastructure health has been the Liveness Probe (or PID check): Is the process running? Is the port open? Is the heartbeat being sent? If yes, the system is green.

This model works perfectly for stateless web servers and simple background daemons. However, it fails catastrophically for autonomous agents (LLM-powered workflows, RAG pipelines, or multi-step tool-using bots). An agent can be "alive"—consuming memory, holding open connections, and running its event loop—yet be completely stuck in an infinite reasoning loop, waiting on a dead external API, or silently producing garbage outputs due to a model hallucination.

To engineers building agent observability stacks, this creates a critical blind spot. A "Liveness" signal tells you the container didn't crash. It does not tell you the agent is actually working.

This article explores a shift in monitoring philosophy: moving from Process Liveness to Artifact Age. We argue that the most reliable indicator of agent health is not "Is it running?" but "How old is its last useful output?"

Understanding the Problem Domain

The Nature of Agent Stalls

Autonomous agents operate on probabilistic logic and external dependencies. Unlike a deterministic state machine, an agent's internal state can degrade in subtle ways:

  1. Logical Stalls: The LLM enters a reasoning loop, re-reading its own context without making progress. The process is CPU-active but non-productive.
  2. Dependency Hangs: The agent calls a third-party API (e.g., a search engine or a database) that times out silently or returns malformed data, causing the agent to wait or retry indefinitely.
  3. Context Overflow: As the conversation history grows, the token window fills up. The model begins to lose coherent thread, leading to repetitive or nonsensical tool calls that keep the process "alive" but move the task forward zero percent.

In all three cases, the PID exists. The heartbeat is sent. But the task is dead.

Why Liveness Probes Fail Here

A standard Kubernetes livenessProbe or a custom health check endpoint (/health) typically verifies:

  • The gRPC/HTTP server is accepting connections.
  • The database connection pool is not saturated.
  • The main event loop is not blocked.

None of these checks verify semantic progress. If your agent is stuck calling search_tool 500 times with the same query because it failed to update its internal plan, the /health endpoint will still return 200 OK. The infrastructure is healthy; the agent is not.

Introducing "Artifact Age"

Artifact Age is a metric defined as the time elapsed since the agent last produced a verifiable, forward-moving output.

An "Artifact" is not just a log line. It is a concrete result of work. In the context of agent observability, artifacts include:

  • Data Outputs: A successfully parsed JSON response from an API.
  • State Transitions: A change in the agent's planning state (e.g., from "Gathering Data" to "Synthesizing").
  • Tool Successes: A tool execution that returned status success.
  • Final Answers: The completion of a user-facing task.

Artifact Age is calculated as:

$$ Age = T_{current} - T_{last_valid_artifact} $$

If Age exceeds a threshold $\tau$ (derived from the expected SLA of the task), the agent is considered stalled, regardless of its process liveness.

Why This Metric is Superior

  1. Task-Centric: It measures progress toward the goal, not just existence.
  2. Model-Agnostic: It works whether the agent is running on a local LLM, a cloud API, or a hybrid setup. It doesn't care how the computation happens, only what it produced.
  3. Actionable: If Artifact Age spikes, you know immediately that the workflow is stuck. You can trigger a SIGKILL, reset the context, or route the task to a different model, rather than just restarting a healthy-but-idle container.

Architecting the Artifact Monitoring System

Implementing this requires a shift in how we instrument agent code. We need to move from passive logging to active "Heartbeats of Work."

The Instrumentation Layer

We recommend wrapping the agent's core execution loop with an Artifact Emitter. This component is responsible for tagging events that constitute valid progress.

import time
from datetime import datetime, timezone
from typing import Any, Dict, Optional

class ArtifactEmitter:
    def __init__(self, agent_id: str):
        self.agent_id = agent_id
        self.last_artifact_time = datetime.now(timezone.utc)
        self.artifact_counter = 0
        self.status = "alive" # alive, stalled, healthy

    def emit_artifact(self, artifact_type: str, payload: Optional[Dict[str, Any]] = None):
        """
        Call this when the agent produces a meaningful unit of work.
        """
        self.last_artifact_time = datetime.now(timezone.utc)
        self.artifact_counter += 1
        self.status = "healthy"

        # In a real system, publish this to a metrics store (Prometheus/Datadog)
        self._publish_metric("agent_artifact_emitted", {
            "agent_id": self.agent_id,
            "type": artifact_type,
            "timestamp": self.last_artifact_time.isoformat()
        })

    def get_age_seconds(self) -> float:
        """
        Calculate the Artifact Age.
        """
        now = datetime.now(timezone.utc)
        delta = now - self.last_artifact_time
        return delta.total_seconds()

    def check_stall_status(self, threshold_seconds: float) -> bool:
        """
        Determine if the agent is stalled based on Artifact Age.
        """
        age = self.get_age_seconds()
        if age > threshold_seconds:
            self.status = "stalled"
            return True
        return False
Enter fullscreen mode Exit fullscreen mode

Defining "Valid" Artifacts

Not all logs are artifacts. To avoid false positives (thinking an agent is stuck when it's just thinking), you must define what counts.

For a data-extraction agent:

  • Valid Artifacts: Successful API call, successful regex match, updated internal state.
  • Invalid Artifacts (Ignored): LLM token generation (unless it results in a tool call), retry attempts that failed, internal debug logs.

If an agent is waiting for a human approval, the "Artifact" is the request for approval. Once the human acts, that is a new artifact. The age clock resets.

Implementing Health Checks Based on Artifact Age

In a production environment, you rarely want to kill an agent immediately when it pauses. Instead, you use a tiered response system based on Artifact Age.

The Monitoring Loop

You can implement a sidecar monitor or a central orchestrator that polls the agent's ArtifactEmitter.

class AgentHealthMonitor:
    def __init__(self, emitter: ArtifactEmitter, thresholds: dict):
        self.emitter = emitter
        self.thresholds = thresholds
        # thresholds = {
        #     'warning': 30,   # seconds
        #     'critical': 120, # seconds
        #     'kill': 300      # seconds
        # }

    def evaluate(self):
        age = self.emitter.get_age_seconds()

        if age > self.thresholds['kill']:
            self.action("kill", reason="Artifact age exceeded critical threshold")
        elif age > self.thresholds['critical']:
            self.action("alert", reason="Agent likely stalled")
            self.action("inject_context", reason="Attempting to nudge agent")
        elif age > self.thresholds['warning']:
            self.action("log_warning", reason="Agent slower than expected")
        else:
            pass # Healthy

    def action(self, action_type: str, reason: str):
        # Implement side effects: log, send signal, update dashboard
        print(f"[MONITOR] {action_type}: {reason} (Age: {self.emitter.get_age_seconds():.2f}s)")
Enter fullscreen mode Exit fullscreen mode

Advanced Technique: Context Injection for Stalls

One of the unique advantages of monitoring Artifact Age is the ability to intervene intelligently.

If an agent is stalled (Artifact Age > 120s), a traditional restart loses all context. With Artifact Age monitoring, you can attempt a "nudge":

  1. Detect that the last artifact was a tool_call_failed or no new artifacts for 120s.
  2. Inject a system prompt message: "You have been stuck for 2 minutes. Please check your progress and try a different approach or summarize where you are blocked."
  3. Reset the Artifact Age clock.

If the agent produces a new artifact (even just a status update), the clock resets and it may recover. If it fails again, you escalate to a hard kill. This preserves work-in-progress and reduces cost by avoiding full re-initialization of the agent context.

Comparison: Liveness vs. Artifact Age

Feature Liveness (PID/Port) Artifact Age Use Case
Measures Process existence Task progress Infrastructure vs. Logic
False Positives High (stuck loops look alive) Low (stuck loops look old) Accuracy of "Working"
Implementation Standard (/health) Custom Instrumentation Effort to implement
Actionability Restart Container Nudge / Reset Context / Kill Granularity of control
Cost Awareness Ignores Token Burn Can correlate with token spend Efficiency monitoring
Best For Stateless Services Stateful, Probabilistic Agents

Practical Scenarios and Edge Cases

Scenario 1: The "Thinking" Agent

An LLM agent is allowed to "think

to perform multi-step reasoning without immediate output. In a traditional liveness model, this appears as a dead or stalled node, triggering false-positive restarts. With Artifact Age, we monitor the generation of intermediate "thought" artifacts (chain-of-thought traces, tool call logs, or draft responses). If the age of the latest reasoning_step artifact exceeds the threshold for its specific phase (e.g., 5 seconds for logic, 10 seconds for tool waiting), we flag it as STUCK, distinguishing it from ACTIVE_WAITING. This precision prevents the wasteful killing of agents that are simply performing complex cognitive tasks.

Scenario 2: The Zombie Worker

Consider an agent that successfully completes its main task (generating a report) but fails to clean up its temporary files or release its memory pool. Its HTTP server might still respond to health checks, maintaining a "Live" status. However, its last_artifact is a report generated 10 minutes ago, while the current time is 10 minutes 30 seconds later. The system detects that no new deliverables are being produced. By correlating this with rising memory usage, the orchestrator can mark the agent as ZOMBIE and gracefully terminate it, reclaiming resources that a liveness probe would never identify.

Scenario 3: The Flaky Dependency

An agent relies on an external API that is rate-limiting. The agent enters a retry loop with exponential backoff. A liveness check might see the process running and assume health. Artifact Age, however, reveals that the last successful api_response artifact is 2 minutes old, while the expected refresh rate is every 15 seconds. The observability stack tags this as DEGRADED_DEPENDENCY. The system can now proactively adjust the agent's priority or scale it down to prevent resource exhaustion from rapid retries, rather than waiting for a total failure.

Implementing Artifact Age Monitoring: A Code Deep Dive

To move from theory to practice, we need a lightweight instrumentation layer that does not impose significant overhead on the agent’s runtime. The key is to decouple the agent's business logic from its observability heartbeat. Instead of generic heartbeats, we use semantic tagging.

Below is a Python implementation using a standard library approach that can be adapted to any language. It demonstrates how to wrap agent tasks to automatically track artifact age and publish this data to a monitoring bus.

import time
import threading
from dataclasses import dataclass
from typing import Dict, Any, Callable
from enum import Enum

class ArtifactType(Enum):
    INPUT_RECEIVED = "input_received"
    THOUGHT_GENERATED = "thought_generated"
    TOOL_INVOCATION = "tool_invocation"
    FINAL_OUTPUT = "final_output"
    HEARTBEAT = "heartbeat"

@dataclass
class ArtifactEvent:
    agent_id: str
    artifact_type: ArtifactType
    timestamp: float
    payload: Dict[str, Any] = None
    # The 'age' is calculated on the consumer side, but we store the absolute time.

class AgentArtifactMonitor:
    def __init__(self, agent_id: str, publish_fn: Callable[[ArtifactEvent], None]):
        self.agent_id = agent_id
        self.publish = publish_fn
        self.last_artifacts: Dict[ArtifactType, float] = {}
        # Define expected freshness thresholds per artifact type
        self.thresholds = {
            ArtifactType.INPUT_RECEIVED: 60.0,  # Should see new input within 60s
            ArtifactType.FINAL_OUTPUT: 30.0,    # Should produce output within 30s of task
            ArtifactType.HEARTBEAT: 5.0         # Internal tick every 5s
        }

    def record(self, artifact_type: ArtifactType, payload: Dict[str, Any] = None):
        """
        Record an artifact. This is the primary hook for the agent's code.
        """
        now = time.time()
        self.last_artifacts[artifact_type] = now
        event = ArtifactEvent(
            agent_id=self.agent_id,
            artifact_type=artifact_type,
            timestamp=now,
            payload=payload
        )
        self.publish(event)

    def check_health(self) -> Dict[str, Any]:
        """
        Returns a health status based on the age of critical artifacts.
        """
        now = time.time()
        status = {
            "agent_id": self.agent_id,
            "is_healthy": True,
            "details": {}
        }

        # Ensure we have a baseline heartbeat
        if ArtifactType.HEARTBEAT not in self.last_artifacts:
            status["is_healthy"] = False
            status["details"]["reason"] = "No heartbeats recorded"
            return status

        for art_type, threshold in self.thresholds.items():
            last_time = self.last_artifacts.get(art_type, 0)
            age = now - last_time

            if art_type == ArtifactType.HEARTBEAT:
                # Heartbeats are critical for liveness
                if age > threshold:
                    status["is_healthy"] = False
                    status["details"][art_type.value] = {
                        "age": age,
                        "status": "STALE"
                    }
            else:
                # For functional artifacts, staleness indicates state
                if age > threshold:
                    status["details"][art_type.value] = {
                        "age": age,
                        "status": "AT_RISK" if age < threshold * 2 else "STALLED"
                    }
                    # Note: Stale artifacts don't necessarily mean the agent is dead,
                    # but they indicate a lack of progress.

        return status
Enter fullscreen mode Exit fullscreen mode

Integrating with an Orchestrator

The orchestrator consumes these check_health responses or the raw event stream. Here is how a simple orchestrator loop might interpret these signals to make scaling decisions.

def orchestrator_decision(agent_id: str, health_status: Dict[str, Any]):
    """
    Logic for the orchestrator to act based on artifact age.
    """
    if not health_status["is_healthy"]:
        # Immediate action: Restart or kill if heartbeat is missing
        print(f"[{agent_id}] CRITICAL: Heartbeat missing. Triggering restart.")
        return "RESTART"

    details = health_status["details"]

    # Check for "Stalled" functional artifacts
    stalled_types = [
        k for k, v in details.items() 
        if isinstance(v, dict) and v.get("status") in ["STALLED", "AT_RISK"]
    ]

    if "final_output" in stalled_types:
        # The agent is processing something but hasn't finished.
        # This is distinct from being "dead".
        print(f"[{agent_id}] WARNING: No final output in a while. Monitoring token spend.")
        # Action: Tag for observability dashboards, potentially adjust SLA
        return "MONITOR"

    if "tool_invocation" in stalled_types:
        print(f"[{agent_id}] INFO: Stuck in tool invocation. Possible external dependency issue.")
        return "CHECK_DEPENDENCIES"

    return "HEALTHY"
Enter fullscreen mode Exit fullscreen mode

Performance Considerations and Best Practices

Implementing Artifact Age introduces new vectors for performance degradation if not done carefully. Adhere to these principles:

  1. Async Publishing: The record() method must never block the agent's main thread. Use a thread-safe queue or an asynchronous message bus (e.g., Kafka, Redis Streams) to decouple artifact production from consumption.
  2. Payload Sanitization: Never send full payloads (like large document bodies or API keys) in artifact events. Send metadata (hash, size, type) instead. This reduces bandwidth and prevents leakage of sensitive data into observability logs.
  3. Tiered Thresholds: Not all agents are the same. Define thresholds dynamically. A "simple" agent that returns a lookup result should have a strict 50ms threshold. A "complex" agent that performs multimodal analysis might have a 5-second threshold. Allow the agent to self-report its expected turnaround time.
  4. Clock Skew: If agents run in distributed environments (containers, serverless), NTP drift can cause false positives. Use a central time source for the orchestrator's calculations, or allow a small epsilon (e.g., 100ms) in threshold checks.

Conclusion: The Shift from "Are You Alive?" to "Are You Useful?"

The transition from Liveness to Artifact Age represents a maturation in how we observe intelligent systems. Liveness is a binary, blunt instrument designed for the deterministic world of 1995. Artifact Age is a nuanced, continuous spectrum designed for the probabilistic world of 2025.

By monitoring what an agent produces rather than just if it runs, we gain the ability to:

  • Detect "zombie" processes that consume resources but deliver no value.
  • Differentiate between complex thinking and hung processes.
  • Proactively manage dependencies before they cause total failure.
  • Correlate technical health with business impact (token spend, latency, accuracy).

As agent architectures become more complex, with multi-agent systems and long-running tasks, "Artifact Age" will become the standard metric for operational reliability. It bridges the gap between infrastructure observability and application intelligence, giving engineers the context they need to build systems that are not just alive, but alive and doing the right thing.

For those ready to experiment, start by adding a simple last_output_timestamp field to your existing agent state. Then, ask yourself: when was the last time this agent did something that mattered? If the answer is "I don't know," you are ready to build out Artifact Age monitoring.

Top comments (1)

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

Artifact age is a great signal because a hung agent loop still looks alive to a PID check while producing nothing. The tricky case is an agent that keeps writing artifacts that are wrong, so freshness alone can hide a retry storm. Do you pair artifact age with any content check, like diff size or schema validation, before calling a run healthy?

iin1006h05