DEV Community

Cover image for I built maarg: Python experiment tracking with zero logging boilerplate
Moazzam Matin
Moazzam Matin

Posted on

I built maarg: Python experiment tracking with zero logging boilerplate

Every time I work on quick machine learning experiments or parameter sweeps, I hit the same workflow friction: experiment tracking instrumentation.

Setting up traditional experiment loggers usually requires running local daemon servers, managing URIs, and scattering explicit logging statements (log_metric(), log_param(), log_artifact()) across internal function logic.

I built maarg (मार्ग — Hindi for "path") to test a different approach: function-boundary tracking.


💡 The Core Idea: Intercept at Function Boundaries

Instead of forcing you to write logging calls inside your functions, maarg uses a single decorator (@track). It uses inspect.signature to automatically bind positional arguments, keyword arguments, and parameter defaults, while capturing return values and timing upon execution.

from maarg import track, get_runs, top_n

@track(experiment="learning-rate-sweep")
def fit(learning_rate, epochs=100):
    w = 0.0
    for _ in range(epochs):
        grad = sum(2 * (w * x - 3 * x) * x for x in range(1, 6)) / 5
        w -= learning_rate * grad
    return {"error": abs(w - 3)}

# Run hyperparameter trials without writing tracking calls
for lr in (0.001, 0.003, 0.01):
    fit(learning_rate=lr)

# Query results directly from local storage
for run in top_n(get_runs(), "error", n=2, higher_is_better=False):
    print(f"lr={run.inputs['learning_rate']} error={run.metrics['error']:.2e}")
Enter fullscreen mode Exit fullscreen mode

⚡ Technical Design Decisions

  1. Zero External Server Setup: Every run serializes locally into a single SQLite database (.maarg/runs.db).
  2. Zero Required Dependencies: The core library relies strictly on Python's standard library (sqlite3, inspect, json, dataclasses).
  3. Automatic Output Classification: Return dictionaries containing numeric scalars become metrics, Matplotlib figures are serialized to PNG artifacts (if matplotlib is installed), and other returns are captured into an execution ledger.
  4. Transparent Exception Handling: If your function fails, maarg captures the stack trace and failure status into SQLite, then re-raises the original exception so application behavior isn't swallowed.

🛠️ Early Alpha & Building in Public

maarg is currently in early alpha (v0.2.0.post1). The core architecture works and passes its test suite, but building this highlighted several key engineering problems I'm solving for v0.3.0:

  • Failure Isolation: Making tracking errors non-fatal so capture failures never break the underlying user experiment.
  • Serializer Registry: Moving away from hard-coded type checks into an extensible serializer registry for custom types (NumPy, Pandas, PyTorch).
  • Artifact Safety & JSON Hardening: Hardening path sanitization for generated image artifacts and dictionary key inspections.

📦 Try it out

You can test maarg today via PyPI:

pip install maarg
Enter fullscreen mode Exit fullscreen mode

Or with optional Matplotlib artifact rendering:

pip install "maarg[plotting]"
Enter fullscreen mode Exit fullscreen mode

GitHub logo Moazzam-Matin / maarg

Zero-instrumentation experiment tracking for Python — a decorator that auto-captures inputs and outputs, no logging calls required.

maarg

CI License: MIT Python 3.9+ PyPI version

maarg - Zero-Boilerplate Experiment Tracking

Experiment tracking with zero logging code.

Put @track on a function. Every execution—arguments returns, metrics, execution timing, and failures—is automatically saved to local storage for instant querying.


How It Works

maarg sits transparently at function boundaries. It reads signature parameter defaults and runtime return payloads without requiring explicit parameter or metric logging statements inside your function logic.

  ┌────────────────────────┐
  │  @track decorated fn   │  ──► (Intercepts arguments & execution context)
  └───────────┬────────────┘
              │
              ▼
  ┌────────────────────────┐
  │   Function Execution   │  ──► (Captures return dict / scalars / figures)
  └───────────┬────────────┘
              │
              ▼
  ┌────────────────────────┐
  │   SQLite Persistence   │  ──► Saves to .maarg/runs.db (or custom backend)
  └───────────┬────────────┘
              │
              ▼
  ┌────────────────────────┐
  │  Query & Analysis API  │  ──► maarg.get_runs() ──► top_n() / filter_runs()
  └────────────────────────┘

Quickstart

from maarg import track, get_runs, top_n
@track(experiment="learning-rate-sweep")
def fit(learning_rate, epochs=100):
    w = 0.0
    for _ in range
…
Enter fullscreen mode Exit fullscreen mode

I'd love feedback from developers and ML engineers on the API ergonomics or the decorator boundary model! What features or storage backends would you like to see next?

Top comments (0)