I asked an AI agent to analyze 30 days of Ethereum data using an external market-data service. The workflow sounded trivial:
discover → approve → call API → chart → summarize
It wasn’t.
1. Tool access did not mean executable access
The agent found the correct endpoint, but the endpoint required a paid wallet flow that was not available in the runtime. Discovery succeeded; execution failed.
The fix was to separate capability discovery from authorization. In this experiment, Anvita Flow exposed the service scope and per-call price, while execution only continued after a human approved the specific call.
2. A successful response was not trustworthy data
The call returned 721 hourly observations. It also returned two trading-volume values near 1e19—obvious outliers compared with the rest of the series.
If we had rendered the JSON directly, the volume chart would have been useless. We kept the raw values for auditability, flagged them, and excluded them from robust volume statistics.
fetch → validate schema → check units/ranges
→ flag anomalies → compute → render
3. “30 days” had a boundary problem
The final observation arrived at 07:44 UTC, so the last day was incomplete. Comparing it with a full UTC day would distort daily volume and return calculations.
Every report now carries its exact sample window, timezone, observation count, and incomplete-period warning.
4. Correlation tried to become causation
The agent aligned price moves with ETF flows and policy events. But one policy announcement occurred after the dataset ended. Without an explicit cutoff check, it could have become a convincing—but impossible—explanation.
Event enrichment must obey:
event.timestamp <= dataset.cutoff
What the pipeline produced
After validation, ETH showed a +26.23% 30-day return, but a -4.47% seven-day return and an 8.71% pullback from the period high. Hourly data exposed a strong repricing followed by weakening short-term momentum—detail hidden by a simple start/end comparison.
My takeaway: giving an agent external-service access is the easy part. Production reliability comes from approval boundaries, schema validation, anomaly handling, temporal checks, and preserving the raw response for review.

Top comments (0)