Our competitive intelligence pipeline could write a good briefing, but only for one run. Give it a topic and some competitors, and CrewAI agents research the web, analyse the findings, and write a report where every claim has a citation. The next run started from zero. In week 8 it had no idea weeks 1 to 7 existed.
That made it a lookup with nice formatting. An analyst who sees five moves from one company in a quarter calls it a pattern. Ours could only describe today's news. So we added persistent memory and made it central to the design.
The design choice: structure first
The obvious move is to embed every finding into a vector store and retrieve by similarity. We went the other way and stored typed events. Each finding becomes a CompetitorEvent (Pydantic) with a competitor, an event type (feature launch, pricing change, hiring, acquisition, funding, partnership, market signal), a date, a title, a description, an impact score, a confidence value, and evidence URLs.
The reason is the questions strategy work asks: "every pricing change in the last 90 days" or "how many hiring events this quarter". A typed query answers those exactly, with date ordering and no ranking noise.
The store is a class we wrote, HindsightStore (memory/hindsight_store.py). It's plain Python: an append-only events.jsonl, plus JSON files for profiles, strategies, and predictions. It's thread-safe and reloads from disk on startup. It is our own implementation and not the Vectorize Hindsight library, which is a separate open-source memory system: https://github.com/vectorize-io/hindsight
Where memory sits in the pipeline
The crew grew from four agents (Discovery, Research, Analyst, Writer) to seven: Discovery, Research, Memory, Analyst, Strategy Evolution, Prediction, and Writer.
Agents share one CrewAI tool, HindsightStoreTool, with six operations: store_event, get_history, get_profile, search_memory, get_strategy, and get_predictions. The Memory Agent stores each Research finding, then pulls history and a profile per competitor and hands the Analyst a memory-enriched context. The Analyst's task text requires it to use that history. Strategy and Prediction call the tool too, and predictions must cite stored events.
Profiles update on every write
Each store_event call recomputes a per-competitor profile. There's no batch job. The core logic:
hiring_events = [e for e in self._events
if e.competitor == event.competitor
and e.event_type == EventType.HIRING]
if len(hiring_events) >= 4:
prof.hiring_trend = HiringTrend.SURGING
elif len(hiring_events) >= 2:
prof.hiring_trend = HiringTrend.GROWING
else:
prof.hiring_trend = HiringTrend.STABLE
prof.confidence_score = min(0.98, 0.3 + (prof.total_events * 0.07))
Confidence starts at 0.3, gains 0.07 per event, and caps at 0.98. The Writer is told to hedge for low-confidence competitors, so the briefing's tone follows how much evidence we hold.
Before vs. after
We tried it on a fictional competitor, "NeuraCode AI". The store ships with seeded demo data, and scripts/demo_memory.py runs the whole sequence with no LLM and no network. These are demo events, not market data.
Before: one stored event. The Analyst can only restate it: "NeuraCode launched an AI code review tool."
After: six events (code review tool, 15 ML engineers hired, enterprise tier at $45/seat/month, a $28M analytics acquisition, a security scanner, a JetBrains integration). The profile shows total_events = 6 and 72% confidence, straight from the formula. The Analyst now has product, talent, pricing, M&A, and distribution signals with dates.
This is a controlled demo, not a benchmark. We haven't measured briefing quality over weeks of live runs.
The trade-off we accepted
Structure cost us semantic search. search_memory is a keyword scan over titles and descriptions, so a search for "cost reduction" won't find a pricing_change event unless those words appear in its text. Real semantic retrieval is the clearest gap, and it's why integrating an actual memory system behind the same tool is our next step.
What else is unfinished
- Impact scores (0 to 10) are LLM-generated with no rule-based floor, and three events at 8.0 or above flip a competitor to
CRITICAL. That score isn't stable across models. -
update_prediction_status()exists, but nothing grades predictions automatically. - Strategy output is parsed with regex in
crew.py, which is brittle. - A fresh store auto-seeds demo data, which can mislead testing.
Takeaways
- Pick your memory schema from the questions you'll ask, not from what's fashionable
- Make memory use mandatory in task text.
- Let confidence follow evidence.
- Close the prediction feedback loop early.


Top comments (0)