DEV Community

sword luan
sword luan

Posted on AI-assisted

Tell Me More (叙能): Sanity Challenge Path One Submission

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.

What I Built

The product definition I am preserving is:

Tell Me More (叙能) is an intelligent system that triggers authentic memories through real objects, uses a proactive AI Agent to reconstruct real decision-making processes through interviews, accumulates authentic personal data over time, and uses algorithms to analyze and empower.

The three non-negotiable “real” layers are:

  • real objects / 真实物件
  • real memories / 真实回忆
  • real decisions at that time / 真实当时决策

The core chain is:

object → memory → decision → data → analysis → empowerment

The Agent records, organizes, compares, and presents.

The final choice always remains with the user.

That is the product.

It is not a generic memory chatbot, and it is not a system that simply searches old notes.

The product must make six things actually work together:

  1. Real artifacts
  2. Traceable personal history
  3. Active gap interviews
  4. The user’s “best decision at that time”
  5. Structured sealing
  6. Similar-history Empowerment

1. Real artifacts are the entry point

A historical event begins from a real object or record when one exists: a message, document, image, video, link, spreadsheet, or another item that can anchor a memory.

The first action is to preserve the original.

The original artifact is not rewritten by later interpretation.

If no artifact exists, the user can start from a confirmed life anchor, but that source is explicitly stored as later recall rather than being presented as contemporaneous evidence.

Tell Me More keeps three layers separate:

  • the original artifact;
  • the user’s own recollection;
  • system organization such as classification, indexing, statistics, retrieval, and comparison.

The system never lets one layer silently overwrite another.

2. The Agent interviews for missing history, not for a predetermined story

The user speaks freely first.

The Agent does not interrupt to force every person through the same questionnaire.

It records the user’s words, identifies which historical decision fields are already present, and asks only about genuine gaps.

Every question must come from:

  • the source artifact;
  • a missing field;
  • an information gap;
  • content that cannot yet be classified.

If there is no gap, there is no question.

The Agent is not allowed to invent a motive, emotion, fact, or conclusion.

And uncertainty is valid historical data:

  • unknown;
  • forgotten;
  • cannot judge;
  • not applicable.

“No data” stays “no data.”

3. A historical event is sealed as a traceable decision record

Once the user confirms the event, it becomes a sealed historical record.

The record preserves the user’s own words and the source behind each structured field.

A decision record can include:

  • what happened;
  • the user’s role and position;
  • what was known at the time;
  • what options were visible;
  • goals;
  • constraints;
  • resources;
  • social and technical conditions;
  • the actual choice;
  • why the user made that choice;
  • the user’s best decision at that time;
  • the real outcome;
  • later reflection and evaluation.

“Best decision at that time” is not an AI recommendation.

It means:

Given the information, cognition, visible options, goals, capabilities, resources, and constraints available at that historical moment, what did the user believe was the best or necessary choice?

Later information does not rewrite that historical judgment.

4. History is append-only

If the user remembers something new later, Tell Me More does not silently rewrite the sealed event.

It appends a new record and links it back to the original event.

This preserves a traceable history of:

  • original evidence;
  • original recollection;
  • later supplements;
  • later evaluation.

The system can accumulate years of personal history without turning the archive into one repeatedly rewritten autobiography.

5. Empowerment is not the first answer

The product constitution defines Empowerment very specifically:

Empowerment is not the first answer. It is a second reference.

The user may first use a general AI or another tool to understand the current world.

Tell Me More is the vertical personal-history system.

When a general answer is not enough for a personal choice, the user can press Empower and provide four things:

  1. what happened;
  2. why the choice must be faced now;
  3. what choices are currently visible;
  4. where the user is stuck.

Then the Agent searches the user’s own historical database.

6. What Empower actually returns

The background flow is:

current event
  ↓
extract current conditions
  ↓
retrieve the same / similar historical events
  ↓
read the user’s past choices, reasons, outcomes, and later evaluations
  ↓
compare similarities and differences
  ↓
present traceable personal-history evidence
Enter fullscreen mode Exit fullscreen mode

The output is not “Choose A.”

It provides:

  • how many similar historical events were found;
  • which events are closest;
  • when they happened;
  • what the user chose each time;
  • why the user chose that way;
  • what actually happened afterward;
  • how the user later evaluated the decision;
  • what is similar and different about the current situation;
  • the source event behind every claim.

The Agent can aggregate repeated historical patterns, but it does not turn them into destiny.

History is evidence.

The user decides.

Demo

Live public review build

Live Demo: https://tell-me-more-web.vercel.app/

The public review build now uses my real personal history, with my explicit authorization to disclose it publicly for this challenge.

I chose not to hide behind synthetic examples. I want the Agent to be judged on the material it was designed for.

The public Sanity archive currently contains:

  • 126 sealed real personal-history documents
  • 23 structured historical events
  • 56 verbatim memory statements
  • 16 artifacts / life anchors
  • 26 cognition planes
  • 1 real T0 baseline

The three earlier synthetic partnership events have been removed from the public dataset.

The Archive page exposes the complete saved verbatim corpus, including memory statements that are not yet linked to a structured event. The UI keeps verbatim recall separate from system classification, so judges can inspect what I actually said and what the application derived from it.

Real personal-history archive and full verbatim corpus

A real Empower session

The fixed public Empower question is not a made-up example. It comes from my May 28 personal-history session.

The current situation begins:

今天我把已经开了大约十年的网约车,已经正式结束了。所以我今天也面临着一个重大选择

The options and constraints in that same saved record include returning to Douyin, returning to my hometown, looking for employment, or pursuing YouTube, along with my uncertainty about whether ending ride-hailing was the right choice.

The public request runs through the live Sanity Context / Knowledge Base path over my real historical archive.

Verified public result:

  • retrievalMode: context
  • 18 current Context candidates returned
  • 3 structured matches retained after decision-time reranking

The three current references are:

  1. 2023 — tried publishing edited videos with the goal of reaching ad-revenue conditions
  2. 2022 — ran game-farming and video-editing income attempts in parallel
  3. 2022 — bought a computer and paid to learn video editing

The result then shows, for each historical event:

  • what I chose at the time;
  • why I chose it;
  • what actually happened;
  • how I later evaluated it when that evaluation exists;
  • the source event and provenance;
  • similarities and differences from the current situation.

Real personal-history Empower result

What Empower now does with those three references

Retrieving history is not enough.

After Context recall and structured reranking, the application runs Empower Analysis V1. It converts the matched personal history into two different kinds of support:

Cognition empowerment — improve how the current decision is framed and validated.

Capability empowerment — improve whether the chosen direction can actually be executed, tested, funded, and stopped safely.

For this real session, the current output is:

  • 8 / 8 current decision dimensions recorded;
  • current visible-option breadth: 5 explicit paths, compared with a maximum of 1 in the currently matched historical events;
  • 3 matched historical cognition planes linked to the retrieved events;
  • a direct T0 ↔ T+1 longitudinal comparison across 20 cognition dimensions;
  • 2 evidence-linked decision-improvement actions.

Those two current actions are:

  1. replace “judge from early feedback” with explicit observation periods, validation metrics, and stop conditions;
  2. stage investment and define an investment ceiling / validation period / exit condition.

Each action is generated only when the relevant historical or current evidence exists, and the UI exposes that evidence.

The algorithm also keeps cognition, execution capability, resources, and constraints separate. It does not output one “growth score,” and it does not convert historical similarity into a prediction.

The longitudinal proof: T0 → T+1

The challenge build now tests the part of the product that normally requires time.

Because a competition prototype cannot wait four real months after onboarding, I used a creator test timeline built from my own real history:

  • T0: May 28, 2026 — a reconstructed baseline interview for the day I ended roughly ten years of ride-hailing and faced a major direction change;
  • T+1: September 24, 2026 — a new current-state interview recorded after four months of actual work, experiments, tool use, publishing, and business attempts.

The current T+1 plane is not allowed to silently replace T0. It is a new sealed record.

The 20-dimension comparison currently reports:

  • 7 dimensions where both T0 and T+1 have recorded evidence and the recorded text changed;
  • 13 dimensions that were not recorded at T0 but are recorded at T+1;
  • 0 growth score.

Those 13 newly recorded dimensions are not called “improvement.” They are simply increased evidence coverage.

This matters because Empower must not treat an old version of me as the current version of me. Early T0 statements such as missing tool ability or an undefined first step must not remain current-state constraints after later evidence shows otherwise.

I also found and fixed a sealing bug during this test: the first T+1 plane had been sealed before my explicit confirmation. I did not overwrite or delete it. The archive now keeps that original record, appends the real “确认T+1” confirmation evidence, and uses a confirmed revision that supersedes the premature version. The repository now rejects any future-observed cognition plane that does not reference stored verbatim confirmation evidence.

That correction is itself part of the product proof: history is append-only even when the system corrects itself.

Real Empower Analysis V1: decision improvement plus cognition/capability evidence

This is what “Empower” means in this submission: past decisions do not merely reappear; they change what the current decision process checks before the user chooses.
That is the product proof I wanted:

A current personal choice can call up my own accumulated, traceable history — not a fabricated profile and not a generic world answer.

90-second judge walkthrough

  1. Open the live demo.
  2. Confirm the banner says REAL PERSONAL HISTORY · OWNER AUTHORIZED PUBLIC DISCLOSURE.
  3. Read the product chain: Real Object → Real Memory → Real Decision → Data → Analysis → Empower.
  4. Open Archive.
  5. Inspect the real structured events and the Full Verbatim Corpus with all 54 saved memory statements.
  6. Open a memory statement and compare my exact words with the structured event/classification layer.
  7. Open Empower.
  8. Run the live Context query.
  9. Confirm retrievalMode = context, 18 current Context candidates, and the three retained historical matches.
  10. Read Empower Analysis V1: 8/8 current decision dimensions, 5-vs-1 visible-option breadth, and the two evidence-linked decision-improvement actions.
  11. Inspect T0 ↔ T+1: 7 directly comparable dimensions changed, while 13 dimensions are newly recorded at T+1 and are not claimed as growth.
  12. Then inspect Personal History Basis and expand a match for field-level similarities, differences, and provenance.

Public disclosure and write boundary

This is intentionally unusual for a personal-history product: I explicitly authorized public disclosure of this challenge corpus because I want the prototype to prove itself on real history.

The public deployment is still read-only:

  • judges can read the real archive;
  • judges can inspect all saved verbatim memory statements;
  • judges can run the fixed real Empower question;
  • public writes are rejected server-side;
  • the public Empower result is non-persistent;
  • the Context organization token remains server-side.

The public corpus is evidence for this challenge, not a claim that future users should make their private history public.

Code

GitHub: https://github.com/swordluan8-hash/tell-me-more

Current verification:

  • 63 deterministic / rule tests passing;
  • 4 browser E2E flows passing;
  • TypeScript typecheck passing;
  • ESLint passing;
  • production build passing;
  • GitHub Actions CI passing;
  • public Vercel deployment passing.

Architecture

Real artifact / confirmed life anchor
      ↓
verbatim recall
      ↓
gap audit
      ↓
user-confirmed structured historical event
      ↓
sealed Sanity Content Lake record
      ↓
long-term personal archive

Current choice
      ↓
current-condition extraction
      ↓
Sanity Context + Knowledge Base recall
      ↓
candidate historical events
      ↓
complete Content Lake read-back
      ↓
structured comparison
      ↓
past choices + reasons + outcomes + later evaluation + provenance
      ↓
user makes the final choice
Enter fullscreen mode Exit fullscreen mode

The archive keeps source layers separate

The application models the original artifact, memory statements, historical events, cognition planes, baseline state, and Empower sessions as separate document types.

This is necessary because a personal-history system must be able to say where each claim came from.

A system-derived classification is not allowed to become a new historical fact.

Unknown is a first-class state

The schemas and workflow explicitly preserve:

  • unknown;
  • forgotten;
  • cannot_judge;
  • not_applicable.

The Agent does not complete history just because a field is empty.

Sealed history rejects replacement

The application repository exposes append behavior for historical records and rejects duplicate-ID replacement.

Later recall is written as a new supplement linked to the existing event.

The browser and unit tests verify that sealed history is not silently rewritten through the normal application path.

How I Used Sanity

1. Sanity Content Lake stores the formal personal-history archive

The product uses Sanity for the confirmed structured archive, not as a generic chat log.

The challenge dataset contains six core document categories:

  1. artifact
  2. memoryStatement
  3. historicalEvent
  4. cognitionPlane
  5. baselineT0
  6. empowermentSession

The public review dataset now contains the owner-authorized real corpus: 126 sealed personal-history documents, including 23 historical events, 56 verbatim memory statements, 16 artifacts / life anchors, 26 cognition planes, and 1 T0 baseline.

The current Knowledge Base dataset import reads the 23 real sealed historical events (demo == false). The latest ingestion fetched and distilled 23 / 23 events with 0 failed.

2. Sanity Context and Knowledge Base retrieve relevant personal history

Knowledge Base ID: kbrqT3iILYmW

When the user presses Empower, the application turns the current choice into a temporary structured representation and calls Sanity Context / Knowledge Base for historical recall.

Context is used to find relevant history.

The application then maps those candidates back to the complete sealed Content Lake records before presenting the personal-history evidence.

3. Structured comparison remains explainable

The first version does not use one opaque embedding score as the final answer.

The candidate recall can use Sanity Context’s semantic capability, but the application shows a structured comparison across recorded fields such as:

  • event / problem domain;
  • role and responsibility;
  • goal;
  • constraints and resources;
  • visible options;
  • information available at the time;
  • social environment;
  • technical / era conditions.

Missing fields reduce the amount of comparable evidence.

They are not filled by AI.

4. The result includes the parts the product constitution requires

For each related historical event the interface can show:

  • event date;
  • past choice;
  • reason;
  • outcome;
  • later evaluation;
  • similarity and difference evidence;
  • source references.

The product can therefore answer a personal-history question that a generic current-world model cannot:

“When situations like this happened in my own life, what did I actually choose, why, what happened afterward, and how did I later evaluate those choices?”

It still does not convert that evidence into “therefore choose A.”

Sanity Project Details

What I Learned / Design Decisions

The product is longitudinal, not horizontal

Tell Me More is deliberately about the user’s own longitudinal history.

It does not try to replace the current-world information that another AI, professional, market tool, legal source, or search system can provide.

The product’s role is different:

bring the user’s own recorded history into the moment when the user is making a new choice.

Recording and advising are separate responsibilities

The recording side:

  • preserves evidence;
  • records verbatim recall;
  • identifies gaps;
  • asks neutral gap questions;
  • seals confirmed records.

The Empower side:

  • retrieves relevant personal history;
  • compares recorded facts;
  • presents choices, outcomes, later evaluations, similarities, differences, and provenance.

Neither side is allowed to manufacture facts.

The Agent is a staff adviser, not the final decision-maker

The product can compare and summarize the user’s recorded history.

It cannot decide for the user.

The final constitutional rule remains:

History does not decide the user’s future. History should appear when the user makes a choice and provide real, traceable personal evidence for that choice.

Limitations

This challenge build is deliberately a single-user prototype.

  • The current public corpus is real personal history disclosed with the owner's explicit authorization; it is still only a partial lifetime archive, not a claim that every life event has been captured.
  • 22 provenance references inside one lending/investment event point to two child-record IDs that were never created by the earlier import workflow. I did not fabricate those missing records; Sanity preserves those references as weak references.
  • The Knowledge Base currently reports two review issues even though the real-data build completed; I do not treat “build succeeded” as proof that every historical record is semantically complete.
  • The final reranker is a transparent, conservative structured lexical matcher after Context recall. It is not a predictive scientific model and does not claim that historical similarity determines the best future action.
  • Images are currently sealed with metadata / hashes; complete media understanding is outside this challenge slice.
  • Application-layer sealing is not provider-enforced WORM storage.
  • The Agent does not make the final choice for the user.

Agent Session

Agent Session is optional for this challenge.

I used Codex during implementation and testing, but I am not attaching a raw development transcript because it contains unrelated local operational context.

The public repository, CI, deterministic tests, live read-only demo, and Sanity-backed retrieval path are the reproducible evidence for this submission.

Final links

Live Demo:

https://tell-me-more-web.vercel.app/

GitHub:

https://github.com/swordluan8-hash/tell-me-more

Sanity Project ID:

3tdecpiq

Knowledge Base ID:

kbrqT3iILYmW

Top comments (0)