DEV Community

pawel jankes
pawel jankes

Posted on AI-assisted

Before the Mexico Benchmark: CZARA, Offline Director, Independent ISKRA Paths, and 24 Stages of Web Training

I have just preregistered the next research phase of SSI V5 before the first official robotics benchmark with a research team in Mexico.

This matters because I do not want to describe the system only after seeing a successful result. I want the repository history to show what was planned, what was already implemented, which hypotheses will be tested, and which claims remain open.

SSI V5 is a solo R&D project, but it is organized as a set of separate runtime roles rather than one large chatbot. The public repository is an evidence and review mirror. The proprietary implementation remains private.

The new preregistration separates three different paths:

  1. the core SSI training sequence, S11-S40;
  2. a separate WEB01-WEB24 engineering track;
  3. the Mexico robotics and offline research programme.

They are related, but they are not the same experiment.

1. Core SSI training: S11-S40

The core sequence trains seven independent execution lines:

  • BODY_FROZEN;
  • ISKRA1;
  • ISKRA2;
  • ISKRA3;
  • ISKRA4;
  • ISKRA5;
  • ISKRA6.

Each stage preserves PASS, FAIL and INCONCLUSIVE outcomes. Only the verified subset is eligible for cross-consolidation.

The intended stage transition is:

stage execution
-> preserve all outcomes
-> export verified subset
-> consolidate compatible competence
-> BODY_FROZEN + DIRECTOR report completion
-> COMMITTED gate
-> next stage
Enter fullscreen mode Exit fullscreen mode

BODY_FROZEN and DIRECTOR receive the same compatible skill payload and Champion state, but they do not become the same agent.

BODY_FROZEN is the technical execution core. DIRECTOR is the planning, coordination, resource and research-communication core. They keep separate runtime, memory, identity and decision topology.

Consolidation transfers verified competence, LEGO packages, evidence and provenance. It does not transfer personality, merge private memory, retrain neural weights, or import the BODY core into DIRECTOR.

The public front door currently contains detailed evidence through S12. Internally, the owner-reported runtime status includes committed consolidations through S16 and later S-stage activity including S19. The later evidence package will be published after the uninterrupted sequence closes or stops.

2. Champion-Challenger is not a “latest answer wins” mechanism

A new candidate does not replace an older solution just because it appeared later.

A reusable candidate is expected to carry:

competence
+ input/output contract
+ LEGO or micronetwork artifact
+ LAB result
+ evidence
+ provenance
+ version identity
Enter fullscreen mode Exit fullscreen mode

It may remain in HOLD until there is a real reason to compare it against an existing Champion.

An earlier S16 snapshot reported complete accounting for 875 BODY_FROZEN candidates:

CHAMPION   = 1
CHALLENGER = 1
HOLD       = 873
SPECIALIST = 0
ACCOUNTED  = 875 / 875
Enter fullscreen mode Exit fullscreen mode

This does not mean there were 875 Champions or 875 active duels. It means SSI retained a large pool of traceable potential competence instead of deleting everything except the current winner.

The formal software-engineering Champion in that snapshot was s10_9769b8e74e6230f29df90218568d4ecf. The active Challenger was s10_97a39fdb193470923d6de3f5f0764412. Both had disclosed validation PASS, accuracy around 0.85 and F1 around 0.82.

Because the disclosed metrics were similar, novelty alone was not enough to promote the Challenger.

3. Reuse before expensive reasoning

The routing objective is simple:

problem
-> search existing competence
-> strong match: REUSE_TOP1
-> weak match: ESCALATE_FULL_FLOW
Enter fullscreen mode Exit fullscreen mode

One observed reuse example had confidence 0.71. A weak match around 0.3065 escalated to FULL FLOW instead of pretending an existing Champion was suitable.

This is important economically. During later training I observed decreasing paid-model token use. That reduction is an observed runtime result, not only a future hypothesis.

What is not yet proven is the complete causal explanation. The decrease may involve validated reuse, cache, local models, provider routing, shorter paths or other factors.

The final analysis will therefore connect:

  • paid input and output tokens per case;
  • cost per verified PASS;
  • model/provider class;
  • daily cost limit;
  • REUSE_TOP1 count;
  • intermediate verification count;
  • FULL FLOW escalation count;
  • PASS, FAIL and INCONCLUSIVE outcomes;
  • time and retries.

The stronger result would be:

lower paid-token use
+ stable or improved verified quality
+ more validated reuse
+ fewer unnecessary FULL FLOW escalations
Enter fullscreen mode Exit fullscreen mode

4. The ISKRA lines are independent trajectories

ISKRA is not only a different label attached to the same agent.

Each ISKRA preserves its own:

  • training history;
  • PASS, FAIL and INCONCLUSIVE path;
  • candidate pool;
  • Champion and Challenger evolution;
  • micronetwork connections;
  • reuse and escalation history;
  • dynamic affective-state trajectory.

The affective layer is a computational state mechanism. I am not claiming consciousness or subjective experience.

These state variables can change after decisions, successes, failures, uncertainty and feedback. They may later modify risk tolerance, escalation behavior, planning, reuse or Champion selection.

This creates an important contrast.

BODY_FROZEN accumulates consolidated verified competence. An ISKRA may preserve a narrower but highly specialized solution that never became the global BODY_FROZEN Champion.

A future domain-specific task may therefore be solved better by an independent ISKRA trajectory than by the globally consolidated core.

Frozen checkpoints make longitudinal comparison possible:

initial state
-> post-S16
-> post-S40
-> post-WEB24
-> final replay
Enter fullscreen mode Exit fullscreen mode

Earlier FAIL and INCONCLUSIVE cases can be replayed after the complete curriculum. The test is not only whether the outcome changes, but why:

  • a different Champion;
  • a different micronetwork;
  • a changed affective state;
  • more reuse;
  • fewer paid tokens;
  • a different risk decision;
  • or a genuinely new solution.

Where technically possible, I also want to separate competence evolution from affective-state evolution by replaying matched variants:

initial competence + initial affective state
final competence + final affective state
final competence + initial affective state
initial competence + final affective state
Enter fullscreen mode Exit fullscreen mode

5. WEB01-WEB24 is a separate engineering school

S40 is not the end of training.

A post-S40 hook is installed and self-tested. Its declared gate is:

S40 execution_complete = true
AND S40 consolidation = COMMITTED
AND WEB LEGO = READY
-> WEB01 -> ... -> WEB24
Enter fullscreen mode Exit fullscreen mode

The WEB track currently reports readiness metadata:

24 stages
192 cases
126 LEGO items
15 templates
7 actors
SYNTHETIC_ONLY data policy
Enter fullscreen mode Exit fullscreen mode

The seven actors are BODY_FROZEN and ISKRA1..ISKRA6.

The progression covers reusable web/programming engineering mechanisms, including HTML/CSS decomposition, routing, CRUD, authentication, REST APIs, FastAPI, localization, security and a capstone.

It does not use real company information as training data. The intended boundary is mechanism and architecture extraction with neutralized or synthetic fixtures.

After WEB24, all seven lines can receive the same previously unseen specification for a page, dashboard or application.

The functional objective will be identical. The final implementation may not be.

I want to compare:

  • visual design;
  • information architecture;
  • component structure;
  • accessibility and security;
  • selected LEGO and templates;
  • Champion/Challenger identity;
  • number of iterations;
  • token use, time and cost;
  • functional test results;
  • blind human visual evaluation.

The research hypothesis is that independent training, micronetwork, Champion-Challenger and affective-state histories may produce measurable differences in architecture, interaction design and visual finish despite an identical functional objective.

Different-looking outputs are not guaranteed. They must be observed and measured.

6. Mexico is a separate robotics research path

The Mexico programme is not WEB25 and it is not another name for the core training.

The planned curriculum contains 48 stages:

38 training stages
4 validation stages
6 capstone stages
Enter fullscreen mode Exit fullscreen mode

It covers drones, humanoids, cross-domain coordination, offline execution, LEGO Navigation, LEGO Space, no-network map exchange and later unseen partner-defined scenarios.

The first Mexico benchmark will intentionally use a simpler baseline:

DIRECTOR
+ BODY_FROZEN
+ OFFLINE_DIRECTOR
+ CZARA
+ LEGO Navigation
+ LEGO Space
+ LAB
+ EVIDENCE
Enter fullscreen mode Exit fullscreen mode

ISKRA agents will not participate in the first baseline.

The first question is narrower:

Can OFFLINE_DIRECTOR execute a bounded mission, reuse validated competence, preserve local state and evidence, build or exchange spatial LEGO fragments, and synchronize after reconnection?

Only after this mechanism works will I compare OFFLINE_DIRECTOR and BODY_FROZEN Offline against independent ISKRA Offline variants.

7. CZARA is context, not control

CZARA is a multilingual research-context and translation layer between the Mexico research environment and SSI.

It should preserve:

  • original Spanish or English text;
  • Polish translation;
  • source session;
  • verified speaker identity when available;
  • uncertainty when identity is unknown;
  • requirements, corrections, decisions and open questions;
  • confidence and provenance.

CZARA does not directly command robots and does not replace DIRECTOR.

The important distinction is:

verified professor message
!=
unstructured conversation context
Enter fullscreen mode Exit fullscreen mode

A verified message from the professor has known authorship. A live conversation may be a “context soup” containing the professor, students, technical staff and incidental comments.

SSI must not treat every sentence in that stream as an instruction.

The intended chain is:

source
-> original text
-> translation
-> structured context
-> DIRECTOR interpretation
-> versioned decision
-> BODY_FROZEN execution
-> LAB result
-> evidence
Enter fullscreen mode Exit fullscreen mode

If authority or intent is uncertain, DIRECTOR should ask the professor for confirmation before changing a frozen mission.

This is the human-AI co-learning question behind CZARA: can an external team teach and correct the system through conversation without collapsing human discussion into unrestricted execution authority?

8. Two simultaneous R&D conversations

The Mexico panel is designed around two auditable chats operating at the same time.

Professor to DIRECTOR

The professor discusses objectives, observations, evidence and benchmark revisions with DIRECTOR.

DIRECTOR explains what the system is doing, why a test stopped, what evidence exists, and what revision is proposed.

Authorized technical staff to BODY_FROZEN

An authorized technical operator can communicate directly with BODY_FROZEN for diagnosis, bounded technical requests and immediate STOP, PAUSE or SAFE STATE actions.

This avoids unnecessary latency during an active experiment.

DIRECTOR still observes BODY_FROZEN state and events. If BODY_FROZEN stops, DIRECTOR can explain the reason and consequence to the professor without requiring the technical team to manually relay every detail.

The authority boundary is strict:

STOP / PAUSE / SAFE STATE
= direct when authorized

change of benchmark objective or acceptance criteria
= versioned research revision
Enter fullscreen mode Exit fullscreen mode

The technical chat cannot silently rewrite a frozen benchmark.

This is closer to a real R&D laboratory than a single chat window: the research conversation and technical execution conversation can proceed in parallel while sharing an auditable event chronology.

9. Later offline comparison

After the baseline, the comparison may include:

OFFLINE_DIRECTOR with consolidated competence
vs
BODY_FROZEN Offline
vs
ISKRA1..ISKRA6 Offline
Enter fullscreen mode Exit fullscreen mode

One preregistered question is:

Can an independent ISKRA trajectory, using its own training history, affective-state evolution, micronetworks and Champion-Challenger state, outperform the consolidated baseline in selected LEGO Navigation or LEGO Space tasks without internet access?

For an example ISKRA6 comparison:

H0: consolidated OFFLINE_DIRECTOR/BODY_FROZEN is equal or superior.
H1: ISKRA6 is better on predefined quality, stability or efficiency metrics.
Enter fullscreen mode Exit fullscreen mode

No ISKRA is declared the winner in advance.

10. What already exists and what does not

The current owner estimate for internal Mexico implementation and training readiness is approximately 85%.

That is an engineering estimate, not a benchmark score.

Already implemented or prepared elements include the Mexico panel, separate DIRECTOR and BODY_FROZEN chats, contextual memory, role separation, consolidated-skill compatibility, Champion-Challenger mechanisms, laboratory/evidence foundations and the CZARA/offline architecture.

Remaining work includes final integration, authorization, shared event chronology, stop/pause verification, evidence-completeness tests, offline synchronization checks, an internal rehearsal and version freeze.

The official external Mexico benchmark completion remains 0%. Independent external validation has not yet been performed.

11. Why preregister this now?

If the Mexico collaboration later produces seven or ten benchmark revisions, a future reviewer should be able to go back in Git history and see that:

  • the architecture existed before the result;
  • negative results were expected to be preserved;
  • the baseline excluded ISKRA agents by design;
  • later ISKRA comparisons were planned in advance;
  • CZARA was context rather than execution authority;
  • token reduction was measured separately from its causal explanation;
  • web artifacts were intended for blind comparison after WEB24;
  • FAIL and INCONCLUSIVE were not supposed to disappear.

The strongest evidence will not be a claim that SSI is “professional.”

It will be the chronology:

preregistration
-> frozen protocol
-> execution
-> PASS / FAIL / INCONCLUSIVE
-> evidence
-> revision
-> replay
-> transfer to a new domain
Enter fullscreen mode Exit fullscreen mode

The public preregistration is available here:

https://github.com/jankes72/SSI_V5/blob/main/MEXICO_PREBENCHMARK_RND_PROTOCOL_20260928.md

The Mexico robotics plan is here:

https://github.com/jankes72/SSI_V5/blob/main/MEXICO_ROBOTICS_TRAINING_AND_BENCHMARK_PLAN_20260926.md

The public repository contains protocols, sanitized results, failures, repairs, hashes, provenance and claim boundaries. It does not publish proprietary implementation, private prompts, credentials or reconstructive internals.

This protocol does not claim completed S13-S40 public evidence, completed WEB01-WEB24 training, completed Mexico validation, physical robotics validation, safety certification, AGI, consciousness or superiority over external R&D teams.

It records the questions before the answers are known.

Top comments (2)

Collapse
 
amnezja3 profile image
Amn Tree •

Outstanding architectural discipline. Separating the execution core BODY_FROZEN, from the planning and communication layer DIRECTOR while only sharing validated competency packages is a brilliant way to prevent cognitive drift.
By maintaining distinct runtime topologies and private memories, you avoid the unpredictability that usually happens when merging neural states.

I'm highly interested to see how this modular LEGO package approach handles the physical constraints of the Mexico robotics program without retraining weights.

Best of luck with the benchmark!

Collapse
 
jankes72 profile image
pawel jankes •

Thank you, Amn Tree, for the kind words and for reading the architecture so closely. Keeping BODY_FROZEN and DIRECTOR separate while transferring only traceable, validated competence is central to the design. I appreciate your interest in the LEGO approach. The Mexico work should test how it handles actual robotics constraints under partner-defined criteria. With the partner's approval, I hope to share the results and limitations, including failures or inconclusive outcomes. Thanks again for the encouragement!