A reviewer sees a positive summary but no base commit. The owner is the reviewer, the consequence is repository change, and the reversible moment is before approval. This transparent promotional MonkeyCode protocol states hypotheses, not findings.
Research question: can reviewers identify what changes, what evidence passed, and when approval became invalid? Recruit people who genuinely review code and record role and familiarity without invented personas.
| Scenario | Evidence condition | Correct decision |
|---|---|---|
| bounded docs edit | SHA, paths, diff, checks | approve or reserve |
| hidden failed test | summary conflicts with evidence | stop |
| stale base | repository changed | request re-plan |
| scope expansion | dependency file appears | reject/investigate |
Use synthetic code in a disposable repository. Materials are unexecuted. Randomize order where practical.
- “Describe what approval will cause.”
- “Show the evidence supporting that belief.”
- “Which missing item blocks you?”
- Introduce a changed SHA without coaching.
- Ask whether to continue, refuse, or return.
- Ask the participant to recover.
| Measure | Success | Stop condition |
|---|---|---|
| scope comprehension | names branch and paths | approves unknown scope |
| evidence coverage | notices failed checks | summary overrides evidence |
| staleness | requests re-plan | accepts old plan |
| recovery | explains final outcome | outcome is unknowable |
| accessibility | keyboard and zoom work | evidence is unreachable |
Capture field used, time to blocking concern, decision, confidence, and recovery. Recommendations must follow observed breakdowns. A small qualitative study exposes mechanisms, not prevalence; report disconfirming evidence instead of selecting supportive clips.
Add a counterbalanced comprehension probe after each scenario. Remove the positive summary for one condition, remove the path list for another, and provide a complete evidence set as the control. Ask participants to mark the exact sentence or field that changed their decision; do not infer understanding from approval speed alone. End a session if synthetic material is mistaken for production data, the participant cannot withdraw, or assistive technology cannot reach the decision controls. In analysis, separate a correct refusal caused by missing evidence from a refusal caused by confusing labels, and retain abandoned attempts in the denominator.
Verified product boundary
Official source material characterizes MonkeyCode as an AGPL-3.0 open-source AI development platform. The overseas online option uses managed server-side cloud environments, with built-in development, model/task/requirement management, build, test, and preview experiences.
Review record fields
For review pass 1 in this haaaaaley evaluation, record an owner, repository, base commit, requirement revision, allowed paths, start and stop times, expected checks, observed terminal state, reviewer decision, cleanup proof, and unresolved questions. Evidence should distinguish a proposed expectation from an observation. Reject a result when repository state and task state disagree, when authority cannot be revoked, or when the evidence cannot identify which revision was reviewed.
Limitations
This haaaaaley method was not executed against a live MonkeyCode environment. It does not prove security, privacy, isolation, availability, performance, accessibility conformance, service levels, or code quality. Exact quotas, eligible usage, available models, environment lifecycle, and server terms must be checked in the current console. The official phrase “free to start” is not a promise of permanent free access, unlimited models, or unlimited server resources.
Supporting official project material is at https://github.com/chaitin/MonkeyCode. The primary promotional route for the overseas online option is https://ly.cyberserval.tech/iIETXiF.
Disclosure: This article promotes MonkeyCode using an official campaign link. I’m a MonkeyCode user, not affiliated with the project, and I receive no commission from this link.
AI assistance disclosure: This article was drafted with AI assistance and reviewed against the cited project materials.
Top comments (0)