// Scientific Analysis Instrument · Live Audit

The Instrument

A Mechanist-protocol audit of an AI-assisted product pipeline: 5 extensions, 44 tests, 150 agents, 4 live products, 8 animated architecture flows, and every artifact in between. Each work item is reviewed as behavior → mechanism → intervention → verified result.

Wang M, Fang J, Qiao S, et al. "Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence." arXiv:2608.12036 (cs.AI), 12 Aug 2026. Full text read 2026-08-14.

01

The protocol, applied

STAGE 1

Hypothesis

Name a behavior, its mechanism, or an application. Ground it in evidence, state atomic claims and milestone tests.

workspace: 50 live store searches → 10 ideas scored /25 → 3 empty cells named before any code
STAGE 2

Experiment

Operationalize each claim: inputs, methods, metrics, controls, success criteria. Sanity check before the full run.

workspace: pure engine first, MV3 shell second, success = passing tests
STAGE 3

Verification

Audit validity: traceability, leakage, metric soundness. Test robustness: does it hold when inputs change?

workspace: 44/44 tests vs paper receipts, live URL audits, vision-model QA of artifacts
STAGE 4

Iteration

Route failures to the right stage: scope to hypothesis, design to experiment. Loop until robust or budget ends.

workspace: 5 bugs caught by tests, each fixed at the engine layer, not the docs
↺ LOOP UNTIL RELIABLE AND ROBUST · EXTERNAL MEMORY AT EVERY STAGE
02

The specimen: ten mechanism cards

03

Evaluation panel

Scores follow the paper's dimensions: hypothesis novelty, impact, and testability; experiment reliability. Self-assessed 1-10 per work item.

METHOD NOTE: this is an adaptation. Mechanist audits model internals; this instrument audits an AI-assisted product pipeline. Reliability here means unit tests plus live verification, not statistical power. Scores are self-assessed, not peer-reviewed. Every claim in every card links to a public artifact.
04

Global memory

Settled · verified

    Open · candidates for further study

      06

      Research instruments

      Mechanist Protocol Audit

      ARXIV 2608.12036
      InstrumentThis dashboard: an audit of the AI-assisted product pipeline through the Mechanist 4-stage loop.
      Output10 mechanism cards, 4 evaluation dimensions, global memory, live surfaces.
      Status
      LIVE

      Information Abundance Paradox

      ARXIV 2608.12218
      ClaimLong-context training shifts models from parametric knowledge to context reliance (context addiction).
      ReviewFull 33-page read, 8 evidence sweeps, supporting + counter-evidence audit, 5 original theses.
      BridgeThe scientific backbone for CiteCheck and Loop Detector: contextualization and fabrication share a root cause.
      Output
      17-SLIDE
      05

      Live instrument surfaces