Theory
State professional terminology, lineage, and model-limit questions. Stage intent: V&V matrix.
For teachers
Everything you need to run the lesson: the prompts for this academic depth, the diagnostic answer key, the misconceptions to watch for, and the truth boundary to hold.
A model can be perfectly correct and still let you reach a wrong conclusion, if you stop watching too early. Your team will run the same spacecraft twice, changing nothing but how long you look, and test whether a conclusion drawn from the first run survives the second.
These prompts come from the academic layer, so they change with the depth you selected.
State professional terminology, lineage, and model-limit questions. Stage intent: V&V matrix.
Predict where the model is expected to be useful and where it is not.
Inspect allowed lineage/hash channels only if the profile discloses them. Runtime remains the frozen Twin; this plan does not execute physics.
Separate simulated, derived, reference, and (if ever present) measured evidence.
Design, integrate, verify, and validate within frozen Core V1 limits. Evidence intent: model/hardware correlation report + discrepancy analysis + V&V conclusion.
Produce a V&V-style conclusion that does not claim flight qualification.
Name at least one frozen-model limitation that this experiment cannot answer.
Keep simulated, simulated_sensor, estimator_state, derived, reference, and measured distinct. Never label simulated as measured.
A measurement is repeatable and correct. Does that make it sufficient?
One session of 55–70 minutes. Adjust freely — the sequence matters more than the clock.
| When | Stage | What you are doing |
|---|---|---|
| 0 → 4–5 min | Mission | Set the role, objective, mission question, and success criterion. |
| 4–5 → 14–18 min | Preparation | Diagnostic, theory, and a written prediction before any run. |
| 14–18 → 17–22 min | Readiness | Learners confirm the local formative gate after preparation passes. |
| 17–22 → 31–40 min | Operate | Run the bounded baseline, then the candidate where comparison is disclosed. |
| 31–40 → 47–60 min | Evidence | Inspect provenance, select evidence, decide, state a limitation, and complete the formative assessment. |
| 47–60 → 53–68 min | Complete | Review the result band, reflect, and finalize local practice at any band. |
| 53–68 → 55–70 min | Recognition | Explain the local record and the separate future verified-recognition boundary. |
Authored lesson design — what a class reliably gets wrong here, and where you can catch it. Not a claim about any learner.
“The measurement was repeatable and correct, so it was enough.”
Correctness and sufficiency are different questions. Ask what claim was being made, then ask whether this evidence could have detected the claim being false.
Watch: the "sufficiency" diagnostic · Code: evidence_provenance_or_verification_gap
“The short run must have been faulty, since the long run disagreed.”
Neither the model nor the short run is wrong. What changed is how much of the behaviour was in view — sufficiency, not correctness, is the idea here.
Watch: evidence — The short-run against long-run comparison · Code: evidence_provenance_or_verification_gap
“The long run settles it.”
Guard the opposite error too: the long run is better evidence, not proof. Require the bound to be written alongside the claim.
Watch: the decision option "A claim bounded by the longer evidence, stated with its limit" · Code: evidence_provenance_or_verification_gap
How did the longer run change the claim you were willing to make, and what would count as enough evidence next?
Propose an observation window you would defend for this claim, and state what would have to be true about the system for your window to be insufficient.
Take a conclusion you drew in an earlier mission and ask whether its run was long enough to support it. Rewrite the claim with a bound if it was not.
Did the learner commit to a conclusion first, and can they state what the long run does and does not establish?
Home mission: too early to tell
Think of something judged too early - a film after ten minutes, a game after one round. Write what the early evidence suggested, what it turned out to be, and what would have counted as enough.
Ten questions, answered locally. Nothing is submitted or tracked — you download the file and send it if you want to.
This local-first form contains the ten approved pilot-review questions. It does not submit, track, or store data remotely. Optional name/contact should be handled outside this form only if a reviewer volunteers it.