Skip to content

For teachers

Prepare a mission

Everything you need to run the lesson: the prompts for this academic depth, the diagnostic answer key, the misconceptions to watch for, and the truth boundary to hold.

Choose the mission and depth

Did you watch it long enough?

A model can be perfectly correct and still let you reach a wrong conclusion, if you stop watching too early. Your team will run the same spacecraft twice, changing nothing but how long you look, and test whether a conclusion drawn from the first run survives the second.

Pilot lesson

What to say at Grades 9–10

These prompts come from the academic layer, so they change with the depth you selected.

Theory

Define quantities, units, and the comparison you will calculate. Stage intent: differences.

Prediction

Predict a quantitative difference between baseline and candidate.

Running the Twin

Execute baseline/candidate Twin runs and export a table or plot. Runtime remains the frozen Twin; this plan does not execute physics.

Checkpoint

Confirm the calculated difference against the mission criterion.

Analysis

Calculate, model, compare, and test using simulated evidence only. Evidence intent: model/hardware correlation report + discrepancy analysis + V&V conclusion.

Engineering decision

Recommend the candidate with quantitative support and stated uncertainty.

Limitation

Name at least one frozen-model limitation that this experiment cannot answer.

Provenance

Keep simulated, simulated_sensor, estimator_state, derived, reference, and measured distinct. Never label simulated as measured.

Diagnostic answer key

A measurement is repeatable and correct. Does that make it sufficient?

  • Not necessarily - it depends on the claim being made
  • · Yes, correct measurements are always sufficient
  • · Only if it was taken twice

Timing

One session of 55–70 minutes. Adjust freely — the sequence matters more than the clock.

Suggested lesson timing
WhenStageWhat you are doing
0 → 4–5 minMissionSet the role, objective, mission question, and success criterion.
4–5 → 14–18 minPreparationDiagnostic, theory, and a written prediction before any run.
14–18 → 17–22 minReadinessLearners confirm the local formative gate after preparation passes.
17–22 → 31–40 minOperateRun the bounded baseline, then the candidate where comparison is disclosed.
31–40 → 47–60 minEvidenceInspect provenance, select evidence, decide, state a limitation, and complete the formative assessment.
47–60 → 53–68 minCompleteReview the result band, reflect, and finalize local practice at any band.
53–68 → 55–70 minRecognitionExplain the local record and the separate future verified-recognition boundary.

Misconceptions to watch for

Authored lesson design — what a class reliably gets wrong here, and where you can catch it. Not a claim about any learner.

The measurement was repeatable and correct, so it was enough.

Correctness and sufficiency are different questions. Ask what claim was being made, then ask whether this evidence could have detected the claim being false.

Watch: the "sufficiency" diagnostic · Code: evidence_provenance_or_verification_gap

The short run must have been faulty, since the long run disagreed.

Neither the model nor the short run is wrong. What changed is how much of the behaviour was in view — sufficiency, not correctness, is the idea here.

Watch: evidence — The short-run against long-run comparison · Code: evidence_provenance_or_verification_gap

The long run settles it.

Guard the opposite error too: the long run is better evidence, not proof. Require the bound to be written alongside the claim.

Watch: the decision option "A claim bounded by the longer evidence, stated with its limit" · Code: evidence_provenance_or_verification_gap

Reflection and extension

What a good reflection contains

How did the longer run change the claim you were willing to make, and what would count as enough evidence next?

  • Quotes the conclusion written before the long run, and says whether it survived.
  • Treats the short run as insufficient rather than as wrong.
  • States the new claim with its observation-window bound attached.

If they finish early, or go further

  • How long is long enough? (Grades 11–12 and above)

    Propose an observation window you would defend for this claim, and state what would have to be true about the system for your window to be insufficient.

  • Re-examine an earlier mission (University and above)

    Take a conclusion you drew in an earlier mission and ask whether its run was long enough to support it. Rewrite the claim with a bound if it was not.

Facilitation and the truth boundary

While they work

  • Insist the conclusion is written down before the long run. Without that commitment the mission demonstrates nothing.
  • The model is not wrong and the short run is not faulty. Sufficiency, not correctness, is the idea.
  • Guard against the opposite error: the long run is better evidence, not proof.

Did the learner commit to a conclusion first, and can they state what the long run does and does not establish?

Hold this line

  • Both runs are produced by the same model. Neither is measured evidence.
  • A longer software run is still a software run and does not establish flight behaviour.
  • This lesson does not qualify the model against physical hardware.

Home mission

Home mission: too early to tell

Think of something judged too early - a film after ten minutes, a game after one round. Write what the early evidence suggested, what it turned out to be, and what would have counted as enough.

Tell us what did not work

Ten questions, answered locally. Nothing is submitted or tracked — you download the file and send it if you want to.

Informal educator feedback

This local-first form contains the ten approved pilot-review questions. It does not submit, track, or store data remotely. Optional name/contact should be handled outside this form only if a reviewer volunteers it.