Compare and report

Final check: compare and report

Check fair checkpoint comparisons, configuration records, missing evidence, honest limitations, and conclusions supported by one experiment.

Before the quiz, walk through this short checklist against your notebook and report.

Confirm you held prompts, seeds, and sampling settings fixed when comparing checkpoints. Mixed settings make a good story checkpoint look worse than a lucky DPO sample.

Confirm your run manifest lists tokenizer, blueprint, dataset IDs, and step budget. A missing field there usually means a conclusion you cannot reproduce.

Confirm every strong claim has a loss number, a sample, or a table row beside it. Opinions belong in the limitations section unless evidence backs them.

Confirm you wrote at least one honest limitation. Small data, short training, narrow rubric, or skipped DPO all count.

If you skipped a stage, say so in the report instead of leaving the section blank. Blank sections read like missing homework, not honest science.

Open model-lab.md and search for TBD. Either fill those fields or move the unknown into limitations. A coordinator reading your work should never guess what you measured.

Look at your checkpoint comparison table one more time. Every cell with a score should have a pasted sample underneath it. If a row has only a number, go back and add the raw output.

Check that your best checkpoint is named with a step number and a reason. “Final” is not a reason. “Lowest validation loss at step 11,400” is.

If your report claims SFT helped instruction following, confirm you still have the pre-SFT sample for the same prompt and seed. Side-by-side evidence beats a single after screenshot.

Answer each question before continuing. Read the explanation after every answer, including the ones you get right.

Quick check

Result

You got of right.

Lesson completed