Insight / Research software

A practical anatomy of a reproducible scientific workflow

Reproducibility is not a single notebook. It is the ability to trace inputs, environment, decisions, execution, and outputs well enough to repeat or inspect the work.

A workflow becomes trustworthy when another informed person can understand what entered it, what transformed the material, what assumptions were made, and which output belongs to which run.

01

Preserve the original input

Keep source material separate from derived files. Record provenance, retrieval date, version, identifiers, and any manual exclusions. When the source cannot be redistributed, document how an authorised user can obtain it.

02

Make transformation explicit

Turn repeated manual steps into scripts or clearly documented commands where practical. Parameters should live in configuration or visible code, not only in a person’s memory.

Intermediate outputs can be valuable when they help diagnose where an unexpected result entered the pipeline.

03

Describe the environment

Record software versions, dependencies, operating assumptions, and external services. A lockfile, container, or environment file can help, but it should be paired with a short human-readable setup route.

04

Validate and label outputs

Use checks that reflect the risk: schema validation, row counts, unit tests, known examples, range checks, plots, or manual review. Label outputs with the run, parameters, and input version that produced them.

Reproducibility supports scrutiny; it does not guarantee the scientific question or statistical design was correct.

Working checklist

Before you move forward.

  • Source inputs preserved or referenced
  • Parameters and exclusions visible
  • Dependencies and versions recorded
  • Repeated steps scripted where practical
  • Validation checks matched to risk
  • Outputs traceable to a specific run

Continue reading

Have a project where the method and handover need to remain inspectable?

Discuss the work