Episode 12.3: Open Science Across the Research Lifecycle

Last updated on 2026-09-08 | Edit this page

Overview

Questions

  • Where in the research process does “openness” actually happen?
  • Does open science mean doing something extra at the end, or something different throughout?
  • How do open practices at one stage of research make later stages easier?

Objectives

Learners will be able to:

  • Map specific open science practices onto the ten stages of the research process introduced in Episode 1.2.
  • Explain why open practices are easiest to adopt early, and hardest to bolt on after the fact.
  • Describe how an open choice at one stage changes what’s possible at a later stage.

A Wise Scholar Once Said…


“Science is a way of thinking much more than it is a body of knowledge.”
— Carl Sagan

Revisiting Aisha’s Study


In Episode 1.2, we followed Aisha, a final-year public health student and community volunteer, as she investigated whether poor sanitation was contributing to malaria among school-aged children in her town. She moved through the ten stages of the research process: identifying her problem, reviewing what was already known, forming a hypothesis, choosing a design, collecting data from 50 homes, analysing it with her nephew, interpreting the results, drawing conclusions, presenting her findings, and reflecting on the whole process.

Aisha’s study worked. But nothing about it was built to be checked or reused by anyone outside her town. Let’s walk back through those same ten stages and ask: what would it look like if Aisha (or a research team doing something similar at a larger scale) built openness in at every step, instead of only thinking about it once the paper was ready?

Openness at Each Stage


Problem Identification: State the research question clearly and put a timestamp on it (in a lab notebook, a project registry, or a public repository) before any data exists. This matters later: it’s the only way anyone can tell whether a hypothesis was predicted in advance or shaped to fit results that had already come in.

Literature Review: Record which sources were searched, which databases, and which search terms, not just which papers ended up cited. A documented search can be checked and repeated; a vague “we reviewed the literature” cannot.

Objectives and Hypotheses: This is where preregistration happens: publicly and permanently recording the hypothesis and the planned analysis, before collecting data. It’s the single practice with the biggest effect on credibility, because it closes off the option to quietly change the question after seeing the answer.

Research Design: Share the full protocol (sampling plan, instruments, checklists) in enough detail that someone else could run the same study in a different location, the way Aisha’s sanitation checklist could, in principle, be reused in a neighbouring town.

Data Collection: Collect data using standard formats and clear documentation from day one (a “codebook” describing every variable), rather than reorganising messy field notes only once someone asks for the raw data.

Data Analysis: Share the analysis code or syntax alongside the results (the actual steps Aisha’s nephew used in Excel), not just the resulting chart. This is what makes a result reproducible in the strict sense you’ll see defined precisely in Episode 2.2.

Result Interpretation: Distinguish, in writing, between what was predicted in advance (confirmatory) and what was noticed only after looking at the data (exploratory). Both are valuable, but conflating them overstates how strong the evidence really is.

Draw Conclusions: State the limitations plainly, including anything that didn’t go according to plan: a missed household, an unusually hot week, a checklist item that turned out to be ambiguous.

Sharing Findings: Publish somewhere accessible to the people who most need it: in Aisha’s case, that’s the town chief and local health workers, not a journal none of them could read even if they wanted to. Where possible, share the underlying data and materials alongside the write-up, not just the narrative.

Evaluate and Reflect: Document what would be worth doing differently, including openness itself. Aisha reflected that involving others earlier would have improved her data; an open project keeps a written record of exactly that kind of lesson, so the next person doing similar work doesn’t have to rediscover it.

The Pattern: Open Early, Not Just at the End


Notice a theme running through all ten stages above: every open practice is far easier to build in from the start than to reconstruct afterwards. A hypothesis can only be preregistered before you’ve seen your data, never after. A codebook is quick to write while you’re still collecting data, and painfully slow to reconstruct from memory a year later. Openness that’s treated as a final step, tacked on right before submitting a paper, usually ends up thin: a data file with unlabelled columns, a note that says “code available on request” that nobody ever actually requests.

This is also why open science and the research process from Episode 1.2 fit together so naturally. Both are iterative, both connect stage to stage, and a weak link early on (an unregistered hypothesis, an undocumented dataset) limits what’s possible at every stage that follows.

Test Your Knowledge!


Challenge

Challenge 1:

At which stage of the research process does preregistration happen?

A. Data analysis B. Objectives/hypothesis formulation C. Sharing findings D. Evaluate and reflect

B. Preregistration happens when hypotheses and the analysis plan are formed, and it only counts if it’s timestamped before data collection begins.

Challenge

Challenge 2:

A researcher shares their dataset for the first time only after a journal asks for it during peer review, several months after data collection ended. What’s the main risk with leaving documentation this late?

A. The journal will reject the paper automatically. B. The researcher may not accurately remember or reconstruct exactly what each variable and decision meant. C. Late sharing is against copyright law. D. There is no real risk; timing doesn’t matter.

B. Documentation degrades with time and memory. Codebooks and protocols are far easier to write accurately while the work is fresh, which is why open practices work best when built in from the start.

Challenge

Challenge 3:

True or False: Once a study is designed, it’s generally too late to make it meaningfully more open.

False, though it gets harder. Data collection and analysis can still be documented as they happen, and findings can still be shared accessibly. But some practices (like preregistering a hypothesis) genuinely can only happen before data collection starts, which is why “open early” beats “open eventually.”

Key Points
  • Open science isn’t a separate step tacked onto the end of a study; it’s a set of choices available at every stage of the research process.
  • Preregistration, protocol-sharing, documented data collection, shared analysis code, and accessible publishing each map onto a specific stage from Episode 1.2.
  • Open practices are far easier to build in as you go than to reconstruct after the fact, especially anything that depends on a timestamp, like preregistration.
  • A weak link early in the process (an undocumented decision, an unregistered hypothesis) limits how open and how credible everything downstream can be.

To-do: Add infographic. Can use the one in episode 1.2.

Callout

💡 You don’t need to open every single stage to benefit from open science. Even one or two practices (say, preregistering your hypothesis and sharing your dataset) meaningfully raise how much confidence others can place in your work.