Milestone map
Milestone map
3 milestones
Define genetic question and select open dataset
1 week
Define the specific genetic question you will investigate and identify an appropriate freely available dataset to address it. Accessible alternative: all routes for this outcome use open genomic databases — no physical laboratory access is required. NCBI SRA contains millions of publicly deposited sequencing datasets; Ensembl provides genome annotation and comparative genomics data; NCBI dbSNP and published GWAS summary statistics are freely downloadable.
Proof required
Submit your genetic research question (specific enough to produce an answerable result), the dataset you have selected with its accession number and source URL, a description of the biological context (organism, tissue, condition, or population), and a brief justification of why this dataset is appropriate for your question.
What gets checked
- Genetic question is specific and answerable — not 'understand how genes work' but 'compare the variant frequency of rs334 between African and European populations in the 1000 Genomes data'
- Dataset source is authoritative (NCBI SRA, Ensembl, 1000 Genomes, GTEx, GWAS Catalog, or equivalent) with an accessible accession number
- Biological context is described at sufficient depth to interpret the results — organism, tissue or population, and what biological question the analysis addresses
Common mistakes
- Choosing a dataset without first confirming its quality and completeness — population genetics data from low-coverage sequencing requires different methods from high-coverage WGS; metadata inspection before analysis is essential
- Framing a question that requires controlled experimental data but selecting an observational dataset — association studies cannot establish causation; the question must match what the data can actually show
Resources
Foundationstart here
Depthgo deeper
Masteryfor the dedicated
What a verifier looks for
- Ask the submitter to explain why their dataset is appropriate for their question — does the sequencing depth, coverage, or population structure allow the question to be answered?
- Ask what the potential confounders in their dataset are — population stratification, batch effects, or ascertainment bias are classic genetic analysis issues.
- Ask what quality control steps they will apply before analysis — a geneticist with real bioinformatics experience will ask about QC thresholds for depth, missingness, and heterozygosity.
Execute bioinformatic analysis and produce results
2–4 weeks
Run the bioinformatic analysis on your selected dataset using appropriate free computational tools and produce interpretable results. The analysis pipeline must be documented so that it is reproducible from the documentation alone. Results must include appropriate statistical testing, not only visual patterns.
Proof required
Submit your analysis pipeline documentation (the tools used, the commands run, and the parameters chosen with justification), the output figures (at minimum two: a quality plot and a result visualisation), and a summary table or text stating the statistical result with effect size and confidence interval or p-value as appropriate for your question.
What gets checked
- Pipeline documentation is specific enough to reproduce — every major step names the tool, version, and the parameters used; not 'ran BLAST' but 'blastn -query sequences.fasta -db nt -evalue 1e-5 -outfmt 6'
- Statistical result is reported completely — effect size, test statistic, degrees of freedom, and p-value or confidence interval; visual patterns alone are not statistical evidence
- Quality control output is shown — a plot or table demonstrating that low-quality data have been removed before the main analysis
Common mistakes
- Reporting only visual patterns without statistical tests — a phylogenetic tree that 'looks like' one clade groups together is not a statistical finding without bootstrap support values
- Using default tool parameters without justification — key parameters like e-value thresholds, minimum sequence identity, or population genetics window sizes all have biological implications that must be considered
Resources
Foundationstart here
Depthgo deeper
Masteryfor the dedicated
What a verifier looks for
- Ask the submitter to run a specific analysis step in front of you — confirms the pipeline is real and not reconstructed after the fact.
- Ask what their QC thresholds were and why — the choice of quality cutoffs has a direct effect on results; a bioinformatician who understands the data will justify them.
- Ask what would happen to the key result if one QC parameter was changed — tests understanding of the relationship between data quality and biological conclusions.
Write genetics analysis report and present with Q&A
1–2 weeks to write and schedule review
Complete a genetics analysis report communicating the research question, analytical methods, results, and biological interpretation. Present to a geneticist or bioinformatician with relevant research or clinical genetics experience for a Q&A that challenges both the analytical choices and the biological interpretation of the findings.
Proof required
Submit your complete genetics analysis report (2500–4000 words: introduction with biological context, methods including tool versions and key parameters, results with figures and statistical summaries, discussion interpreting the biological significance and limitations of the analysis) plus a Q&A record showing specific technical and biological challenges from the reviewer and your responses. The reviewer must be named and their genetics or bioinformatics research background stated.
What gets checked
- Methods section includes tool names and version numbers — a critical requirement in bioinformatics, where results can vary between tool versions
- Discussion distinguishes between what the analysis actually shows and the broader biological conclusion — results from a single dataset do not generalise without appropriate caveats
- Q&A record shows at least two technical or biological challenges and substantive responses
Common mistakes
- Overclaiming from a single dataset — population genetics studies with limited sample sizes cannot establish causality; the discussion must scope conclusions to what the data can support
- A reviewer with only computational programming experience but no genetics knowledge — the Q&A must probe biological interpretation, which requires genetics domain expertise
Resources
Depthgo deeper
What a verifier looks for
- Ask the submitter to explain the most important limitation of their dataset for their specific question — not generic limitations but why this dataset specifically may not fully answer their question.
- Ask what an alternative biological interpretation of the same results would be — tests whether the conclusion is the only or most parsimonious explanation of the data.
- Verify the reviewer has genetics research or clinical genetics experience — computational experience without genetics domain knowledge is insufficient for biological interpretation review.
Part of