All outcomes
Skills

Genetic Data Interpretation

8 weeks · 0 milestones

Interpret real genetic data from a named publicly available source: annotate a set of variants using ClinVar and Ensembl Variant Effect Predictor (VEP — free), interpret a GWAS result from a published study by tracing a named variant through the statistical and functional evidence, or analyse sequencing results for a named genetic question using published population data. Document the evidence base for each interpretation claim: variant frequency, functional consequence predictions, clinical significance classifications, and limitations of the evidence. Reviewed by a geneticist or bioinformatician who presents an unseen variant of uncertain significance during the review session and asks you to classify it and justify the classification using ACMG/AMP variant interpretation criteria — requiring real-time reasoning about evidence quality, not recall of known variant classifications. All data sources and annotation tools are freely accessible.

Milestone map

Milestone map

3 milestones

Define genetic question and select open dataset

1 week

Define the specific genetic question you will investigate and identify an appropriate freely available dataset to address it. Accessible alternative: all routes for this outcome use open genomic databases — no physical laboratory access is required. NCBI SRA contains millions of publicly deposited sequencing datasets; Ensembl provides genome annotation and comparative genomics data; NCBI dbSNP and published GWAS summary statistics are freely downloadable.

Proof required

Submit your genetic research question (specific enough to produce an answerable result), the dataset you have selected with its accession number and source URL, a description of the biological context (organism, tissue, condition, or population), and a brief justification of why this dataset is appropriate for your question.

What gets checked

  • Genetic question is specific and answerable — not 'understand how genes work' but 'compare the variant frequency of rs334 between African and European populations in the 1000 Genomes data'
  • Dataset source is authoritative (NCBI SRA, Ensembl, 1000 Genomes, GTEx, GWAS Catalog, or equivalent) with an accessible accession number
  • Biological context is described at sufficient depth to interpret the results — organism, tissue or population, and what biological question the analysis addresses

Common mistakes

  • Choosing a dataset without first confirming its quality and completeness — population genetics data from low-coverage sequencing requires different methods from high-coverage WGS; metadata inspection before analysis is essential
  • Framing a question that requires controlled experimental data but selecting an observational dataset — association studies cannot establish causation; the question must match what the data can actually show

Resources

Foundationstart here

Depthgo deeper

Masteryfor the dedicated

What a verifier looks for

  • Ask the submitter to explain why their dataset is appropriate for their question — does the sequencing depth, coverage, or population structure allow the question to be answered?
  • Ask what the potential confounders in their dataset are — population stratification, batch effects, or ascertainment bias are classic genetic analysis issues.
  • Ask what quality control steps they will apply before analysis — a geneticist with real bioinformatics experience will ask about QC thresholds for depth, missingness, and heterozygosity.

Execute bioinformatic analysis and produce results

2–4 weeks

Run the bioinformatic analysis on your selected dataset using appropriate free computational tools and produce interpretable results. The analysis pipeline must be documented so that it is reproducible from the documentation alone. Results must include appropriate statistical testing, not only visual patterns.

Proof required

Submit your analysis pipeline documentation (the tools used, the commands run, and the parameters chosen with justification), the output figures (at minimum two: a quality plot and a result visualisation), and a summary table or text stating the statistical result with effect size and confidence interval or p-value as appropriate for your question.

What gets checked

  • Pipeline documentation is specific enough to reproduce — every major step names the tool, version, and the parameters used; not 'ran BLAST' but 'blastn -query sequences.fasta -db nt -evalue 1e-5 -outfmt 6'
  • Statistical result is reported completely — effect size, test statistic, degrees of freedom, and p-value or confidence interval; visual patterns alone are not statistical evidence
  • Quality control output is shown — a plot or table demonstrating that low-quality data have been removed before the main analysis

Common mistakes

  • Reporting only visual patterns without statistical tests — a phylogenetic tree that 'looks like' one clade groups together is not a statistical finding without bootstrap support values
  • Using default tool parameters without justification — key parameters like e-value thresholds, minimum sequence identity, or population genetics window sizes all have biological implications that must be considered

Resources

Foundationstart here

Depthgo deeper

Masteryfor the dedicated

What a verifier looks for

  • Ask the submitter to run a specific analysis step in front of you — confirms the pipeline is real and not reconstructed after the fact.
  • Ask what their QC thresholds were and why — the choice of quality cutoffs has a direct effect on results; a bioinformatician who understands the data will justify them.
  • Ask what would happen to the key result if one QC parameter was changed — tests understanding of the relationship between data quality and biological conclusions.

Write genetics analysis report and present with Q&A

1–2 weeks to write and schedule review

Complete a genetics analysis report communicating the research question, analytical methods, results, and biological interpretation. Present to a geneticist or bioinformatician with relevant research or clinical genetics experience for a Q&A that challenges both the analytical choices and the biological interpretation of the findings.

Proof required

Submit your complete genetics analysis report (2500–4000 words: introduction with biological context, methods including tool versions and key parameters, results with figures and statistical summaries, discussion interpreting the biological significance and limitations of the analysis) plus a Q&A record showing specific technical and biological challenges from the reviewer and your responses. The reviewer must be named and their genetics or bioinformatics research background stated.

What gets checked

  • Methods section includes tool names and version numbers — a critical requirement in bioinformatics, where results can vary between tool versions
  • Discussion distinguishes between what the analysis actually shows and the broader biological conclusion — results from a single dataset do not generalise without appropriate caveats
  • Q&A record shows at least two technical or biological challenges and substantive responses

Common mistakes

  • Overclaiming from a single dataset — population genetics studies with limited sample sizes cannot establish causality; the discussion must scope conclusions to what the data can support
  • A reviewer with only computational programming experience but no genetics knowledge — the Q&A must probe biological interpretation, which requires genetics domain expertise

Resources

Depthgo deeper

What a verifier looks for

  • Ask the submitter to explain the most important limitation of their dataset for their specific question — not generic limitations but why this dataset specifically may not fully answer their question.
  • Ask what an alternative biological interpretation of the same results would be — tests whether the conclusion is the only or most parsimonious explanation of the data.
  • Verify the reviewer has genetics research or clinical genetics experience — computational experience without genetics domain knowledge is insufficient for biological interpretation review.

We use analytics to improve Powstik. No ads, ever.