Open-Source Computational Biology Hub

Interactive Tutorials & Research Pipelines

Access free, reproducible computational biology walkthroughs, interactive Google Colab notebooks, and video masterclasses covering transcriptomics, single-cell analysis, AlphaFold, and AI drug discovery.

3 Core

Interactive Guides

8+

Colab Pipelines

9+

Video Playlists

100%

Free & Open-Source

Step-by-Step Walkthroughs

Featured Interactive Guides

Read comprehensive, in-depth documentation and code examples built directly into our learning platform.

Ready-to-Run Code

Production-Grade Computational Pipelines

Verified bioinformatics and cheminformatics workflows ready for your research projects.

RNA-SeqOpen-Source Asset

End-to-End Bulk RNA-seq Quantification Pipeline using Salmon

A complete pipeline for pseudo-alignment and quantification of bulk RNA-seq data using Salmon — from raw FASTQ reads to transcript-level abundance estimates ready for downstream DESeq2 analysis.

Key Tools
PythonSalmontximetaDESeq2
Data Specifications

In: Raw FASTQ files, reference transcriptome
Out: Transcript/gene-level count matrix, quantification summary

RNA-SeqOpen-Source Asset

nf-core/rnaseqmeta: Nextflow Pipeline for RNA-seq Meta-Analysis

A Nextflow pipeline for reproducible meta-analysis of multiple RNA-seq cohorts — automating sample retrieval, quality control, batch effect correction, and cross-study differential expression.

Key Tools
Nextflownf-coreSalmonDESeq2
Data Specifications

In: Multiple RNA-seq datasets (FASTQ or SRA accessions), sample sheets
Out: Integrated count matrix, cross-study DE results, batch-corrected expression

Single-CellOpen-Source Asset

Fast Preprocessing of scRNA-seq with kallisto | bustools | kb-python

End-to-end pipeline for scRNA-seq preprocessing using kallisto, bustools, and kb-python — transforming raw sequencing data into filtered count matrices for Scanpy and Seurat.

Key Tools
Pythonkallistobustoolskb-pythonScanpy
Data Specifications

In: Raw FASTQ files (10x Chromium)
Out: Filtered cell × gene count matrix, QC metrics

Single-CellOpen-Source Asset

Practical Guide for Single-Cell Data Analysis with scverse Ecosystem

Comprehensive walkthrough of the scverse workflow — covering quality control filtering, highly variable genes, dimensionality reduction, Leiden clustering, and marker gene annotation.

Key Tools
PythonScanpyscVI-toolsAnnData
Data Specifications

In: Count matrix (AnnData .h5ad or 10x format)
Out: Annotated cell clusters, UMAP embeddings, cell type labels

Protein ModelingOpen-Source Asset

Predicting Protein Structures with ColabFold & AlphaFold2

Predict 3D protein structures from amino acid sequences using ColabFold's accelerated AlphaFold2 pipeline with MMseqs2 MSA generation and interactive in-browser 3D structure visualization.

Key Tools
PythonColabFoldAlphaFold2MMseqs2py3Dmol
Data Specifications

In: Amino acid sequence (FASTA format)
Out: Predicted 3D structures (PDB), pLDDT confidence scores, PAE plots

Protein ModelingOpen-Source Asset

Boltz2-Notebook: Diffusion-Based Protein-Ligand Structure Prediction

Predict protein-ligand complex conformations and binding affinities using the state-of-the-art Boltz2 diffusion model — enabling rapid in silico docking and interaction mapping without heavy MD setups.

Key Tools
PythonBoltz2RDKitpy3Dmol
Data Specifications

In: Protein sequence and ligand SMILES/SDF
Out: Predicted complex structures (PDB), binding affinity scores

Drug DiscoveryOpen-Source Asset

AI in Drug Discovery: Molecular Property Prediction & Virtual Screening

End-to-end pipeline for AI-driven small molecule screening — covering Morgan fingerprint featurization, toxicity prediction, ADMET property modeling, and virtual screening of compound libraries.

Key Tools
PythonRDKitDeepChemscikit-learn
Data Specifications

In: Compound libraries (SMILES), molecular descriptors
Out: Toxicity predictions, ADMET profiles, ranked hit compounds

Drug DiscoveryOpen-Source Asset

In Silico Toxicology & Safety Modeling with Machine Learning

Build machine learning models to predict compound toxicity from 2D molecular structures — covering molecular fingerprint generation, endpoint classification, and structure-activity relationships (SAR).

Key Tools
PythonRDKitscikit-learnMordred
Data Specifications

In: Chemical compounds (SMILES), toxicity endpoint labels
Out: Toxicity classification models, SAR insights, safety predictions

Video Masterclasses

Curated YouTube Video Courses

Structured video series covering programming, computational genomics, and academic publication.

StatisticsR Programming

R for Research & Bioinformatics

Video lectures covering R programming fundamentals, tidyverse data wrangling, statistical testing, and publication-ready visualization.

Watch Course Series
Machine LearningGenomics

Machine Learning for Bioinformatics

Hands-on video tutorials applying machine learning techniques (Random Forest, XGBoost, Neural Nets) to biological and genomic datasets.

Watch Course Series
TranscriptomicsBioconductor

RNA-Seq Analysis with R & Bioconductor

Step-by-step video tutorials on bulk RNA-seq analysis using R, DESeq2, and Bioconductor, from raw counts to biological pathway interpretation.

Watch Course Series
AI TherapeuticsCADD

AI for Drug Discovery & Cheminformatics

Video series introducing AI-powered approaches to drug discovery, including molecular dynamics, toxicology modeling, and virtual screening.

Watch Course Series
Cancer GenomicsTCGA

Cancer Bioinformatics & Multi-Omics

Video tutorials on cancer genomics analysis, covering TCGA data mining, mutation profiles, Kaplan-Meier survival modeling, and immune infiltration.

Watch Course Series
Scientific WritingPublishing

Academic Writing & Manuscript Preparation

Video lectures on scientific writing, publication ethics, structured methodology descriptions, and peer-review preparation.

Watch Course Series
PythonHealth Analytics

Python for Health Data Analytics

Video tutorials on using Python, pandas, and seaborn for clinical data wrangling, epidemiologic trends, and biostatistical models.

Watch Course Series
NextflowWorkflows

Bioinformatics Workflow Automation Bootcamp

Intensive video series on building automated, reproducible bioinformatics workflows using Linux, Conda, Nextflow, and Docker.

Watch Course Series
Single-CellSeurat

Single-Cell Analysis with R & Seurat

Comprehensive step-by-step video guide on scRNA-seq analysis using R and Seurat, from QC filtering to t-SNE/UMAP cluster annotation.

Watch Course Series

Accelerate Your Research with 1-on-1 Mentorship

Take your skills further with direct weekly guidance, live debugging, real-world multi-omics datasets, and personalized career roadmaps.