Microbiome glossary.
Definitions for key terms across sequencing technologies, bioinformatics, and microbial community analysis.
Sequencing Fundamentals
- Amplicon Sequencing
- A targeted DNA sequencing approach that amplifies a specific genomic region (such as the 16S rRNA gene for bacteria or ITS region for fungi) using PCR prior to sequencing. It enables cost-effective profiling of microbial community composition without sequencing every gene in the sample.
- Metagenome Sequencing (Shotgun Metagenomics)
- A culture-independent approach that sequences all DNA present in a sample without prior amplification of specific targets. It provides comprehensive information on microbial community composition, functional potential, and metabolic pathways at higher resolution than amplicon sequencing.
- Whole Genome Sequencing (WGS)
- The process of determining the complete DNA sequence of an organism's genome at a single time. In microbiology, WGS is used to fully characterise isolates, identify virulence factors, detect antibiotic resistance genes, and perform phylogenetic analyses.
- Coverage (Sequencing Depth)
- The average number of sequencing reads that align to a given position in a reference genome or assembled sequence. Higher coverage increases confidence in variant calls and reduces the likelihood of missing low-abundance organisms. Typically expressed as a multiplier (e.g., 30×).
- Paired-End Reads
- A sequencing strategy where both ends of a DNA fragment are sequenced, producing two reads per fragment. Paired-end reads improve alignment accuracy, enable detection of structural variants, and enhance genome assembly quality compared to single-end sequencing.
- Read Length
- The number of base pairs sequenced in a single read. Longer reads (e.g., from Oxford Nanopore or PacBio platforms) improve assembly of repetitive regions and structural variant detection, while shorter reads (e.g., Illumina) offer higher accuracy and throughput per cost.
Sample Analysis Techniques
- qPCR (Quantitative PCR)
- A real-time PCR method that quantifies the amount of a specific DNA or RNA target in a sample. In microbiome research, qPCR is used to validate sequencing findings, quantify specific taxa or functional genes (e.g., 16S rRNA gene copies), and assess total microbial load.
- Primer Sets
- Short synthetic DNA sequences designed to flank and amplify a specific target region during PCR. In microbiome research, primer selection critically determines which microbial groups are captured. Common examples include the 515F/806R pair for bacterial 16S V4 region amplification.
- Detection Methods
- Experimental approaches used to identify and characterise microorganisms. These include culture-based methods, fluorescence in situ hybridisation (FISH), mass spectrometry (MALDI-TOF), and DNA sequencing. Each method differs in sensitivity, specificity, throughput, and cost.
- DNA Extraction
- The process of isolating genomic DNA from biological samples. The choice of extraction method significantly affects downstream results: different protocols favour lysis of different cell types (e.g., gram-positive bacteria require more vigorous disruption). Standardised extraction is critical for reproducible microbiome studies.
Bioinformatics & Data Processing
- Bioinformatics Pipeline
- An automated series of computational steps for processing raw sequencing data into interpretable biological results. Typical steps include quality control, adapter trimming, read alignment or assembly, taxonomic classification or variant calling, and statistical analysis. GeneDance pipelines are fully automated and configurable.
- Quality Control (QC)
- The process of assessing and filtering sequencing data to remove low-quality reads, adapter sequences, and contaminants before downstream analysis. Common metrics include Phred quality scores, GC content distribution, duplication rates, and per-base sequence quality.
- De Novo Assembly
- The process of reconstructing genomic sequences from short sequencing reads without a reference genome. Assembly algorithms (e.g., SPAdes, MEGAHIT) join overlapping reads into longer contiguous sequences (contigs) and then into scaffolds.
- Genome Annotation
- The process of identifying and labelling functional elements (genes, regulatory regions, non-coding RNAs, repeat elements) within a genome sequence. Tools like Prokka and RAST are commonly used for bacterial genome annotation. Annotation enables interpretation of genomic content in biological context.
- OTU / ASV Clustering
- Methods for grouping similar sequences to reduce data complexity and noise. Operational Taxonomic Units (OTUs) cluster sequences above a similarity threshold (typically 97% for 16S). Amplicon Sequence Variants (ASVs) represent exact sequences after error correction, offering higher resolution and reproducibility.
Microbial Community Analysis
- Alpha Diversity
- A measure of species diversity within a single sample or environment. Common metrics include species richness (number of distinct taxa), Shannon entropy (accounting for evenness), and Faith's Phylogenetic Diversity (incorporating evolutionary relationships). Higher alpha diversity is generally associated with ecosystem stability.
- Beta Diversity
- A measure of the difference or dissimilarity in microbial community composition between samples. Commonly quantified using UniFrac (phylogenetically-weighted), Bray-Curtis dissimilarity, or Jaccard index. Beta diversity is used to assess how communities cluster by treatment, time point, or environment.
- Taxonomic Classification
- The assignment of sequencing reads or assembled sequences to taxonomic groups (kingdom, phylum, class, order, family, genus, species) using reference databases such as SILVA, NCBI, or GTDB. Accuracy depends on database completeness, reference quality, and classifier algorithms.
- Metabolic Pathway Reconstruction
- The prediction of functional metabolic potential from genomic or metagenomic data by mapping identified genes to known biochemical pathways (e.g., KEGG, MetaCyc). Enables inference of what biochemical reactions a microbial community is capable of performing, even without direct metabolomic measurements.
- Relative Abundance
- The proportional representation of each taxon within a sample, expressed as a percentage or fraction of total reads. Microbiome data is inherently compositional, relative abundance values sum to 1, which requires specialised statistical methods (e.g., compositional data analysis, log-ratio transforms) for robust interpretation.
Functional Analysis
- Antimicrobial Resistance (AMR) Genes
- Genes that encode mechanisms enabling bacteria to survive exposure to antimicrobial agents (antibiotics, biocides). Detection uses curated databases such as CARD or ResFinder. AMR gene profiling is critical in clinical, veterinary, food safety, and environmental microbiome studies.
- Virulence Factors
- Molecular components that enable pathogens to colonise a host, evade immune defences, and cause disease. Virulence factor databases (e.g., VFDB) allow identification of toxin genes, adhesins, invasion factors, and immune evasion mechanisms from genome sequences.
- Secondary Metabolite Biosynthetic Gene Clusters (BGCs)
- Clusters of co-located genes responsible for biosynthesis of specialised metabolites including antibiotics, antifungals, and bioactive compounds. Tools such as antiSMASH identify BGCs in genome sequences, supporting natural product discovery efforts.
- Enzymatic Capabilities
- The catalogue of enzyme-encoding genes present in a microbial genome or community, inferred from sequence data and enzyme commission (EC) number databases. Functional enzyme profiling supports applications in industrial biotechnology, bioremediation, and gut microbiome research.
Specialized Terms
- Mobile Genetic Elements (MGEs)
- DNA sequences capable of moving between genomes or genomic locations, including plasmids, transposons, insertion sequences, integrons, and phages. MGEs play a major role in horizontal gene transfer (HGT) and the spread of antimicrobial resistance and virulence genes between bacteria.
- Phenotype Prediction
- The inference of observable biological traits from genomic sequence data using machine learning models or rule-based systems. Applications include predicting antibiotic susceptibility phenotypes, growth conditions, metabolic capabilities, and ecological roles from WGS data alone.
- Metagenome-Assembled Genomes (MAGs)
- Draft genomes reconstructed from metagenomic sequencing data by binning assembled contigs based on coverage, GC content, and tetranucleotide frequency. MAGs enable genomic characterisation of unculturable microorganisms and greatly expand our understanding of microbial diversity.
- Pathogen Detection
- The identification of disease-causing microorganisms in clinical or environmental samples using sequencing-based approaches. Metagenomics enables unbiased detection of all potential pathogens in a sample simultaneously, including novel or unexpected agents, making it a powerful tool for diagnostics and surveillance.
- Horizontal Gene Transfer (HGT)
- The transfer of genetic material between organisms outside of vertical parent-to-offspring inheritance. HGT is widespread in bacteria and occurs via three main mechanisms: transformation (uptake of free DNA), transduction (phage-mediated), and conjugation (direct cell-to-cell transfer). HGT is a major driver of microbial evolution and AMR spread.
- Microbiome
- The collective community of microorganisms (bacteria, archaea, viruses, fungi, protists) inhabiting a defined environment, together with their genomes and the products of their interactions. The human gut microbiome comprises trillions of cells and has been linked to immunity, metabolism, neurological function, and disease susceptibility.