Resources · Glossary
Microbiome glossary
Key terms across sequencing technologies, bioinformatics and microbial community analysis.
Sequencing fundamentalsSample analysis techniquesBioinformatics and data processingMicrobial community analysisFunctional analysisSpecialised terms
Sequencing fundamentals
- Amplicon sequencing
- A targeted approach that amplifies a specific genomic region (such as the 16S rRNA gene for bacteria or the ITS region for fungi) by PCR before sequencing. Cost‑effective profiling of community composition without sequencing every gene.
- Metagenome sequencing (shotgun metagenomics)
- Sequences all DNA in a sample without amplifying specific targets. Gives composition, functional potential and metabolic pathways at higher resolution than amplicon sequencing.
- Whole genome sequencing (WGS)
- Determines the complete DNA sequence of an organism. Used to characterise isolates, identify virulence factors, detect resistance genes and build phylogenies.
- Coverage (sequencing depth)
- The average number of reads aligned to a position in a reference or assembly. Higher coverage raises confidence and reduces the chance of missing low‑abundance organisms; expressed as a multiplier, e.g. 30×.
- Paired‑end reads
- Both ends of a DNA fragment are sequenced, giving two reads per fragment. Improves alignment accuracy, structural‑variant detection and assembly quality.
- Read length
- Number of base pairs in a single read. Long reads (Oxford Nanopore, PacBio) resolve repetitive regions and structural variants; short reads (Illumina) offer higher accuracy and throughput per cost.
Sample analysis techniques
- qPCR (quantitative PCR)
- Real‑time PCR that quantifies a specific DNA or RNA target. Used to validate sequencing findings, quantify taxa or functional genes, and assess total microbial load.
- Primer sets
- Short synthetic DNA sequences that flank and amplify a target region. Primer choice determines which microbial groups are captured; a common pair is 515F/806R for the bacterial 16S V4 region.
- Detection methods
- Culture‑based methods, FISH, MALDI‑TOF mass spectrometry and DNA sequencing, differing in sensitivity, specificity, throughput and cost.
- DNA extraction
- Isolating genomic DNA from a sample. The protocol strongly affects results (gram‑positive bacteria need more vigorous lysis); standardised extraction is essential for reproducible studies.
Bioinformatics and data processing
- Bioinformatics pipeline
- An automated series of steps turning raw reads into interpretable results: quality control, trimming, alignment or assembly, classification or variant calling, statistics.
- Quality control (QC)
- Assessing and filtering sequencing data to remove low‑quality reads, adapters and contaminants. Metrics include Phred scores, GC content, duplication rate and per‑base quality.
- De novo assembly
- Reconstructing genomic sequences from reads without a reference. Assemblers such as SPAdes and MEGAHIT join overlapping reads into contigs and scaffolds.
- Genome annotation
- Identifying and labelling functional elements (genes, regulatory regions, non‑coding RNAs, repeats) in a genome, with tools such as Prokka and RAST.
- OTU / ASV clustering
- Grouping similar sequences to reduce noise. OTUs cluster above a similarity threshold (typically 97 % for 16S); ASVs are exact sequences after error correction, with higher resolution and reproducibility.
Microbial community analysis
- Alpha diversity
- Diversity within one sample: richness, Shannon entropy (accounting for evenness), Faith’s phylogenetic diversity. Higher alpha diversity is generally associated with ecosystem stability.
- Beta diversity
- Dissimilarity in composition between samples, quantified with UniFrac, Bray–Curtis or Jaccard. Used to see how communities cluster by treatment, time point or environment.
- Taxonomic classification
- Assigning reads or assembled sequences to taxonomic groups using reference databases such as SILVA, NCBI or GTDB.
- Metabolic pathway reconstruction
- Predicting functional potential by mapping identified genes to known pathways (KEGG, MetaCyc).
- Relative abundance
- The proportion of each taxon within a sample. Microbiome data is compositional (values sum to 1), which calls for compositional statistics such as log‑ratio transforms.
Functional analysis
- Antimicrobial resistance (AMR) genes
- Genes that let bacteria survive antimicrobials. Detected with databases such as CARD or ResFinder; central to clinical, veterinary, food‑safety and environmental studies.
- Virulence factors
- Molecular components that let pathogens colonise a host, evade immunity and cause disease. Identified from genomes with databases such as VFDB.
- Secondary metabolite biosynthetic gene clusters (BGCs)
- Co‑located genes that synthesise specialised metabolites such as antibiotics, antifungals, bioactive compounds. Identified with tools such as antiSMASH.
- Enzymatic capabilities
- The catalogue of enzyme‑encoding genes in a genome or community, inferred from sequence data and EC‑number databases.
Specialised terms
- Mobile genetic elements (MGEs)
- Plasmids, transposons, insertion sequences, integrons and phages that move between genomes and drive horizontal gene transfer, including the spread of resistance and virulence genes.
- Phenotype prediction
- Inferring observable traits (antibiotic susceptibility, growth conditions, metabolic capabilities, ecological roles) from genome sequence using machine learning or rules.
- Metagenome‑assembled genomes (MAGs)
- Draft genomes reconstructed from metagenomic data by binning contigs on coverage, GC content and tetranucleotide frequency; give access to unculturable organisms.
- Pathogen detection
- Identifying disease‑causing organisms by sequencing. Metagenomics detects all potential pathogens in a sample at once, including unexpected ones.
- Horizontal gene transfer (HGT)
- Transfer of genetic material between organisms outside parent‑to‑offspring inheritance, by transformation, transduction or conjugation. A major driver of microbial evolution and AMR spread.
- Microbiome
- The community of microorganisms (bacteria, archaea, viruses, fungi, protists) in a defined environment, with their genomes and interactions.

Put your microbes to work.
Create a free account and start analysing your data.