Resources · Glossary

Microbiome glossary

Key terms across sequencing technologies, bioinformatics and microbial community analysis.

Sequencing fundamentals

Amplicon sequencing
A targeted approach that amplifies a specific genomic region (such as the 16S rRNA gene for bacteria or the ITS region for fungi) by PCR before sequencing. Cost‑effective profiling of community composition without sequencing every gene.
Metagenome sequencing (shotgun metagenomics)
Sequences all DNA in a sample without amplifying specific targets. Gives composition, functional potential and metabolic pathways at higher resolution than amplicon sequencing.
Whole genome sequencing (WGS)
Determines the complete DNA sequence of an organism. Used to characterise isolates, identify virulence factors, detect resistance genes and build phylogenies.
Coverage (sequencing depth)
The average number of reads aligned to a position in a reference or assembly. Higher coverage raises confidence and reduces the chance of missing low‑abundance organisms; expressed as a multiplier, e.g. 30×.
Paired‑end reads
Both ends of a DNA fragment are sequenced, giving two reads per fragment. Improves alignment accuracy, structural‑variant detection and assembly quality.
Read length
Number of base pairs in a single read. Long reads (Oxford Nanopore, PacBio) resolve repetitive regions and structural variants; short reads (Illumina) offer higher accuracy and throughput per cost.

Sample analysis techniques

qPCR (quantitative PCR)
Real‑time PCR that quantifies a specific DNA or RNA target. Used to validate sequencing findings, quantify taxa or functional genes, and assess total microbial load.
Primer sets
Short synthetic DNA sequences that flank and amplify a target region. Primer choice determines which microbial groups are captured; a common pair is 515F/806R for the bacterial 16S V4 region.
Detection methods
Culture‑based methods, FISH, MALDI‑TOF mass spectrometry and DNA sequencing, differing in sensitivity, specificity, throughput and cost.
DNA extraction
Isolating genomic DNA from a sample. The protocol strongly affects results (gram‑positive bacteria need more vigorous lysis); standardised extraction is essential for reproducible studies.

Bioinformatics and data processing

Bioinformatics pipeline
An automated series of steps turning raw reads into interpretable results: quality control, trimming, alignment or assembly, classification or variant calling, statistics.
Quality control (QC)
Assessing and filtering sequencing data to remove low‑quality reads, adapters and contaminants. Metrics include Phred scores, GC content, duplication rate and per‑base quality.
De novo assembly
Reconstructing genomic sequences from reads without a reference. Assemblers such as SPAdes and MEGAHIT join overlapping reads into contigs and scaffolds.
Genome annotation
Identifying and labelling functional elements (genes, regulatory regions, non‑coding RNAs, repeats) in a genome, with tools such as Prokka and RAST.
OTU / ASV clustering
Grouping similar sequences to reduce noise. OTUs cluster above a similarity threshold (typically 97 % for 16S); ASVs are exact sequences after error correction, with higher resolution and reproducibility.

Microbial community analysis

Alpha diversity
Diversity within one sample: richness, Shannon entropy (accounting for evenness), Faith’s phylogenetic diversity. Higher alpha diversity is generally associated with ecosystem stability.
Beta diversity
Dissimilarity in composition between samples, quantified with UniFrac, Bray–Curtis or Jaccard. Used to see how communities cluster by treatment, time point or environment.
Taxonomic classification
Assigning reads or assembled sequences to taxonomic groups using reference databases such as SILVA, NCBI or GTDB.
Metabolic pathway reconstruction
Predicting functional potential by mapping identified genes to known pathways (KEGG, MetaCyc).
Relative abundance
The proportion of each taxon within a sample. Microbiome data is compositional (values sum to 1), which calls for compositional statistics such as log‑ratio transforms.

Functional analysis

Antimicrobial resistance (AMR) genes
Genes that let bacteria survive antimicrobials. Detected with databases such as CARD or ResFinder; central to clinical, veterinary, food‑safety and environmental studies.
Virulence factors
Molecular components that let pathogens colonise a host, evade immunity and cause disease. Identified from genomes with databases such as VFDB.
Secondary metabolite biosynthetic gene clusters (BGCs)
Co‑located genes that synthesise specialised metabolites such as antibiotics, antifungals, bioactive compounds. Identified with tools such as antiSMASH.
Enzymatic capabilities
The catalogue of enzyme‑encoding genes in a genome or community, inferred from sequence data and EC‑number databases.

Specialised terms

Mobile genetic elements (MGEs)
Plasmids, transposons, insertion sequences, integrons and phages that move between genomes and drive horizontal gene transfer, including the spread of resistance and virulence genes.
Phenotype prediction
Inferring observable traits (antibiotic susceptibility, growth conditions, metabolic capabilities, ecological roles) from genome sequence using machine learning or rules.
Metagenome‑assembled genomes (MAGs)
Draft genomes reconstructed from metagenomic data by binning contigs on coverage, GC content and tetranucleotide frequency; give access to unculturable organisms.
Pathogen detection
Identifying disease‑causing organisms by sequencing. Metagenomics detects all potential pathogens in a sample at once, including unexpected ones.
Horizontal gene transfer (HGT)
Transfer of genetic material between organisms outside parent‑to‑offspring inheritance, by transformation, transduction or conjugation. A major driver of microbial evolution and AMR spread.
Microbiome
The community of microorganisms (bacteria, archaea, viruses, fungi, protists) in a defined environment, with their genomes and interactions.

Put your microbes to work.

Create a free account and start analysing your data.