Biarylitides are ribosomally synthesized and post-translationally modified peptides, RiPPs, whose precursor genes span a mere 18 base pairs, the shortest known coding sequences across the tree of life. That extreme miniaturization is also the central obstacle to discovering new members of the family: conventional BLAST searches miss distantly related cytochrome P450 cross-linking enzymes, PSI-BLAST accumulates false positives, and automatic gene annotation pipelines cannot detect an open reading frame this small. The result has been a systematic undercount of biarylitide biosynthetic gene clusters, BGCs, leaving the true breadth of precursor sequence diversity, cross-link chemistry, and co-encoded tailoring enzymes largely unknown. A broader map of the family would expose new P450 substrate tolerance, new building-block chemistry for antibiotic synthesis, and new avenues into organisms, including human-associated bacteria, whose biarylitide capacity has never been recognized.
Researchers in the Crüsemann and Cryle Groups at Goethe University Frankfurt and Monash University, Clayton, Australia, published in JACS Au, adapted AtropoFinder, a random-forest genome-mining pipeline originally built for atropopeptide BGCs, into a dedicated BiarylitideFinder. The classifier was trained on a curated set of 106 confirmed or strongly predicted biarylitide P450 sequences against nearly 10,000 negative examples, and a companion CoreFinder searched the 3 kb genomic neighborhood of each P450 hit for open reading frames of five to eight amino acids carrying aromatic residues at positions 3 and 5, the signature of a biarylitide core peptide. Applied to the full NCBI GenPept database, the workflow identified 277 unique BGCs, including 124 not detected in any prior study, and extended the known phylogenetic range of biarylitide producers to human oral microbiome bacteria in the genus Rothia. To validate the bioinformatic predictions, the team used in vitro cross-linking assays, deuterium-labeled substrates, and multi-dimensional NMR spectroscopy to characterize P450 enzymes associated with three previously uncharacterized core peptide motifs, confirming distinct cross-link regiochemistries for each and showing that substrate tolerance at the central pentapeptide position correlates with the type of biaryl bond formed.
The work demonstrates that machine-learning-assisted genome mining can recover biosynthetic diversity that sequence-similarity methods systematically miss when genes are this short and P450s this divergent. The 277-BGC landscape, with its newly mapped precursor motifs and extended producer phylogeny, provides a structured foundation for exploring biarylitide cross-link chemistry and its potential in biocatalytic production of cross-linked tripeptide antibiotic precursors. Full experimental details, NMR assignments, and the complete BGC table are in the original publication.