PlasmidFinder currently identifies plasmids using a database of 488 replicon sequences that has not been expanded since 2014, even though the amount of plasmid sequencing data available has grown enormously since then. Here I describe my work extending this database with new sequences mined from GenBank.
Starting from around 381,000 GenBank plasmid records, I identified replication genes by matching their names and functional descriptions against a curated set of terms. This was not straightforward, since gene names in public databases are not always reliable; the same name can sometimes label two completely unrelated proteins in different organisms, so false matches had to be filtered out along the way. After that, I removed exact duplicates, mostly the same replicon found repeatedly across different bacterial strains, along with other redundant sequences, leaving a final database of 24,254 replicon sequences.
The sequences were also organized into a nested hierarchy of similarity, at four increasingly strict levels, working much like a taxonomic lineage from broad groups down to closely related, specific types. This structure will let the tool report not just whether a plasmid replicon is found, but how closely it relates to already known types, and it will be integrated into the PlasmidFinder tool itself in the next phase of the project.
Alessandro Caula’s presentation