Hi Chunyu!
I built my own custom database with MIDASv3 using an external reference collection (MGBC), published in Cell Host & Microbe. The MGBC study had already clustered genomes together into species clusters using a 95% ANI threshold, and selected a representative genome for each species cluster (the exact method is detailed in the section of the methods of the Cell Host & Microbe paper entitle "Taxonomic classification of genomes and species clustering"). Before constructing the database, I built the 'clean_imports' folder based on their clustering and representative genome selection results.
My problem: When building the database, some representative genome gene IDs are not included in the clusters_99_info.tsv file produced during the "Build pangenome" step of the database construction pipeline (lines 37-42 of the test_db.sh script). In my downstream analyses, I need to be able to map representative genome gene IDs to other pangenome gene IDs.
My Question: Why is this happening? I never noticed this behavior when using the MIDASv1 database building feature. I was wondering if this behavior is a feature of MIDASv3 I should be aware of, or potentially something wrong with the reference collection I'm using to build the database.
Thank you in advance for you help!
Best,
Michael Wasney
Hi Chunyu!
I built my own custom database with MIDASv3 using an external reference collection (MGBC), published in Cell Host & Microbe. The MGBC study had already clustered genomes together into species clusters using a 95% ANI threshold, and selected a representative genome for each species cluster (the exact method is detailed in the section of the methods of the Cell Host & Microbe paper entitle "Taxonomic classification of genomes and species clustering"). Before constructing the database, I built the 'clean_imports' folder based on their clustering and representative genome selection results.
My problem: When building the database, some representative genome gene IDs are not included in the clusters_99_info.tsv file produced during the "Build pangenome" step of the database construction pipeline (lines 37-42 of the test_db.sh script). In my downstream analyses, I need to be able to map representative genome gene IDs to other pangenome gene IDs.
My Question: Why is this happening? I never noticed this behavior when using the MIDASv1 database building feature. I was wondering if this behavior is a feature of MIDASv3 I should be aware of, or potentially something wrong with the reference collection I'm using to build the database.
Thank you in advance for you help!
Best,
Michael Wasney