Skip to content

Not all reference genome gene IDs appear in the clusters_99_info.tsv #159

Description

@michaelwasney

Hi Chunyu!

I built my own custom database with MIDASv3 using an external reference collection (MGBC), published in Cell Host & Microbe. The MGBC study had already clustered genomes together into species clusters using a 95% ANI threshold, and selected a representative genome for each species cluster (the exact method is detailed in the section of the methods of the Cell Host & Microbe paper entitle "Taxonomic classification of genomes and species clustering"). Before constructing the database, I built the 'clean_imports' folder based on their clustering and representative genome selection results.

My problem: When building the database, some representative genome gene IDs are not included in the clusters_99_info.tsv file produced during the "Build pangenome" step of the database construction pipeline (lines 37-42 of the test_db.sh script). In my downstream analyses, I need to be able to map representative genome gene IDs to other pangenome gene IDs.

My Question: Why is this happening? I never noticed this behavior when using the MIDASv1 database building feature. I was wondering if this behavior is a feature of MIDASv3 I should be aware of, or potentially something wrong with the reference collection I'm using to build the database.

Thank you in advance for you help!

Best,
Michael Wasney

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions