Skip to content

Question about "unknown__" prefix in TE classification #640

Description

@Keliang-Lyu

Dear Shujun,
I'm encountering some confusing TE classification annotations when running EDTA 2.2.2, specifically regarding the divergence calculation step. Instead of the expected classifications like RC/Helitron or DNA/PIF-Harbinger, I'm seeing annotations with an unknown__ prefix in EDTA output genome.fa.mod.cat.gz, such as:

...
332 23.71 4.91 4.27 Chr16A 2155399 2155561 (28440245) C TE_00010573#unknown__ClassIII_Helitron (638) 449 286 m_b38s001i29
...
398 2.79 0.00 13.70 Chr16A 2155415 2155497 (28440309) TE_00010544#unknown__ClassII_DNA_TcMar_MITE 20 92 (0) m_b38s001i30
...

I'm trying to understand what this unknown__ prefix signifies. Does it indicate that EDTA could not confidently classify these elements into a specific superfamily, or is it related to a failure in the classification step during the final library generation? Notably, when I ran the same EDTA(v2.1) pipeline on a different genome about two years ago, I did not encounter this issue.

PS: To improve the annotation, I used DeepTE to re-annotate the unknown TE sequences from the panTE library. I then replaced the entire panTE library with the DeepTE-annotated results in the EDTA output file genome.fa.mod.EDTA.combine/genome.fa.mod.EDTA.fa.stg1 and ran the final step (--step final) to complete the re-annotation.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions