Skip to content

Question about data augmentation in Flex-MoE: omitted by design or harmful to specific modules? #10

Description

@shOh-ai

Hello authors! thanks for releasing your awesome works with the clear paper and appendix!
I could not find any mention of data augmentation in the paper, appendix, or code. Could you clarify the rationale?

Was data augmentation intentionally excluded to isolate the contribution of the missing modality bank and the G-router / S-router training strategy, or did you observe degradation when trying augmentation?

Or are there specific modules that are particularly sensitive to augmentation? (Missing modality bank, S-router specialization with top-1 gating, G-router generalization, and so on...)

Thank you!

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions