Skip to content

Questions regarding training loss behavior and convergence #7

Description

@li129031

"Hi everyone,
I am currently working on reproducing the training process for this repository and would like to clarify some observations regarding the loss behavior.

  • Training Experience: Has anyone successfully completed the training process from scratch? I would appreciate it if you could share your experience.

  • Optimal Performance: At which stage (epoch/iteration) does the model typically achieve its best results?

  • Loss Plateau: Is it normal for the training loss to remain stagnant for a long period? In my current run, the loss curve seems to have plateaued (horizontal trend) and hasn't shown significant changes for quite a while.

Is this plateauing behavior expected in the early or middle stages of training for this model? Any insights on the typical convergence pattern would be very helpful.
Thanks in advance for your help!"

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions