Hi, thanks for your contribution, I am reproducing your work and found that audio data is missing in the finetune, so I changed the code to run a visual only model. I found that I can't reproduce the results in your paper (CAT (Visual) 59.55 53.67 91.81 51.93). is it possible to release the checkpoint in the paper?
Hi, thanks for your contribution, I am reproducing your work and found that audio data is missing in the finetune, so I changed the code to run a visual only model. I found that I can't reproduce the results in your paper (CAT (Visual) 59.55 53.67 91.81 51.93). is it possible to release the checkpoint in the paper?