The HowTo100M + VidChapters-7M + ViTT model is performing poorly on dense video captioning.
yt-dlp -P $TRANSFORMERS_CACHE -o video.mp4 https://www.youtube.com/watch?v=WJPyrQqLTl4
to download this specific video.
[{'sentence': 'Dress up for the ceremony.', 'timestamp': [350.1055353535354, 395.77147474747477]}, {'sentence': 'Take a photo with me.', 'timestamp': [395.77147474747477, 426.21543434343437]}, {'sentence': 'Take a photo with me.', 'timestamp': [426.21543434343437, 456.659393939394]}, {'sentence': 'Take a photo with me.', 'timestamp': [471.88137373737374, 487.10335353535356]}, {'sentence': 'Take a photo with me.', 'timestamp': [487.10335353535356, 517.5473131313131]}, {'sentence': 'Take a picture with me.', 'timestamp': [547.9912727272728, 578.4352323232324]}, {'sentence': 'Take a picture with me.', 'timestamp': [593.6572121212122, 608.879191919192]}, {'sentence': 'Take a picture with me.', 'timestamp': [639.3231515151516, 654.5451313131314]}, {'sentence': 'Take a picture with me.', 'timestamp': [669.7671111111111, 684.9890909090909]}, {'sentence': 'Take a picture with me.', 'timestamp': [684.9890909090909, 700.2110707070708]}, {'sentence': 'Take a picture with me.', 'timestamp': [730.6550303030302, 745.8770101010102]}, {'sentence': 'Take a picture with me.', 'timestamp': [745.8770101010102, 761.0989898989899]}, {'sentence': 'Take a picture with me.', 'timestamp': [791.5429494949495, 806.7649292929293]}, {'sentence': 'Take a picture with me.', 'timestamp': [806.7649292929293, 821.9869090909092]}, {'sentence': 'Take a picture with me.', 'timestamp': [837.208888888889, 852.4308686868687]}, {'sentence': 'Take a picture with me.', 'timestamp': [852.4308686868687, 867.6528484848486]}, {'sentence': 'Take a picture with me.', 'timestamp': [882.8748282828284, 913.318787878788]}, {'sentence': 'Take a picture with me.', 'timestamp': [913.318787878788, 928.5407676767677]}, {'sentence': 'Take a picture with me.', 'timestamp': [928.5407676767677, 943.7627474747475]}, {'sentence': 'Take a picture with me.', 'timestamp': [958.9847272727274, 974.2067070707071]}, {'sentence': 'Take a picture with me.', 'timestamp': [989.4286868686869, 1019.8726464646466]}, {'sentence': 'Take a picture with me.', 'timestamp': [1035.0946262626262, 1065.538585858586]}, {'sentence': 'Take a picture with me.', 'timestamp': [1080.7605656565656, 1111.2045252525254]}, {'sentence': 'Take a picture with me.', 'timestamp': [1111.2045252525254, 1141.648484848485]}, {'sentence': 'Take a picture with me.', 'timestamp': [1141.648484848485, 1156.8704646464648]}, {'sentence': 'Take a picture with me.', 'timestamp': [1156.8704646464648, 1172.0924444444445]}, {'sentence': 'Take a picture with me.', 'timestamp': [1172.0924444444445, 1202.536404040404]}, {'sentence': 'Take a picture with me.', 'timestamp': [1202.536404040404, 1217.758383838384]}, {'sentence': 'Take a', 'timestamp': [1217.758383838384, 1232.9803636363638]}]
The HowTo100M + VidChapters-7M + ViTT model is performing poorly on dense video captioning.
Reproduction:
Run
to download this specific video.
Follow the steps in the demo using the HowTo100M + VidChapters-7M + ViTT checkpoint.
Output captions: