hello,
not sure if this is interesting for you,i have been testing parakeet it seems similar to whisper-medium and 3x faster (on my testing run of jfk 'ich bin ein berliner' (hard to find/no idea on some long open speech to use)
---some testing---
i'm testing on 4 cores of i7-9700k/downclocked xmp memory on proxmox lxc:
crispasr(model parakeet-tdt-0.6b-v3-q4_k.gguf):
./crispasr --backend parakeet -m auto --vad --flush-after 1 -osrt --split-on-punct -f ./jfk_1963_0626_berliner.wav >parakeet.srt
~47s transcribe/52s with prep(after first auto-download)
faster-whisper-xxl:
(didn't save cmdline)
on medium model 2m45s
i'm comparing few special-cases:
-
at the end,"ich bin ein berliner",large-v2-distill:"Is bin Iin Birina" :)
large-v3-turbo:"Ich bin ein violiner."
small:"ish bin ein bierlina"
parakeet and whisper-medium got it right
-
"civis Romanus sum"
whisper-medium:"Kiwis Romanus Sum."
parakeet:"Kewis Romanus sum."
basically no model got it right
otherwise english seems pretty good and speedup is substantial on parakeet (or i can use larger non-quantized parakeet)
---end testing---
tldr:
is this interesting enough to include parakeet in subgen?or maybe do some more tests on more material? (i've read some comparisons on reddit and this pushed me to test it,it is nothing spontaneous,just some reviews talked about how awesome speedup this is)
thank you,
k.
edit:i wanted to test more models/backends,but hf has some problems on those models,i was not able to download canary backend quantized model for example
edit2:just tested on 42min mono mp3(old tv series),subtitle edit,on rtx3080 cuda parakeet-tdt-0.6b-v3.gguf ~30s including ~6s download from nas,on vulkan (from memory) twice that long(around 48s) and cpu ~2m30s (default 4 threads on i5-13600k 'almost stock',xmp)
quality(just the feeling):some wound scene it missed few badly spoken words,otherwise translation is spot on,would like to hear from others if i'm crazy to use it,or if they find it crazy good
hello,
not sure if this is interesting for you,i have been testing parakeet it seems similar to whisper-medium and 3x faster (on my testing run of jfk 'ich bin ein berliner' (hard to find/no idea on some long open speech to use)
---some testing---
i'm testing on 4 cores of i7-9700k/downclocked xmp memory on proxmox lxc:
crispasr(model parakeet-tdt-0.6b-v3-q4_k.gguf):
./crispasr --backend parakeet -m auto --vad --flush-after 1 -osrt --split-on-punct -f ./jfk_1963_0626_berliner.wav >parakeet.srt
~47s transcribe/52s with prep(after first auto-download)
faster-whisper-xxl:
(didn't save cmdline)
on medium model 2m45s
i'm comparing few special-cases:
at the end,"ich bin ein berliner",large-v2-distill:"Is bin Iin Birina" :)
large-v3-turbo:"Ich bin ein violiner."
small:"ish bin ein bierlina"
parakeet and whisper-medium got it right
"civis Romanus sum"
whisper-medium:"Kiwis Romanus Sum."
parakeet:"Kewis Romanus sum."
basically no model got it right
otherwise english seems pretty good and speedup is substantial on parakeet (or i can use larger non-quantized parakeet)
---end testing---
tldr:
is this interesting enough to include parakeet in subgen?or maybe do some more tests on more material? (i've read some comparisons on reddit and this pushed me to test it,it is nothing spontaneous,just some reviews talked about how awesome speedup this is)
thank you,
k.
edit:i wanted to test more models/backends,but hf has some problems on those models,i was not able to download canary backend quantized model for example
edit2:just tested on 42min mono mp3(old tv series),subtitle edit,on rtx3080 cuda parakeet-tdt-0.6b-v3.gguf ~30s including ~6s download from nas,on vulkan (from memory) twice that long(around 48s) and cpu ~2m30s (default 4 threads on i5-13600k 'almost stock',xmp)
quality(just the feeling):some wound scene it missed few badly spoken words,otherwise translation is spot on,would like to hear from others if i'm crazy to use it,or if they find it crazy good