The big problem is the cost compared with the length created. AI text-to-voice is pretty good these days, although you could be excused for thinking otherwise. It can take a while to get the pronunciation of some words correct. Mainly because there are so many words that have the same spelling, but different pronunciation depending on the context. Some of these tools can be fine-tuned, but not all. ElevenLabs seems to do a pretty good job across the board, but it’s not cheap. Getting an audio clip to lip-sync with a face is challenging, and I’ve found that HeyGen does a great job at a reasonable price. Most of the videos I’m making at the moment are short. 1 – 2 minutes typically. If I wanted to make longer videos, they would be very expensive. That would mean I’d change my tactics for constructing them. I’d still do the audio in ElevenLabs, but I’d construct the video from still images in Camtasia or CapCut. Lip-syncing wouldn’t be an issue, and the timing of image changes isn’t as critical. Of course, the least expensive way to make longer-form videos is to do the voice-over yourself. Use a slideshow tool to lay out the images and the visual text, then read the author notes as you record each slide individually. If recording your own voice bothers you, listen back to your first take. Almost everyone hates the sound of their own voice at first. After the first few, it won’t bother you either. Regards, |
