I Made A TTS Voice With 78 Sentences On Mary TTS. That's The Result. What Do You Think About It?
Loading reactions…
Loading reactions…
Script: Bdl: Hello, Daniel. Daniel: Hi neighbor. Bdl: Since you're an unofficial voice made with 78 audios on Mary TTS, which is far less than needed, we'll do an interview with you to show how you sound like, Okay? Daniel: Okay. I'm ready. Bdl: What's your little sister's name? Daniel: Margaret. Bdl: Do you like her? Daniel: I like Margaret. Bdl: What is your favorite toy? Daniel: Tigey. Bdl: Who is your best friend? Daniel: Prince Wednesday. Bdl: What do you say when you feel mad? Daniel: Roar. Bdl: Do you like school? Daniel: I like school. Bdl: Who is your teacher? Daniel: Teacher Harriet. Bdl: What do you eat for lunch? Daniel: Apple. Bdl: Is it yummy? Daniel: Apples are yummy. Bdl: Thank you, Daniel. Bye Bye. Daniel: You're welcome. Ugga Mugga.
For a TTS made with 78 sentences, it’s ok. I think. What do you think?
Loading reactions…
That’s the basics of concatenated synthesis, so it has to chop up the recordings that are in the database. You’ll see that in a lot of today’s systems, even Cepstral voices were made this way.
Loading reactions…
I know. At least the basics. I learned some things while making this with Claude AI. Since Daniel had 78 sentences which were 7-8 minutes total, it's sounding very bad, but still understandable, at least with some words.
Loading reactions…
Yeah. Unfortunately, half of the time I have a hard time understanding systems like that, but the concept is pretty straightforward. It’s just when the systems are chopping it up, that’s what makes it a little harder. But it’s interesting to play around with it though.
Loading reactions…
I think so. And, since Claude’s mobile app’s chats are being in small virtual environments, you can do some things like that with it after uploading files needed for the Linux System it uses. For example, I have no idea about the code things, at least in a way like someone who knows coding. Claude made all the work, except checking the dataset.
Loading reactions…
That’s cool! Although I probably won’t use it myself, I think it’s good to know about these methods. Does it specify the type of audio formatted needs? Sounds like the TTS data needs to be in 16 kHz.
Loading reactions…
Yes, it needs to be in 16khz. I’d like to try piper finetuning for a voice like this instead of Mary TTS, but I don’t have a computer which can handle something like that, the online notebooks are broken, and I don’t have the knowledge to update a notebook’s code to fix it.
Loading reactions…
Yeah. And I think it requires manually tuning the rules for the data set as well if I’m not mistaken. Unfortunately, I don’t really have much knowledge on that neural stuff, but I know it needs a lot of training.
Loading reactions…
I don’t have good knowledge either. Claude made most of the stuff instead of me.
Loading reactions…
Interesting. I’m actually trying to make something from a simple Klatt formant synthesizer using synthesis by rule, but it’s quite complicated. I’d love to make a synthesizer from that someday.
Loading reactions…
It’s hard for someone who doesn’t know things like this. I mean someone like me😀 So, I’m using Claude for most of the things.
Loading reactions…
Exactly. Like for me, I get the concepts, and I have a vision of what it should sound like, but actually making it is harder. I’m definitely not giving up though because I have some pretty interesting ideas in mind.
Loading reactions…
You can back up your work on somewhere and send a copy of it to Claude if you'd like to. It may help, I think. At least you can try working with it since the first back up makes you not lose anything if it messes things up.
Loading reactions…
Yeah, that’s true. I tend to do that quite a lot anyway so yeah. At least I’m having fun playing with it however I can.
Loading reactions…
Loading reactions…