128 points swolpers 2 hours ago 66 comments
thangalin 1 hour ago | parent
https://www.youtube.com/watch?v=WAeHgE94rVo
No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.
Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.
[1]: https://deepmind.google/models/gemma/gemma-4/
[2]: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design
Multicomp 1 hour ago | parent
The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.
Jordan-117 49 minutes ago | parent
loremm 1 hour ago | parent
I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion
talon8635 1 hour ago | parent
simonw 1 hour ago | parent
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
Multicomp 1 hour ago | parent
and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
gruez 59 minutes ago | parent
Probably the latter. Cat's already out of the bag to the extent that you can synthesize with a specific voice in one go and it sounds decent. Even if you need commercial models for better intonation or whatever, you can probably get the commercial models to first generate with a generic voice, then use a local model to transfer that to voice you're cloning. That'll probably get rid of any C2PA watermarks too.
miltonlost 1 hour ago | parent
imjonse 1 hour ago | parent
bakies 1 hour ago | parent
gegtik 1 hour ago | parent
jolan 57 minutes ago | parent
simonw 52 minutes ago | parent
gruez 56 minutes ago | parent
kmoser 55 minutes ago | parent
kingstnap 25 minutes ago | parent
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
So yeah the cat is out of the bag for sure.
perrohunter 1 hour ago | parent
112233 1 hour ago | parent
burkaman 1 hour ago | parent
Multicomp 1 hour ago | parent
Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.
So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.
exhilaration 1 hour ago | parent
Also the Qwen3-TTS demo is cool, you can describe the voice you want: https://huggingface.co/spaces/Qwen/Qwen3-TTS
I came across both on this subreddit, it's very active: https://www.reddit.com/r/TextToSpeech/
I'm personally using this locally: https://github.com/mateogon/pdf-narrator (it's a Python frontend for Kokoro) on my M1 Macbook Air (from 2020, with 8GB RAM) and it's incredible. I make my own audiobooks now - for free!
My favorite voice is am_michael and here's a sample: https://voicerankings.com/voice/kokoro-82M/male/am_michael/s...
talon8635 1 hour ago | parent
sgc 1 hour ago | parent
thevinter 1 hour ago | parent
Price per hour:
- 3.8 Flash TTS, standard: $0.81
- 3.8 Flash TTS, batch: $0.41
- 3.8 Flash‑Lite TTS, standard: $0.54
- 3.8 Flash‑Lite TTS, batch: $0.27
sgc 38 minutes ago | parent
xnx 1 hour ago | parent
mamudo 1 hour ago | parent
laweijfmvo 1 hour ago | parent
andrewstuart 1 hour ago | parent
They all sound like Americans putting in their best fake British accent.
maelito 1 hour ago | parent
Having a voice under 1Mo is crazy, even if it sounds robotic.
drewbitt 1 hour ago | parent
burkaman 1 hour ago | parent
avazhi 1 hour ago | parent
Um, what?
burkaman 1 hour ago | parent
Imagine someone showing you that they've trained their dog to hold a paintbrush and paint. There would be no contradiction between "this is incredible" and "these paintings suck".
nater5000 1 hour ago | parent
Also weird that there are no "neutral gender" voices in the English language. There's also limited "use cases," like the "Gaming" use case is empty?
And there's no pricing listed anywhere.
I don't know, I guess their roll out is a bit sloppy. It's a bit of a shame, though, since the voices which are available all sound like generic Gemini voices to me. Nothing stands out is being particularly interesting or impressive about this.
m3kw9 1 hour ago | parent
seemaze 1 hour ago | parent
Is there a good browser extension that does this with a flexible TTS backend? I know Qwen, Kokoro, and VibeVoice all have decent quality..
kyrra 1 hour ago | parent
janalsncm 35 minutes ago | parent
unglaublich 27 minutes ago | parent
It was especially nice during a bike trip along the Rhine, I listened to a lot of the history of the industrial area and its cities.
xmorse 1 hour ago | parent
https://storage.googleapis.com/gweb-uniblog-publish-prod/ori...
droidjj 1 hour ago | parent
hmokiguess 1 hour ago | parent
xmorse 1 hour ago | parent
konart 38 minutes ago | parent
Onomatopoeia? Sure it is there, and some fillers (or whatever you call those little sounds). But moans?
barrell 28 minutes ago | parent
I would not want that in my product.
fullstackwife 1 hour ago | parent
nitroedge 1 hour ago | parent
$0.50 per hour pricing could last a long time with back and forth conversation use.
cainxinth 55 minutes ago | parent
OutOfHere 55 minutes ago | parent
simonw 53 minutes ago | parent
https://tools.simonwillison.net/gemini-tts-playground#compos...
iAMkenough 50 minutes ago | parent
dadoum 49 minutes ago | parent
LarsDu88 43 minutes ago | parent
dainiusse 24 minutes ago | parent
Thaxll 20 minutes ago | parent