91 6

Hexgrad PRO

hexgrad

https://hf.co/hexgrad/Kokoro-82M

hexgrad

AI & ML interests

Favorite water is distilled

Recent Activity

updated a Space about 11 hours ago

hexgrad/Misaki-G2P

updated a Space about 11 hours ago

hexgrad/Kokoro-TTS

new activity about 11 hours ago

hexgrad/Kokoro-82M:Can I replace the Inference Widget backend on this model page with a Gradio Spaces API?

View all activity

Organizations

None yet

Posts 14

Post

1344

I wrote an article about G2P: https://hf.co/blog/hexgrad/g2p

G2P is an underrated piece of small TTS models, like offensive linemen who do a bunch of work and get no credit.

Instead of relying on explicit G2P, larger speech models implicitly learn this task by eating many thousands of hours of audio data. They often use a 500M+ parameter LLM at the front to predict latent audio tokens over a learned codebook, then decode these tokens into audio.

Kokoro instead relies on G2P preprocessing, is 82M parameters, and thus needs less audio to learn. Because of this, we can cherrypick high fidelity audio for training data, and deliver solid speech for those voices. In turn, this excellent audio quality & lack of background noise helps explain why Kokoro is very competitive in single-voice TTS Arenas.

View all Posts