Kokoro Female Voices is a collection of audio LoRAs trained on Kokoro TTS voices.
Currently, 3 voices are available: Nicole, Heart, Bella
Usage
Recommended strength: 0.5–0.7.
Higher strengths may still work, but they can cut out other sounds in the video and may cause lip-sync issues at high values.
Works with t2v (text-to-video), i2v (image-to-video), and ref2v (reference-to-video).
Each LoRA includes 3 example videos demonstrating these usages.
ComfyUI workflows are embedded in the videos—drag and drop them into ComfyUI to load them.
The example workflows use the default ComfyUI workflows plus the Comfy Kitchen node.
Settings used: 28 steps, res_multistep sampler, simple scheduler.
Turbo / Refinement
Lightly tested with Turbo; it works, but I recommend the ComfyUI-H3-AudioRefine node to improve results.
Use Cases
Consistent voices between scenes
Same voice across generations
A predictable voice for a character you want
Credits / License
Original repo used for training the voices:
https://github.com/nazdridoy/kokoro-tts
https://huggingface.co/hexgrad/Kokoro-82M
apache-2.0 License. Credit to the authors for this great tool.
Description
FAQ
Comments (7)
Yeah this is a good idea
any chance of a texas country 'gal? or deep south milf?
Wow! Nicole's voice really makes my kokoro go doki doki.
isnt voice reference actually cheaper in VRAM than lora?
reference only on ref model and I don't think it is cheaper
What's the use case of the lora instead of just using the 30 second audio sample for the REF model? Genuinely curious
i2v, t2v for obvious reasons
ref2v native is pretty glitchy if you want full preservation - stutters/cuts if your length not perfect, also using this is more easy. Mostly for ease of use.