This Lora was trained on Japanese language audio clips of Futaba Sakura from Persona 5 Royal. I tried to include a wide range of her emotions in the dataset (shy, nervous, confident, relaxed, happy, excited, loving).
The dataset was captioned in English & transcribed in Japanese, examples below. All of the words spoken in the dataset are Japanese. You can give her English dialogue and she will say it, but I assume it will sound different from her official English voice actor.
Sample training captions & phrases from dataset:
p5futaba, Futaba from Persona 5 is confidently saying "私もみんなと一緒にちょっとずつ変わっていきたい"
p5futaba, Futaba from Persona 5 is embarrassed, her voice shaking, nervously saying "狭い部屋だからこれが不可抗力だから許されるべき"
p5futaba, Futaba from Persona 5 is excitedly saying "おぉ、やっぱり…!"Description
FAQ
Comments (3)
Impressive model! Would you consider making a more detailed post on training an audio lora? Also, have you considered experimenting with multiple characters, so that people could have I2V voiced perfectly?
Thanks! I'm still refining my training approach, less than 25% of the audio loras I've baked have been remotely acceptable. I'll probably write something up once I dial it in a bit more.
I have not. I suppose it's theoretically possible if you caption the dataset properly, but as you can probably tell from my profile I'm generally more interested in exploring single characters thoroughly rather than casting a wide net.
@OneMoreLurker Best of luck improving the method, a write up would be very helpful for anyone who feels inclined to try the multi-character approach. Thanks for your response!