DEV Community

#tts

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
18 Insights from Mass-Producing Voice Models — From Diffusion TTS Voice Design to Training Corpus Creation and Quality Gate Pitfalls

18 Insights from Mass-Producing Voice Models — From Diffusion TTS Voice Design to Training Corpus Creation and Quality Gate Pitfalls

Comments
5 min read
OmniVoice โมเดล TTS 600+ ภาษา โคลนเสียงจากคลิป 3 วินาที เปิดซอร์สฟรี

OmniVoice โมเดล TTS 600+ ภาษา โคลนเสียงจากคลิป 3 วินาที เปิดซอร์สฟรี

Comments
1 min read
Defects Missed in Transcription — AI Speaks After 0.5-Second Silence

Defects Missed in Transcription — AI Speaks After 0.5-Second Silence

Comments
7 min read
Dropping Candidates Due to Fixable Defects — A Story of Measuring Factory Flaws as Product Personality

Dropping Candidates Due to Fixable Defects — A Story of Measuring Factory Flaws as Product Personality

Comments
10 min read
The '3-Character' Catchphrase That the Quality Gate Allowed Became the Model's Habit

The '3-Character' Catchphrase That the Quality Gate Allowed Became the Model's Habit

Comments
10 min read
Code for preventing hallucinations only failed during hallucinations

Code for preventing hallucinations only failed during hallucinations

Comments
11 min read
When '少々' becomes 'しょも' — Japanese characters removed by whitelist

When '少々' becomes 'しょも' — Japanese characters removed by whitelist

Comments
6 min read
Where Did the AI Learn to Stretch Its Greetings Like 'Kon-nichiwaa'?

Where Did the AI Learn to Stretch Its Greetings Like 'Kon-nichiwaa'?

Comments
6 min read
A single rough clip can ruin the entire style — 5 ways to choose the right ones

A single rough clip can ruin the entire style — 5 ways to choose the right ones

Comments
6 min read
Introducing Paxa Labs

Introducing Paxa Labs

5
Comments 2
3 min read
I Built a Free MIT Audio Engine for Expo Because Background Audio Should Not Require a Commercial License

I Built a Free MIT Audio Engine for Expo Because Background Audio Should Not Require a Commercial License

Comments
4 min read
TTS that Changes 'Recording Location' Each Time It's Generated — Inconsistent Audio Quality Ruins Style

TTS that Changes 'Recording Location' Each Time It's Generated — Inconsistent Audio Quality Ruins Style

Comments
6 min read
Speech Speed Cannot Be Changed After Learning — Parameters Are Accepted but Ignored

Speech Speed Cannot Be Changed After Learning — Parameters Are Accepted but Ignored

Comments
6 min read
The Tighter the Quality Gate, the More the Flat Reads Survive — Selection Bias from Verification

The Tighter the Quality Gate, the More the Flat Reads Survive — Selection Bias from Verification

Comments
8 min read
Sesi modelden geri almak: karakter videolarında telaffuzu deterministik yapmak (Bölüm 3)

Sesi modelden geri almak: karakter videolarında telaffuzu deterministik yapmak (Bölüm 3)

Comments 1
7 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.