alpha0/duet

← back

Paste a transcript. Two characters split it between them and act it out, taking turns, with voices made on your device. Nothing leaves your device, and the video exports straight to a file.

building the valley…
speed

1 · transcript

Speaker labels like Sam:, SPEAKER_01:, [00:12] Sam: or a name on its own line are picked up automatically, and the two most frequent become the cast. Anyone else is folded in. No labels? Every line is a turn, so a YouTube transcript with a timestamp per line works as is. A single block of text is split by sentence. Timestamps and things like [laughs] are dropped. The voices come from Kokoro, a small model that runs in this page; it downloads once and stays on your device.

2 · cast

the voices are made on this device by a small model that downloads once

3 · lines