Six stages, about half a second

Dialogue goes in one end and comes out the other in your language, while the picture carries on untouched. This is what happens in between.

A single signal passes left to right through six stages: listen, understand, translate, re-voice, sync, enjoy. Original dialogue enters at the left and your language leaves at the right about half a second later.
One signal, six stages, about half a second end to end.

Listen

SUNOLO hears the dialogue.

Speech is separated from music and effects, so the score and the sound design stay where they belong.

Understand

AI works out what was said.

Speech recognition runs per speaker, keeping track of who is talking.

Translate

Meaning moves into your language.

Translation is done on whole thoughts rather than word by word, so idiom survives the trip.

Re-voice

A natural voice brings it back.

A synthetic voice matched to the speaker delivers the line with the original timing and emphasis.

Sync

Speech is aligned to the scene.

The new audio is locked to the picture so lips, cuts and reactions still land.

Enjoy

You just watch.

Nothing about the video changes. Only what you hear does.

What does not happen

As much of the design is about what SUNOLO leaves alone as about what it changes.

The video is never touched

No frame is decoded, copied, re-encoded or stored. SUNOLO works on the audio path only, which is why it does not interfere with the content protection your streaming service applies.

Music and effects stay put

Speech is separated from the rest of the mix before anything is translated. The score, the rain, the footsteps and the room tone are the ones the sound designer chose, at the level they chose.

Nothing is kept

Audio is processed in the moment and discarded. There is no archive of what you watched, and no transcript sitting on a server with your name on it.

Subtitles do not go away

If you prefer reading, keep the subtitles your service provides. SUNOLO adds a soundtrack; it does not take anything off the screen.

About the delay

Translation needs a complete thought before it can be accurate, so there is a short buffer between the original line and the re-voiced one. On Standard mode that is typically under half a second — close enough that lip movement and reactions still read correctly. Low-latency mode shortens it further at some cost to phrasing, which suits live sport and news more than drama.

Both modes are a setting, not a plan tier. You can change it mid-scene.

Questions about the process

Does SUNOLO change the video?
No. The picture, the subtitles and the player all stay exactly as they are. SUNOLO works on the audio only, so what you are watching does not change, only what you hear.
How long is the delay?
Typically under half a second on the Standard setting. SUNOLO BOX also offers a Live mode with a shorter buffer for sport and news, and a Cinema mode that buffers longer for the most natural delivery.
Is my audio sent anywhere?
Audio is processed to produce the translated speech and is not stored afterwards. We do not keep recordings of what you watch. The privacy page explains what is processed and for how long.
Which languages can I hear?
Twelve languages are live now and five more are in beta. The languages page lists every one, and you can vote for the language you want next.

Reading about it
is not the same as hearing it.

Half a second,
and the world opens.