Six stages, about half a second
Dialogue goes in one end and comes out the other in your language, while the picture carries on untouched. This is what happens in between.
Listen
SUNOLO hears the dialogue.
Speech is separated from music and effects, so the score and the sound design stay where they belong.
Understand
AI works out what was said.
Speech recognition runs per speaker, keeping track of who is talking.
Translate
Meaning moves into your language.
Translation is done on whole thoughts rather than word by word, so idiom survives the trip.
Re-voice
A natural voice brings it back.
A synthetic voice matched to the speaker delivers the line with the original timing and emphasis.
Sync
Speech is aligned to the scene.
The new audio is locked to the picture so lips, cuts and reactions still land.
Enjoy
You just watch.
Nothing about the video changes. Only what you hear does.
What does not happen
As much of the design is about what SUNOLO leaves alone as about what it changes.
The video is never touched
No frame is decoded, copied, re-encoded or stored. SUNOLO works on the audio path only, which is why it does not interfere with the content protection your streaming service applies.
Music and effects stay put
Speech is separated from the rest of the mix before anything is translated. The score, the rain, the footsteps and the room tone are the ones the sound designer chose, at the level they chose.
Nothing is kept
Audio is processed in the moment and discarded. There is no archive of what you watched, and no transcript sitting on a server with your name on it.
Subtitles do not go away
If you prefer reading, keep the subtitles your service provides. SUNOLO adds a soundtrack; it does not take anything off the screen.
About the delay
Translation needs a complete thought before it can be accurate, so there is a short buffer between the original line and the re-voiced one. On Standard mode that is typically under half a second — close enough that lip movement and reactions still read correctly. Low-latency mode shortens it further at some cost to phrasing, which suits live sport and news more than drama.
Both modes are a setting, not a plan tier. You can change it mid-scene.
Questions about the process
Does SUNOLO change the video?
How long is the delay?
Is my audio sent anywhere?
Which languages can I hear?
Reading about it
is not the same as hearing it.
Half a second,
and the world opens.