The technology, without the marketing
SUNOLO is five systems in a row, each of which can be the weak link. This page describes what each one does and where it struggles.
Source separation
Before anything is understood, the mix is split into speech and everything else. This is what allows the original score and effects to survive intact, and it is the step that decides how natural the result sounds.
Speech recognition with speaker tracking
Recognition runs continuously and attributes each line to a speaker, so a two-hander does not collapse into one voice. Speaker identity is held for the duration of playback and then discarded.
Translation on whole thoughts
Words are buffered into complete clauses before translation, because word-by-word output produces grammatical nonsense in languages whose verbs arrive last. This buffering is the main source of the delay.
Voice synthesis
A synthetic voice is selected to match the speaker's register and delivers the translated line with the original emphasis and pacing. Voices are our own; we do not clone the voice of the performer.
Re-synchronisation
The new audio is stretched or compressed slightly and locked back to the picture, so cuts, reactions and lip movement still line up with what you hear.
Where it is weakest
These limits are real and we would rather you heard them from us than found them in the middle of a film.
- Songs are left alone. Translating lyrics mid-performance produces something worse than not translating them.
- Heavy accents, crosstalk and very noisy scenes reduce accuracy, in that order.
- Wordplay and puns usually do not survive. Nothing translates those well, including people.
- Live content is harder than recorded content, because there is no chance to look ahead.
Where the processing happens
SUNOLO BOX does separation and synchronisation on the device itself, and calls out only for recognition and translation. WEB and TV do more of the work in our infrastructure because a browser tab and a television processor cannot spare the budget. In all three cases audio moves over TLS, is processed in memory, and is not written to disk.
What we keep
Your account, your language preferences and your device registrations. Counts of how often features are used, without content. We do not keep the audio, a transcript of it, or a record of the titles you watched.
Content protection and copyright
SUNOLO produces an alternative audio track for personal listening while the original is playing. It does not decrypt, capture, copy, store or redistribute the video or the original audio, and it does not circumvent the protection applied by the services you subscribe to. Where a platform's terms restrict third-party audio processing, those terms apply and we will say so on the compatibility page rather than quietly working around them.
Building something that needs this?
There is an SDK, and there are OEM integrations for platforms and manufacturers.