Describing your video
Starting…
The run stopped
How it works
- Listen. Sarvam Saaras transcribes the audio in 1.5-second chunks and finds every window where nobody is speaking — even through continuous music.
- Look. A vision model watches the frames and writes down what happens, moment by moment.
- Speak. Each quiet window gets one spoken line in your language via Sarvam Bulbul — measured, re-paced and fitted so it never talks over the film.
- Ask. Pause any moment and ask what's happening — in any language. The answer is grounded in the paused frame and the story so far, never a moment ahead, and spoken back to you.