DRISHTI AUDIO DESCRIPTION · BUILT ON SARVAM VOICE

Hear the film you can't see.

Add a video. DRISHTI finds the silences between dialogue and speaks what's happening on screen — in your language.

  1. Ujumps here. Up to 162 seconds (120 MB max).

  2. Ljumps to the list, Vto the mic — say “Hindi please”, or answer in any language and DRISHTI recognises which one you spoke. On auto, it listens to the video instead. A silent film has no voice to listen to — choose a language for those.

  3. Wjumps here. Give the names and DRISHTI says “Mr Bean reaches for the kettle” instead of “a man reaches for the kettle”. Names only ever come from you — DRISHTI never guesses one, and if it can't tell who's who it keeps the description instead.

  4. Have the script? optional

    Sjumps here. A screenplay, a scene, or an IMDb-style synopsis all work. DRISHTI reads it when you press Describe, and uses it only to recognise what is already on screen.

  5. Dstarts. ?reads these instructions aloud at any time. Mmutes the voice.

How it works

  1. Listen. Sarvam Saaras transcribes the audio in 1.5-second chunks and finds every window where nobody is speaking — even through continuous music.
  2. Look. A vision model watches the frames and writes down what happens, moment by moment.
  3. Speak. Each quiet window gets one spoken line in your language via Sarvam Bulbul — measured, re-paced and fitted so it never talks over the film.
  4. Ask. Pause any moment and ask what's happening — in any language. The answer is grounded in the paused frame and the story so far, never a moment ahead, and spoken back to you.