Deni.AI Labs100% in-browser · no uploads
LIVEBackground Remover · ~26 MB Image to Text (OCR) · ~10 MB Speech to Text · ~40 MB Object Detector · ~170 MB AI Image Describer · ~250 MB Text Summarizer · ~300 MB Sentiment Analyzer · ~70 MB Text to Speech · 0 MB
AI Labs › AI Tools › Speech to Text

Speech to Text

Transcribe English speech from an audio file or your microphone using a Whisper model that runs on your own device.

AudioModel Whisper tiny.enDownload ~40 MBSpeed ~1× realtimePrivacy stays on your device
Ready. The model downloads the first time you run it (~40 MB).

This tool converts spoken English into written text. You can upload a recording such as a voice memo, a lecture clip or an interview, or you can press record and speak directly into your microphone. A few moments later you get a plain transcript you can copy, edit and save.

Unlike most transcription services, the audio is not sent to a remote server. The speech recognition model is downloaded to your browser and does all of its work locally, which means recordings of meetings, personal notes or private conversations stay on your computer.

How to use it

  1. Choose an audio file (MP3, WAV, M4A or similar) or click the record button and allow microphone access when your browser asks.
  2. On first use, wait while the model downloads; it is about 40 MB and is cached afterwards.
  3. Start transcription and keep the tab open while the audio is processed in chunks.
  4. Read through the transcript, fix names or technical terms the model misheard, and copy the text.

How it works

The transcription is performed by Whisper tiny (English), an open speech recognition model released by OpenAI, in the converted version Xenova/whisper-tiny.en. It runs through Transformers.js, a JavaScript library that executes machine learning models in the browser. Your audio is resampled to 16 kHz and turned into a spectrogram, a picture of which frequencies are present over time. The model's encoder reads that picture and its decoder writes out the most likely sequence of words, a short section at a time.

Good for

Limitations

FAQ

Does it work with languages other than English?

No. This version uses the English-only Whisper tiny model, which is tuned for English and will produce poor output for other languages.

Is my recording stored or uploaded?

No. Audio from a file or your microphone is processed in memory inside this tab and is discarded when you close or reload the page.

Why does my microphone not work?

Your browser must be given microphone permission, and most browsers only allow it on secure HTTPS pages. Check the permission icon in the address bar and make sure no other app is holding the microphone.

More tools

1

Text to Speech

Listen to any text read aloud using the voices built into your browser and operating system, with no download required.

AudioModel Device voicesDownload 0 MBIn Any textOut Spoken audio (live)Speed Instant
2

Background Remover

Cut the subject out of a photo and save it as a transparent PNG, processed entirely on your own device.

VisionModel MODNetDownload ~26 MBIn JPG / PNG / WebPOut Transparent PNGSpeed 1–5 s
3

Image to Text (OCR)

Turn screenshots, scanned pages and photos of printed text into editable text without sending the image anywhere.

VisionModel Tesseract (English and more)Download ~10 MBIn Screenshot, scan, photoOut Plain textSpeed 2–10 s
4

Object Detector

Find common objects in a photo and draw labeled boxes around them, using a detection model that runs locally.

VisionModel DETR ResNet-50Download ~170 MBIn JPG / PNG / WebPOut Labeled boxes, countsSpeed 3–10 s