Deni.AI Labs100% in-browser · no uploads
LIVEBackground Remover · ~26 MB Image to Text (OCR) · ~10 MB Speech to Text · ~40 MB Object Detector · ~170 MB AI Image Describer · ~250 MB Text Summarizer · ~300 MB Sentiment Analyzer · ~70 MB Text to Speech · 0 MB
AI Labs › AI Tools › Object Detector

Object Detector

Find common objects in a photo and draw labeled boxes around them, using a detection model that runs locally.

VisionModel DETR ResNet-50Download ~170 MBSpeed 3–10 sPrivacy stays on your device
Drop, paste or click to choose an image
Ready. The model downloads the first time you run it (~170 MB).

Object detection goes a step beyond image recognition: rather than saying what a picture is about, it locates each item it recognises and marks it with a box and a label such as person, car, cup or dog, along with a confidence score. This tool lets you try that on your own images in a few clicks.

All processing happens in your browser. The photo is not uploaded, which makes the tool suitable for pictures of your home, workplace or family that you would rather not share with an online service.

How to use it

  1. Choose or drop an image file onto the tool.
  2. The first time, wait for the model to download; it is about 170 MB and is cached for future visits.
  3. Click detect. After a few seconds, coloured boxes with labels and confidence percentages appear over the image.
  4. Adjust the confidence threshold to hide uncertain detections, then save or screenshot the annotated result.

How it works

The detector is DETR with a ResNet-50 backbone (Xenova/detr-resnet-50), run through Transformers.js. ResNet-50 is a convolutional network that converts the photo into a map of visual features. A transformer then looks at the whole map at once and proposes a fixed set of possible objects, each with a box position and a class. The model was trained on the COCO dataset, so it recognises about 80 everyday categories, and the tool draws only the predictions above your chosen confidence level.

Good for

Limitations

FAQ

What do the percentages mean?

Each number is the model's confidence that the box contains that object. Higher is more certain, but a high score is not a guarantee of correctness.

Can it recognise specific people?

No. It labels a person as a person and has no face recognition or identity features.

Does the image leave my computer?

No. Detection runs in this tab and the annotated result is drawn locally on a canvas. Only the model files are downloaded.

More tools

1

Background Remover

Cut the subject out of a photo and save it as a transparent PNG, processed entirely on your own device.

VisionModel MODNetDownload ~26 MBIn JPG / PNG / WebPOut Transparent PNGSpeed 1–5 s
2

Image to Text (OCR)

Turn screenshots, scanned pages and photos of printed text into editable text without sending the image anywhere.

VisionModel Tesseract (English and more)Download ~10 MBIn Screenshot, scan, photoOut Plain textSpeed 2–10 s
3

AI Image Describer

Get a short, plain-English sentence describing what appears in a photo, generated entirely on your own device.

VisionModel ViT-GPT2 captioningDownload ~250 MBIn JPG / PNG / WebPOut One-sentence captionSpeed 2–6 s
4

Speech to Text

Transcribe English speech from an audio file or your microphone using a Whisper model that runs on your own device.

AudioModel Whisper tiny.enDownload ~40 MBIn MP3 / WAV / M4A / micOut Transcript, timestampsSpeed ~1× realtime
5

Text Summarizer

Condense a long English article or report into a few sentences, generated by a model that runs on your device.

TextModel DistilBART CNN 6-6Download ~300 MBIn English text, 60+ wordsOut Short summarySpeed 10–60 s
6

Sentiment Analyzer

Paste English text and see whether each line reads as positive or negative, with a confidence score for each.

TextModel DistilBERT SST-2Download ~70 MBIn Text, one item per lineOut Positive / negative + scoreSpeed <1 s per line