How to OCR an Image (and What It Gets Wrong)
Read the printed text out of a photo, screenshot or scan in your browser. Learn what OCR actually does, why a big phone photo loses characters, and what to fix before you trust a result.
What OCR is, in one paragraph
Optical character recognition does not read. It guesses. The engine looks at shapes of dark pixels on a light background, compares each one against a model of how thousands of printed characters are shaped, and returns whichever characters it thinks it saw, along with a score for how sure it is about its own guesses. There is no understanding of your document, no dictionary of your company's names, and no second pass that notices a total is wrong. That is why every OCR tool on the market, including this one, needs you to read the result over.
Image OCR runs Tesseract as a WebAssembly build inside your browser tab. The image is decoded and recognized on your own device and is never uploaded, and the engine, its WebAssembly core and the language model are all served from this site's own origin rather than a third-party CDN. The model for the language you pick is downloaded the first time you use that language and cached by your browser afterwards.
What the tool does before it reads a single character
It checks three things, and each of them saves you a confusing failure later. The format is confirmed from the file's own header bytes rather than its name, so a .png that is really a text file or a truncated download is refused immediately with that reason instead of dying somewhere inside the recognizer. The file size is compared against a 25 MB cap. And the dimensions are read from the header and checked against a 16-megapixel budget and an 8192 px limit on the long side.
All three caps are refusals, not adjustments. A 20-megapixel photo is rejected with its real numbers, not quietly shrunk to fit — which means the file that gets read is the file you chose, at the size it actually is. If you hit a cap, the fix is on your side: crop to the text you want, or downscale a very large photo, and re-open it.
Why a 12-megapixel photo reads worse than a cropped scan
The single biggest cause of bad OCR is not a bad engine, it is a big input. Tesseract works on a page image scaled to roughly 300 DPI. Hand it a 4000-by-3000 photo of a page and it downsamples to its working size before it looks at anything — which throws away exactly the fine detail that distinguishes a comma from a full stop, a 1 from a 7, or an O from a 0. The characters that survive the downsample are the ones that were large to begin with; small print in a footer is the first thing to go.
So the ranking of inputs, best first, is: a scan or export at about 300 DPI, cropped to the text; a flatbed or phone-scanner scan at that resolution; a screenshot taken at native resolution; a phone photo of a printed page, which is fine if it is square-on and evenly lit and poor if it is not. The engine also assumes the text is horizontal, dark on light and not rotated — deskew a scan and it gets noticeably better.
The five things that reliably break it
Skew and perspective: a page photographed at an angle is the most common failure, because the engine models a flat, straight baseline. Shadows and uneven lighting produce gradients across the page, and a gradient is a character as far as the recognizer is concerned. Low contrast — grey text, a faded photocopy, a screenshot with light-grey helper text. Stylised and decorative fonts, where the recognizer has never seen the shapes in training. And handwriting, which is a different problem entirely: the models shipped here are trained on printed text, and cursive or joined-up writing is not what they are looking for.
One more that is not the image's fault: EXIF rotation is not applied. A photo taken in portrait and stored with a rotation flag is handed to the recognizer sideways, and it will return sideways text. Rotate the image before you open it, or accept that the output needs turning.
Reading the result, including the confidence number
Every run reports the engine's own confidence score alongside the text, with a plain description of what the band means. Read that number for what it is: the recognizer's opinion of how plausible its guesses were, averaged over the page. It is not a percentage of characters that are correct, and there is no OCR tool anywhere that can give you that, because the engine has no idea which of its guesses were wrong. A high score on a page with unusual words is not a guarantee, and a middling score on a very clean page often still reads perfectly.
The result panel gives you character, word and line counts and the elapsed time, and the text is cleaned up before you see it — Windows and form-feed line breaks become plain newlines, trailing spaces are trimmed and runs of blank lines collapse. Copy puts that text on your clipboard; Download writes the same bytes to a file named <image>-ocr-<language>.txt. If the result comes back empty, that is reported as a result rather than a silent failure, with the fixes that usually help: a larger, straighter, better-lit scan of the same page.
What this is not
It is not a document scanner, a form filler or a data extractor. It does not find invoice totals, read a table into columns, sort the lines into reading order for a multi-column page, or tell you which of two candidate readings is the right one — Tesseract is LSTM-based here, so it produces one guess per line, not a confidence-ranked list. It does not do handwriting. It does not OCR inside a PDF (that is a separate job, because a PDF needs its pages rasterized first), and it does not accept HEIC, AVIF, TIFF or SVG, which are not in the five supported formats.
All of the work happens in your browser. The image is read locally, recognized locally and saved locally, which matters most for exactly the documents people most want scanned: contracts, payslips, medical letters, ID cards and anything under review before it goes out. If you need higher fidelity than a single offline pass can give you — archival scanning, batch processing, layout-aware extraction — that is a different tool, and this one will not pretend otherwise.
Try it free — Image OCR & Text Extraction
Read printed text out of a PNG, JPEG, GIF, WebP or BMP image in 12 languages, with Tesseract running as WebAssembly on your device. OCR is a guess, so the engine's own confidence score is shown.
Open Image OCR & Text Extraction