Skip to content
Este site está disponível em português
Toolbench

Try: compress pdf, iphone photo, 15% of 240

    Image to Text (OCR)

    The words in a picture, as text you can copy.

    Drop a photo, screenshot or scanned PDF

    The words in it come out as text you can copy, search and edit.

    Guide: From a picture of words to words3 min read

    From a picture of words to words

    A photo of a receipt, a screenshot of a message, a scanned contract: they all show text, but hold only pixels, so nothing can be copied, searched or pasted. Optical character recognition reads the shapes and writes out the letters. Drop one or more images, or a scanned PDF, and the text appears beside it, page by page, ready to copy or save as a .txt file.

    Recognition is done by Tesseract, the open-source engine behind most free OCR, compiled to run in the browser. Choose the language of the text: Portuguese and English are both built in, and the both option helps with a document that mixes the two.

    Getting a good reading

    The engine reads printed text best. Photograph the page straight on, in even light, filling the frame, at the camera's full resolution — a shadow across the page or a steep angle costs more accuracy than anything else. Screenshots read almost perfectly because the letters are already sharp. Handwriting, decorative fonts and very small print are where OCR still struggles, and the page warns when the engine was unsure, so names and numbers get a second look.

    A PDF made from a word processor already contains its text: for that, PDF to Text is exact and instant. This tool is for the PDFs that are pictures of pages.

    The document never leaves your device

    The things people need to read are rarely public: IDs, bank letters, receipts, medical reports. Here nothing is uploaded — the engine and the language model are downloaded once from this site, about 5 MB, and then the reading happens in this tab, even offline. To count the words you got, use the word counter.

    Frequently asked questions

    How accurate is it?

    On a clear photo or scan of printed text, very: Tesseract, the open-source engine used here, reads the large majority of words correctly and flags how sure it is. Accuracy drops with blur, low resolution, strong shadows, pages photographed at an angle, decorative fonts and small print. It is not built for handwriting. Always check names, numbers and amounts before relying on them.

    Why does it download something the first time?

    The recognition engine and the language model — about 5 MB together — are loaded the first time you use it, from this site, not from anyone else. After that they are kept on your device, so later readings start straight away and work offline.

    My PDF already has text. Should I use this?

    No — use PDF to Text, which copies the real text exactly and instantly. OCR is for PDFs that are pictures of pages: scans, photos saved as PDF, faxes. If you are not sure, try selecting a sentence in the PDF; if nothing highlights, it is a scan.

    Does it keep the layout, like columns and tables?

    No. It gives you the text in reading order, line by line, which is what you need to copy a paragraph, search a document or paste figures somewhere else. Tables come out as lines of text, and multi-column pages may need tidying.

    Is my document uploaded?

    No. Recognition runs inside this tab, which matters for the documents people usually need to read: IDs, receipts, contracts, medical letters. Nothing is sent to a server.

    Same drawer