OCR: pulling text out of images and scans
OCR accuracy depends almost entirely on the input image. Here is what to fix before you blame the tool.
OCR reads text out of pixels. It is what turns a photographed page into something you can search, copy and edit, and it is the missing step whenever a PDF to Word conversion produces nothing but images.
Accuracy is decided by the input far more than by the software. Straight, evenly lit, in-focus, high-contrast text reads almost perfectly. Skewed, shadowed, low-contrast or handwritten text does not. Handwriting in particular is a different and much harder problem, and general OCR handles it poorly regardless of what any tool claims. The OCR tool is in the ToolHub image section.
Image to Text (OCR): the facts that matter
| Input | JPG, PNG, and scanned pages |
| Best accuracy | Printed text, straight, well lit, high contrast |
| Poor accuracy | Handwriting, low contrast, heavy skew |
| Language setting | Matters. Set it before running |
| Output | Plain text for copying or editing |
How to do it, step by step
- 1Improve the image before you run OCR. This has more effect than any setting.
- 2Crop to the text area, straighten it, and make sure the contrast is strong.
- 3Open the OCR tool in ToolHub and upload.
- 4Set the language. Getting this wrong is a common cause of garbled output, particularly when mixing Hindi and English.
- 5Proofread the result. Always. OCR confuses similar characters and a number read wrong is worse than no number.
Three things worth knowing
- Shoot documents in even natural light. Flash creates glare that destroys OCR accuracy.
- If a PDF gives you no selectable text, OCR is the step you are missing.
- Always proofread numbers, dates and reference codes. Those are where OCR errors do real damage.
Common problems, solved
Where to run it
This tool is part of the ToolHub tool library, a library of over 1,000 browser-based tools covering PDF, image, video, audio, text and developer work. Everything runs in the browser and none of it needs an account.
ToolHub is built by codaiman.com, an AI-first software company in Ahmedabad. Related to this page: CodeAiMan ai-first web, app and software development company based in ahmedabad, working with clients across india and internationally.
Image to Text (OCR), in the browser: no install, no sign-up, no upload queue. Part of the ToolHub library from CodeAiMan.
Open ToolHubRelated tools
PDF to Word
Some PDFs convert perfectly and some come out as garbage. The difference is whether the text is real text or a picture of text.
Read guidePDF to JPG
The whole job comes down to one setting: DPI. Pick it wrong and your text is unreadable or your file is enormous.
Read guideJPG to PDF
Mostly used for phone-camera document scans. Here is how to stop them coming out crooked, huge and unreadable.
Read guide