You photographed the contract and fed the photo to a scan to text tool. What came back was soup. Words fused together, and the second column read before the first. Half the numbers were wrong. The page was perfectly readable. To you. OCR does not read the way you do, and the gap between the two is where almost every failed conversion lives.
Optical character recognition (OCR) is the process of converting an image of text — a photo, a scan, a screenshot — into machine-encoded characters a computer can search, copy, and index. It doesn't read the way a person does. It measures shapes on a grid of pixels, compares them against trained patterns, and outputs its best guess, character by character. Everything below follows from that one fact.
Before recognition starts, the engine binarizes the image: every pixel is forced to pure black or pure white. Your eyes forgive a shadow falling across the page. A threshold does not. On the bright side of the shadow, thin strokes get pushed to white and vanish. On the dark side, paper texture gets pushed to black and turns into phantom characters. The photo that looks fine produces garbage for a mechanical reason, not a mysterious one.
The Engine Sees a Different Page
Tilt is the same story. OCR finds text by slicing the page into horizontal lines, so a tilted page breaks the slicing. A 5-degree skew — an angle nobody notices shooting one-handed — increases word error rate by 10 to 15 percent for a traditional engine; push past 10 degrees and the loss can exceed 30 percent. Tesseract's own documentation is blunter about why: skew "dramatically harms the quality of Tesseract's line segmentation", because past a certain angle, text from two adjacent lines falls into the same horizontal band and the lines merge into each other.
Resolution gets squandered by framing, not by the camera. Tesseract's documentation asks for at least 300 DPI. A 12-megapixel phone photo is commonly 4032 pixels on the long edge — the output resolution behind most recent iPhones and mid-range Android cameras — so if a letter-size page fills that edge, you land around 474 effective DPI, comfortably clear. Let the page occupy half the frame with the same phone in the same light and you're at 237: below the floor, for no reason except how far back you stood.
The arithmetic scales predictably with sensor size:
| Camera resolution | Long-edge pixels | Page fills the frame | Page fills half the frame |
|---|---|---|---|
| 8 MP | ~3264 px | ~384 DPI | ~192 DPI |
| 12 MP | ~4032 px | ~474 DPI | ~237 DPI |
| 48 MP (default binned output) | ~4032 px | ~474 DPI | ~237 DPI |
| 108 MP (dedicated high-res mode) | ~12000 px | ~1412 DPI | ~706 DPI |
Most phones advertising a 48 MP or 108 MP sensor still save ordinary photos at a binned 12 MP by default — the full sensor count only kicks in inside a dedicated "high-res" or Pro capture mode. The takeaway isn't "buy a better phone." Almost any camera from the last several years already clears 300 DPI once the page fills the frame; the pixels aren't missing, they're being wasted on ceiling and desk.
Compression finishes the job quietly. JPEG throws away detail around hard edges to shrink the file, and a character stroke is nothing but a hard edge. As a working threshold, image quality around 0.8 on a 0-to-1 scale doesn't meaningfully hurt OCR accuracy — but a document photo forwarded through a messaging app has usually been compressed once by the camera app and again by the app that sent it, and each pass compounds the last. By the time a scanned image to text tool sees the file, it can be several generations removed from the original capture.
Four Capture Habits, Each With a Number
Four habits fix most of what goes wrong before the engine ever sees the page:
- Fill the frame with the page. That's the 474-versus-237 DPI arithmetic above — the same camera, wasted or not, depending only on how close you get.
- Shoot square-on with the document flat. A 5-degree skew costs double-digit accuracy, and a curved page near a spine or a stapled corner is a built-in skew-and-blur generator.
- Use diffuse light and keep your own shadow off the paper. Binarization needs one clean threshold to find, not two competing ones created by a hard shadow edge.
- Feed the tool the original file from your camera roll, not the copy a chat app or social platform already re-compressed on the way through.
Preprocessing can claw back some of what a bad capture loses, but it has limits worth knowing. In the classic controlled OCR study run at the University of Nevada, Las Vegas, swapping a thresholded black-and-white scan for a greyscale scan of the identical page — letting the engine choose its own cutoff region by region instead of accepting one global threshold — took character accuracy from 97.02% to 98.50% on that page set, roughly halving the error rate. Deskewing and denoising passes move the needle the same direction. But no filter reinvents pixels that were never captured in the first place. Thirty seconds of care with the phone beats any amount of software heroics after.
When the Page Itself Fights You
Some documents cause trouble no amount of careful framing solves on its own.
- Glossy laminate or a plastic sleeve. Overhead light bounces straight back into the lens as a blown-out patch, and OCR reads a flare the same way it reads a coffee stain: as missing information. Angle the camera 10 to 15 degrees off the light source, or turn the flash off and shoot in indirect daylight instead.
- Faded thermal receipts and carbon-copy forms. Low contrast between ink and paper defeats the binarization threshold before anything else goes wrong. A slight exposure boost before the shot helps more than any setting applied after it.
- Multi-column layouts. Column detection is a separate problem from character recognition, and it fails first when a page is even slightly skewed — a tilt that would barely dent a single-column page can interleave two columns into nonsense.
- A page photographed near a book spine. The curve isn't cosmetic. It changes the local pixel density across the page, so text near the gutter is effectively lower-resolution than text at the outer margin. Press the binding as flat as it allows, or photograph facing pages as two separate flat shots instead of one curved spread.
- A screenshot instead of a photo. Screenshotting a PDF viewer at less than 100% zoom, or re-photographing a fax someone already downsized, hits the same DPI floor from a different direction. There's no framing fix once the pixels are already gone — only a rescan or a better source copy.
Copy the Words Out, or Keep the Page
Scan to text covers two different jobs, and it helps to know which one you have.
Sometimes you want the words out of the document: a paragraph to quote, an address to paste. That is plain extraction, and the result is text on your clipboard.
Sometimes you want the document itself, but searchable: the page keeps its exact appearance while an invisible text layer sits behind the image, so Ctrl+F, copy, highlighting, and screen readers all work. That is the right shape for archives. The IRS has accepted digital copies as legally equivalent to paper originals since Revenue Procedure 97-22, provided the storage system captures the document completely and accurately, indexes it for retrieval, and can reproduce a legible copy on demand. A shoebox of receipts can become one searchable file, and the paper can go.
Both outputs fall out of the same recognition pass. The difference is what you keep.
The Upload Is the Risk
The documents people run through scan to text are the sensitive ones: IDs, medical records, signed contracts, bank statements, tax returns. The tools people reach for have a track record, and it's worth knowing what's in it before you pick one.
CamScanner, with over 100 million installs on Google Play, shipped an advertising SDK in 2019 that Kaspersky identified as Trojan-Dropper.AndroidOS.Necro.n — a component that could sign users up for paid subscriptions without consent and execute code downloaded from a remote server. Google pulled the app, and a cleaned build reappeared on the Play Store that September with the ad SDK stripped out. ABBYY, an OCR vendor whose engine sits behind a wide range of scanning apps, left 203,896 scanned documents — contracts, NDAs, memos, some dating back to 2012 — sitting on an unsecured, publicly reachable database in 2018. In February 2025, Kaspersky disclosed SparkCat, the first OCR-based stealer confirmed inside Apple's App Store: bundled into otherwise ordinary-looking apps since at least March 2024, it runs Google's ML Kit OCR across a victim's photo gallery hunting for crypto wallet recovery phrases, and had reached almost 250,000 downloads on Google Play alone by the time it was found. None of this makes scanning apps uniquely reckless. It makes document photos, specifically, worth taking seriously as an attack surface, because the payoff for compromising one is a folder of IDs and bank statements rather than a browsing history.
For anything counting as protected health information, there's a second layer. Under HIPAA, a service that processes it on your behalf is a business associate, and disclosure to one requires "satisfactory assurance... documented through a written contract" — most free consumer OCR tools don't offer one, because they were never built for that job. Under GDPR, a scanned medical letter or ID document is Article 9 special-category data, which sits in a stricter enforcement tier than an ordinary personal record.
A tool that never receives the file removes the whole category, rather than promising to handle it responsibly. OxygenPDF's default OCR engine, Tesseract, runs compiled to WebAssembly inside your browser tab — the page image goes into a canvas in your own tab, not into an HTTP request body. Turn the photo into a PDF, then run OCR on it; the output is a searchable PDF with the invisible text layer written in. Need the words alone? Extract the text and copy or download it. Need an editable document? OCR straight to Word. A handful of opt-in engines are also available for particularly complex layouts, and those do send page images to a server for processing — they're labeled as such rather than blended in quietly, so stick to a client-side engine (Tesseract, or the on-device PaddleOCR and Florence-2 models) when the document is sensitive. Either way, this is checkable, not a policy promise: open your browser's network tab, run a client-side engine, and watch the request list. Nothing goes out.
If you're starting from a document that's already a scanned PDF rather than a phone photo, our guide to making a scanned PDF searchable covers verifying the OCR actually worked — a step nearly everyone skips. And if you've been using Google Drive's built-in OCR or Adobe Acrobat's OCR, those guides cover the specific tradeoffs of each against a browser-based tool.
What to Expect
On clean, high-resolution print, a modern engine reads the overwhelming majority of characters correctly. The benchmark vendors quote for printed text — clean, high-resolution, standard fonts — is a character error rate below 1%, meaning 99% or better character accuracy under good conditions, per one OCR vendor's published figures. Read that claim carefully: even at 99%, a 500-word page still has roughly a dozen wrong characters scattered through it, which is why you proofread names and numbers before anything signs or ships, not after.
Handwriting is a different sport, and "OCR accuracy" stops meaning one number once cursive enters the picture. On handwritten print — block letters, not joined-up script — major engines score in the high 80s to mid-90s; one vendor's comparison put ABBYY FineReader at 95.2% and Adobe Acrobat at 88.6% on the same handwritten-print sample. Switch to genuine cursive and the same engines drop into the high 70s to low 90s, because classical OCR segments a page expecting a visible gap between letters, and cursive doesn't leave one. Grandma's letters will need a human proofreading pass no matter whose tool you use — that's an architectural limit of character-by-character recognition, not a resolution problem a sharper photo fixes.
The shadow, the tilt, the framing, the compression: all of it is yours to control before an engine ever sees the file.
Frequently Asked Questions
What's the difference between "scan to text" and OCR?
"Scan to text" is the everyday name for the task; OCR, optical character recognition, is the technology that performs it. Every scan-to-text tool — whether it's converting a phone photo, a PDF, or a fax — is running an OCR engine underneath. In casual use the terms get swapped freely, but OCR is the mechanism and scan to text is the job it's used for.
Can scan to text tools read handwriting?
Sometimes, with real caveats. Classical engines like Tesseract are built for printed text and struggle badly with handwriting, especially cursive, because they expect a visible gap between letters. Modern vision-based models do noticeably better on clean, legible handwriting, but they can fail differently — producing confident, fluent text that is simply wrong, rather than obvious garbage — at a measurable rate even on models built specifically for document parsing. Handwritten output always earns a closer proofreading pass than typed text does.
Does the photo need to be black-and-white, or is color better?
Color or greyscale, not pure black-and-white. A scanner or phone's "document mode" often outputs a 1-bit black-and-white image, meaning the device already decided, permanently, which pixels are text and which are background. OCR does that job better itself, adapting the threshold region by region instead of applying one global cutoff across the whole page.
Will scan to text work on a photo taken through a plastic sleeve or lamination?
It's harder, but usually workable if you avoid glare. The reflective surface bounces overhead light straight into the lens, which reads as a blown-out patch of missing text. Angle the camera slightly off the light source, turn the flash off, and shoot in even daylight instead of under a single overhead bulb.
Is a searchable PDF the same thing as an editable one?
No. A searchable PDF keeps your original scan exactly as it was and adds an invisible text layer behind it — what you see is still a picture, but Ctrl+F, copy, and highlighting all work against the hidden text underneath. An editable document replaces the scanned image with real, reflowable text and fonts, a different and lossier operation since it's now guessing at layout. OCR to Word produces the editable kind; OCR PDF produces the searchable kind.
Does OxygenPDF's scan-to-text tool work offline?
Once the page and its OCR language model have loaded, yes, for the client-side engines — Tesseract, PaddleOCR, and Florence-2 all run recognition in your browser rather than on a server, so there's nothing to upload and no connection needed after that initial load. The opt-in cloud engines, reserved for unusually complex layouts, do need a live connection, since they send the page image out for processing.
What languages does it support?
Twenty for the built-in Tesseract engine, including Chinese (Simplified and Traditional), Japanese, Korean, Arabic, Russian, Hindi, Hebrew, and Thai, alongside the major European languages. Setting the correct language before running OCR isn't cosmetic — it changes both the character set and the language model the engine uses to disambiguate similar-looking characters.
Scan to text in your browser, with the photo staying on your device the whole way through.
Rohman

