[{"data":1,"prerenderedAt":122},["ShallowReactive",2],{"glossary-en-ocr":3},{"id":4,"title":5,"body":6,"category":112,"description":113,"draft":114,"extension":115,"meta":116,"navigation":117,"path":118,"seo":119,"stem":120,"__hash__":121},"glossary_en\u002Fglossary\u002Focr.md","OCR (optical character recognition)",{"type":7,"value":8,"toc":103},"minimark",[9,13,18,30,33,37,40,57,66,76,79,83,86,91],[10,11,12],"p",{},"OCR (optical character recognition) is the technology that converts the text visible in an image — a scanned page, a photographed receipt, an image-only PDF — into machine-readable text that software can search, copy, and analyze. Without OCR, a scan is just pixels: a computer can display it but has no idea what it says.",[14,15,17],"h2",{"id":16},"why-scans-are-the-problem-files","Why scans are the problem files",[10,19,20,21,25,26,29],{},"Documents arrive in two forms. Native digital files (a Google Doc, a PDF exported from an invoicing tool) carry their text internally — search finds them, software reads them directly. Scans and photos don't. A passport photographed with a phone, a paper invoice run through a scanner, a receipt snapped at a restaurant: all text-as-image. These are precisely the files that pile up with names like ",[22,23,24],"code",{},"scan_0042.pdf"," or ",[22,27,28],{},"IMG_4523.HEIC",", because nothing about them tells you — or your tools — what's inside.",[10,31,32],{},"OCR closes that gap. Run on the receipt photo, it produces the merchant name, the date, the line items, the total — as text. Modern OCR handles skewed photos, multiple languages, and mediocre scan quality far better than the technology's reputation suggests, though it remains probabilistic: a crumpled thermal receipt or a low-light photo can still come back partly garbled.",[14,34,36],{"id":35},"ocr-in-a-sorting-pipeline","OCR in a sorting pipeline",[10,38,39],{},"For file organization, OCR is the first of two steps:",[41,42,43,51],"ol",{},[44,45,46,50],"li",{},[47,48,49],"strong",{},"OCR"," extracts the raw text from the scan or photo.",[44,52,53,56],{},[47,54,55],{},"Classification"," interprets it: this is an invoice, from supplier Acme, dated 2026-06-12, for 840 €.",[10,58,59,60,65],{},"The second step is what an ",[61,62,64],"a",{"href":63},"\u002Fglossary\u002Fai-file-organizer","AI file organizer"," adds on top. Together they turn an unreadable filename into a placed, renamed file:",[67,68,73],"pre",{"className":69,"code":71,"language":72},[70],"language-text","scan_0042.pdf  →  Invoices \u002F 2026 \u002F 06 - June \u002F acme-2026-06-12.pdf\n","text",[22,74,71],{"__ignoreMap":75},"",[10,77,78],{},"This pipeline is built into Sorters for Google Drive files: when you run a sort, each picked file — PDF, JPG, PNG, HEIC and more — goes through OCR and AI classification so the rename-and-file rules can use what the document actually says, not what the scanner named it.",[14,80,82],{"id":81},"do-you-need-ocr","Do you need OCR?",[10,84,85],{},"Only if image-form documents are part of your mess. A Drive full of native Google Docs and exported PDFs is searchable as-is — Drive even applies some OCR to scans in its own search index, which is why searching a supplier name sometimes surfaces an unnamed scan. What Drive's built-in OCR won't do is act on the result: it finds the scan, but renaming and filing it remains manual or remains a tool's job.",[87,88,90],"h3",{"id":89},"related","Related",[10,92,93,97,98,102],{},[61,94,96],{"href":95},"\u002Fglossary\u002Forganize-invoices-in-google-drive","How to organize invoices in Google Drive"," shows the most common use of this pipeline end to end, and ",[61,99,101],{"href":100},"\u002Fglossary\u002Fautomatic-file-sorter","automatic file sorter"," covers where OCR-based sorting sits among the other sorting approaches.",{"title":75,"searchDepth":104,"depth":104,"links":105},2,[106,107,108],{"id":16,"depth":104,"text":17},{"id":35,"depth":104,"text":36},{"id":81,"depth":104,"text":82,"children":109},[110],{"id":89,"depth":111,"text":90},3,"Definitions","OCR turns the text in scanned documents, PDFs, and photos into machine-readable text — the step that makes it possible to sort OCR documents by content.",false,"md",{},true,"\u002Fglossary\u002Focr",{"title":5,"description":113},"glossary\u002Focr","RwqafUgmQhLHvJvfCbymIV8zoWnfmkRR89Ey3O8NZlY",1785276343831]