The lecturer posts slide screenshots instead of the original deck, and when I revise I want the key points as my own notes — typing them out from an image one character at a time is painfully slow. I tried OCR to turn the screenshots straight into editable text a few times, and it's far faster; I also learned what it can and can't do.
The flow is simple. Drop the slide screenshot into ImgIng's (imging.ai) text recognition, it reads locally, and gives a version laid out in the original rows; copy it straight into my notes and the lines and paragraphs are mostly right, no re-layout needed. The image didn't leave the machine, which I like for slides I shouldn't be spreading around.
Input quality decides accuracy more than anything. A screenshot always beats a photo — a screenshot is pixel-clean, a photo brings blur, glare and shadow. If you can only photograph, shoot straight on with even light and don't let the page curl warp the letters. I tried a hand-shaken photo of a whiteboard once and half of it came out wrong; a clean screenshot of the same thing was near-perfect.
A couple of settings help. It defaults to the Pro tier on desktop, fine for most slides; if the slides have Japanese you need Pro or Ultra, the fast tier doesn't read Japanese. Very small, dense text scans more reliably on Ultra. There's also a Markdown export I like for slides with headings and bullet lists — the recognised hierarchy comes out as Markdown syntax, so pasted into a Markdown note app the headings are headings and lists are lists, no manual formatting. Plain paragraphs go to TXT, structured content to Markdown.
Don't use it blind, scan a few spots after: numbers and formula symbols are the most error-prone — a 0/8 or l/1 swap in your notes misleads you at revision time; technical terms and English abbreviations also slip. Low-confidence spots get flagged, and checking along the marks beats re-reading everything.
And a boundary to know so you don't waste effort: handwriting and the semantics of math formulas, it can't do or can't do reliably. Printed slide screenshots are its home ground; a lecturer's handwritten derivation or complex formula layout is past current OCR's scope — sort those out yourself. Korean's out too, the current model doesn't include it.
The biggest win isn't even the typing time. Image notes can't be searched — to find where a concept was, you flip through screenshots one by one; as text, Ctrl+F lands on it instantly. That's the real reason I OCR a batch of slides first, not just laziness about typing.
Top comments (0)