Text layers are used directly
Every page is checked for embedded text first. Anything exported from Word, InDesign or LaTeX converts in seconds with perfect accuracy — no recognition guesswork at all.
Most PDFs already carry their text, and those convert in seconds. Only genuinely scanned pages go through OCR — in parallel, across every core your machine has.
Read on your device. Nothing is uploaded anywhere.
These only affect pages that actually need OCR.
Leave the tab open — the work happens here, not on a server.
Text layer OCR Waiting
A converter that reads every page as a picture spends minutes doing work the PDF had already done for it.
Every page is checked for embedded text first. Anything exported from Word, InDesign or LaTeX converts in seconds with perfect accuracy — no recognition guesswork at all.
Scanned pages are shared across several recognition workers at once instead of queuing behind a single one, which is roughly a four-fold difference on a typical laptop.
The next page is drawn while the current one is being read, so the processor is never sitting idle waiting for the other half of the job.
Proper table of contents, unique identifier, split chapters and an uncompressed mimetype entry — so Calibre, Apple Books and Kindle Previewer all open it without complaint.
Tamil, Hindi, Malayalam, Telugu and Kannada alongside English. Picking English only when that is all you have roughly halves the recognition time.
Change your mind on page 300 and the pages already read are still assembled into a usable EPUB, rather than the whole run being thrown away.
If the PDF already contains text — most do — a three hundred page book takes a few seconds. A genuinely scanned book needs OCR on every page, which is roughly one to three seconds per page spread across your processor cores, so expect a couple of minutes for two hundred pages.
Good enough to read and search, not good enough to publish unproofed. Tamil is a complex script and recognition accuracy on a clean scan is typically in the nineties, which still means several errors per page. Straight, high-contrast scans do much better than photographs of pages. Always read through the result before sharing it.
No. The PDF is read and the EPUB assembled inside your browser tab. The only network traffic is downloading the recognition language data the first time you use it, which is then cached for later runs.
The recognition engine downloads a language model, which for Indian scripts can be several megabytes. It is stored in your browser afterwards, so the second conversion in the same browser starts almost immediately.
No, and that is the point of the format. EPUB reflows text to fit whatever screen it is on, so fixed columns, page numbers and image placement do not carry over. If you need the layout preserved exactly, keep the PDF.
Yes. Put a range like 1-40 in the pages box to test the settings on a chapter before committing to the whole thing — a sensible habit with a scanned book, since it tells you within a minute whether the OCR quality is worth the wait.