OCR PDF

Extract text from a scanned PDF, page by page, using OCR. Runs entirely in your browser.

Extract text from a scanned PDF

Upload a PDF, choose a language, then run OCR on every page.

Drag & drop a PDF file here, or click to browse
Choose a single scanned or image-based PDF file.

    Extracted text

    Your PDF is processed locally in this browser tab using pdf.js and Tesseract.js (WebAssembly OCR). It is never uploaded to Calcooo or any server — only the rendering and OCR engines are fetched from a CDN.

    How to use this tool

    1. Drag and drop your scanned PDF into the box above, or click to browse and select it.
    2. Choose the language that matches the text in the document.
    3. Press Run OCR and wait while each page is rendered and recognized.
    4. Copy the combined text or download it as a .txt file.

    Frequently asked questions

    Is my PDF uploaded anywhere?

    No. Every page is rendered and recognized entirely locally in your browser, using pdf.js to render pages and Tesseract.js to run OCR. Your file is never uploaded.

    When do I need this instead of a normal text-extraction tool?

    Use this tool when your PDF is a scan or photo of a document — the "text" is actually part of a page image, so a normal PDF-to-text tool would return nothing. This tool reads the image content directly via OCR.

    How long does this take?

    OCR runs per page and can take a few seconds per page depending on your device and the page's content — larger, multi-page PDFs take proportionally longer.

    Which languages are supported?

    English, Spanish, French, German and Hindi, selectable before running OCR.