Skip to content

Documents & OCR

From documents to data your AI can use.

Extract text, tables, and figures from PDFs and scans. Review them alongside the original and take the result to your own AI tools.

From PDF to editable text, tables, and figures.

Extract. Review. Reuse.

Keep the original and the extracted content together while you prepare your documents.

  1. 01

    Upload your documents

    Create a document dataset and add PDFs, PNGs, or JPEGs. OCR extracts the content page by page.

  2. 02

    Review the result

    Compare the original with the extracted pages. Correct the Markdown, check figures and tables, and approve the document.

  3. 03

    Preserve your work

    Save a dataset version to keep a stable reference. Export that version or continue working with the current dataset.

Choose how your work continues.

Both exports include your OCR corrections and extracted figures.

Markdown ZIP

A Markdown file for each document, with figures in an images folder. Use it in your own document pipeline or AI tools.

Best for reusing extracted content.

Example archive structure

document.md
images/
  document/
    figure.png

LLM Wiki

An Obsidian-compatible workspace with original documents, corrected transcriptions, an index, and instructions for an AI agent.

Your agent builds and explores knowledge from the exported sources.

Example archive structure

README.md
AGENTS.md
raw/sources/
wiki/
  index.md
  sources/

Models

Coming soon

From reviewed documents to your model.

Prepare the text and domain-specific examples that will form the basis of your own language model. Explore the document workflow for LLM fine-tuning.

Work from your AI agent

Keep your agent connected to the source.

Through MCP, your agent can read document pages, edit OCR text, update review status, and retrieve exports from your Annota AI workspace.

Connect your AI agent
AI agent - Annota AI MCP / Example

A few practical details.

Which files can I upload?

Document datasets accept PDF, PNG, and JPEG files. The upload screen shows the file size limit and OCR credit cost per processed page.

Does an LLM Wiki export answer questions for me?

The export gives you source files and instructions. You open it with your own compatible AI agent, which can then work with those sources. Hosted document Q&A in Annota AI is part of Brain.

Where can I check costs and data use?

Pricing explains plans, credits, and usage examples. Privacy & Security explains hosting, providers, access, and how your data is used.

The next step

Meet Brain.

Brain will bring connected knowledge and document Q&A into Annota AI. It builds on the same idea: useful answers start with source material you can trust.

Explore the vision

Start with a document.

Create a free account, upload a PDF or scan, and review your first OCR result.

Start for free