PaddleOCR Online · Nothing to install

Drop in a scan.
Get back structure.

Point PaddleOCR at a PDF, a photo, or a web page. It reads the text, rebuilds the tables and formulas, and hands back Markdown, HTML, or JSON — in more than 100 languages.

Pages recognized daily
1.2M+
Pages recognized daily
Input formats
30+
File & web inputs
Service availability
99.9%
Service availability

PaddleOCR Workbench

Upload · recognize · export in one place

Drag a document in, or click to choose

Batch friendly · up to 50 MB per file · layout kept intact

Output

Jobs

3 jobs
List of recognition jobs
Task Status Type Model Created Actions
NeurIPS-2024-paper.pdf 2.4 MB · 38 pages
Completed Document PaddleOCR VL 2026/09/17 14:24
arxiv.org/abs/2409.xxxxx Web snapshot
Completed Web PaddleOCR Web 2026/09/17 13:50
organic-chemistry-lab-manual.pdf 18.7 MB · 214 pages
Recognizing Document PaddleOCR VL 2026/09/17 14:58
polymer-characterization-report.docx 6.1 MB · 62 pages
Completed Document PaddleOCR VL 2026/09/17 12:07
nature.com/articles/s41586-024-07… Web snapshot · 14 formulas
Completed Web PaddleOCR Web 2026/09/17 09:32

Output format examples

Examples of structured document output

Static illustrations of chemistry, math, and table output formats. These examples are not measured recognition results.

Case 01 · Chemical structures

Molecular diagrams become searchable chemistry

A chemical structure represented as SMILES, a molecular formula, and LaTeX. The LaTeX command requires mhchem support.

Input illustration
Acetylsalicylic acid OC(=O)CH₃ COOH Input illustration
Output format example
# diagram → structured chemistry
{
  "type": "chemical_structure",
  "smiles": "O=C(O)C1=CC=CC=[C]1[OC(=O)CH3]",
  "formula": "C9H8O4",
  "iupac": "2-acetoxybenzoic acid",
  "latex": "\\ce{CH3COOC6H4COOH}"
}

# inline Markdown output
acetylsalicylic acid (aspirin, $C_9H_8O_4$, SMILES: O=C(O)C1=CC=CC=[C]1[OC(=O)CH3])
Image reference An image reference preserves the illustration without encoding its molecular structure.
Structured representation SMILES describes molecular connectivity; the formula and LaTeX provide additional text representations.
Case 02 · Math

Scanned math → clean LaTeX

Two formulas shown as LaTeX source with display-math delimiters, for use in a compatible math renderer.

Input illustration
Output format example
# LaTeX source for a math renderer
\[e^{i\pi} + 1 = 0\]

\[x = \frac{-b \pm \sqrt{b^{2}-4ac}}{2a}\]

# human-readable text kept too
e^(iπ) + 1 = 0
Case 03 · Complex tables

Merged headers → structured rows

Markdown uses flattened headers; JSON describes the original table and header merges with zero-based row and column indexes. All numbers are illustrative data.

Input illustration
Example data by script (%) Script Printed Handwritten Mean SD Mean SD Latin 99.1 0.21 96.8 0.74 CJK 98.4 0.33 95.2 0.88 Arabic 97.6 0.41 93.9 1.02 Cyrillic 98.8 0.29 94.7 0.95
Output format example
# merged headers flattened to Markdown
| Script | Printed mean | Printed SD | Handwritten mean | Handwritten SD |
| --- | ---: | ---: | ---: | ---: |
| Latin | 99.1 | 0.21 | 96.8 | 0.74 |
| CJK | 98.4 | 0.33 | 95.2 | 0.88 |
| Arabic | 97.6 | 0.41 | 93.9 | 1.02 |
| Cyrillic | 98.8 | 0.29 | 94.7 | 0.95 |

# original table shape and header merges
{
  "rows": 6,
  "cols": 5,
  "merged": [
    {"row": 0, "col": 1, "rowspan": 1, "colspan": 2},
    {"row": 0, "col": 3, "rowspan": 1, "colspan": 2}
  ]
}

Capabilities

An OCR engine built for difficult documents

Plain text is the easy part. PaddleOCR keeps the layout, the semantics, and the reading order of complex pages.

Chemistry · math · tables

01 / SCIENCE

Scientific layout, understood

Built for chemical structures, mathematical notation, complex tables, and multi-column pages — reading order and logical structure preserved.

Chemical formulas LaTeX math Spanning tables
{ "type": "title", "level": 1 "bbox": [72, 96, 540], "text": "3. Methods" }, "type": "table", "html": "<table>…" }, …] Semantics · bbox · hierarchy

02 / AGENTS

Native input for AI agents

Semantically tagged JSON for headings, paragraphs, tables, figures, and formulas — straight into RAG, knowledge bases, and agent workflows.

Structured JSON RAG-ready Semantic tags
Layout 0.42s Formulas 0.61s Tables 0.77s 100 pages · 1.8s total

03 / SPEED

Fast, elastic compute

Elastic scheduling recognizes hundred-page documents in seconds, and batch jobs scale out in parallel for enterprise throughput.

Sub-second jobs Parallel batches Elastic scaling

Who uses PaddleOCR

For the teams who live in documents

From a single reading list to an organisation-wide archive, PaddleOCR sits at the gate where documents enter your systems.

Overhead view of a researcher's desk with papers and a tablet showing PaddleOCR text recognition

Researchers

Research & academia

Digitize papers, patents, and lab notebooks in bulk. Formulas and chemical notation survive, so literature reviews stop being manual transcription.

A developer integrating the PaddleOCR API at a night-time workstation

Developers

Developers & AI teams

One API call returns semantically tagged JSON for your RAG pipeline, knowledge base, or agent workflow — no recognition stack of your own to maintain.

Operations staff scanning invoices while a wall monitor shows PaddleOCR extraction metrics

Enterprise

Operations & finance

Invoices, contracts, and reports come back as rows and fields your systems can use — with on-premise deployment and usage reporting when you need them.

API reference

Three lines of code to structured output

A REST API and SDKs cover upload, status, and download — clear error codes and webhooks included, so integration stays painless.

  • REST endpoints with cURL, Python, Node.js, and Go examples
  • Sync responses or async jobs — your choice
  • Structured error codes with retry guidance
  • Webhooks on completion — no polling loops
Browse the API reference
Python Node.js cURL
# pip install paddleocr
import paddleocr

client = paddleocr.Client(api_key="your-api-key")

job = client.parse.upload(
    file=open("paper.pdf", "rb"),
    output_format="markdown",
    model="paddleocr-vl"
)

# Wait for the job, then save the result
result = job.wait()
result.save("output.md")
print(result.pages, "pages recognized")

Price

Pay for pages, not promises

From solo researchers to enterprise teams. Every plan includes online recognition and result export.

Free

For trying things on real documents

$0/ month

No payment method required

  • 50 pages free, every month
  • PDF, images, and web links
  • Markdown / HTML export
  • 3 concurrent tasks
  • Community support
Start free

Pro · Monthly

For regular workloads

$9.9/ month

Monthly billing, cancel anytime

  • 10,000 pages every month
  • Priority access to PaddleOCR VL
  • Structured JSON and full API access
  • Batch jobs, concurrency, and webhooks
  • Email support
Go monthly

On-premise

Every page stays on your own metal

For finance, healthcare, government, and research, PaddleOCR runs fully on-premise — models and data never leave your network.

100% Pages stay local
Compliant Security framework ready
24/7 Dedicated support
Talk to us about deployment

Running privately today

  • A top-tier securities research institute Private recognition for an internal research knowledge base — 50,000 pages a day, zero data egress.
  • A tertiary hospital research department Structured extraction from medical literature and trial reports, fully within research data rules.
  • A provincial government cloud Digitized archives of official documents, powering intelligent search and Q&A.
A symmetric GPU server aisle with blue LEDs in an on-premise data center
Models and data stay inside your network — built for finance, healthcare, and government compliance.