technologies · difficulty ◆◆
Docling — turn messy documents into clean, AI-ready data
Your documents, ready for the model.
You know the failure. You paste a PDF into a model and it reads a two-column journal article as one long run-on sentence, mangles every table, and quotes the page header as if it were content. Docling is the parsing front-end built to stop that — it turns a messy document into structured, machine-readable data before an LLM ever sees it.
$ doclingWhat it is
Docling is an open-source Python library that converts unstructured documents into structured, AI-ready data. You hand it a PDF, a scanned document, a Word or PowerPoint file, an image, or a URL, and it parses the content, understands the layout (headings, paragraphs, lists, tables, figures, reading order), optionally runs OCR on scanned pages, and exports a clean representation: Markdown, HTML, or a rich JSON that preserves the document’s structure and metadata.
Why it matters
Raw documents are the weak link in retrieval-augmented generation. An LLM fed a raw PDF often cannot tell a headline from a footer, a table column from a paragraph, or a citation from the main text. Docling exists to solve exactly this "garbage in, garbage out" problem. It is the document-parsing front-end that makes content reliable for downstream AI — clean structure, real tables, reading order preserved, with provenance back to the source page. Feed Docling output to a RAG pipeline and you get usable facts instead of mashed text.
How it fits your stack
You can use Docling three ways, depending on how much you want to run. The Python library gives you the DocumentConverter for batch or single-file work inside a script. The CLI wraps the same pipeline for quick conversions from the terminal. And docling-serve exposes the whole thing as a REST API plus a Gradio demo UI, containerized in a ready-to-run Docker image. The output formats overlap everywhere — Markdown for humans and simple RAG, JSON for pipelines that need the full structure, and HTML when you want to keep the document’s look.
Example
$ docling convert myfile.pdf INFO Parsing document myfile.pdf...
INFO Converting with layout model (fast)...
INFO Exporting to Markdown...
output: myfile.mdThe default conversion — a PDF becomes Markdown in the current directory.
$ docling convert a.pdf b.docx c.pptx --to html --output-dir ./out INFO Parsing a.pdf ... exporting to ./out/a.html
INFO Parsing b.docx ... exporting to ./out/b.html
INFO Parsing c.pptx ... exporting to ./out/c.htmlMultiple formats in one pass, exported as HTML into an output directory.
$ docling convert scan.pdf --to md --do-ocr INFO OCR enabled for image-based pages
INFO Parsing scan.pdf ... reading scanned text via OCR
INFO Exporting to Markdown...Flag --do-ocr makes scanned, image-only pages machine-readable.
$ python -c "from docling.document_converter import DocumentConverter; c = DocumentConverter(); r = c.convert(\"myfile.pdf\"); print(r.document.export_to_markdown()[:200])" # Myfile
## Section 1
This is the first paragraph of the document, cleanly separated from
the heading above it. The table below survives intact:
| Col A | Col B |
|-------|-------|
| 1 | x |The same pipeline from Python. r.document.export_to_markdown() gives you clean structure, not a text dump.
Common flags
- convert <file>
- convert a document (default output: Markdown)
- --to <fmt>
- output format: md (default), html, json, doctags, text, pdf
- --output-dir <dir>
- write output files into a directory
- --do-ocr
- run OCR on scanned / image-only pages
- --table-mode <mode>
- table recognition: fast (default) or accurate
- --layout-engine <engine>
- layout model to use, e.g. rms (docling-parse) or other
- from_format / from_formats
- input format hints (API + CLI)
- to_formats
- comma-separated output formats via the REST API
History
Origin
Docling was created and is maintained by IBM Research. It grew out of the need to feed real-world enterprise documents into AI pipelines that simply could not cope with raw PDFs — text columns, headers, footers, and tables were destroying retrieval quality. IBM open-sourced it under the MIT license, and it has become one of the fastest-growing tools in the document-intelligence and RAG space, passing roughly 64,000 GitHub stars.
The layout-first insight
The core idea is that document understanding is a layout problem before it is a text problem. Docling treats the page as a set of regions — headings, paragraphs, tables, figures — and reconstructs reading order rather than just dumping characters. That is why its Markdown output is structured: the model knows where a heading ends and a paragraph begins, where a table’s cells are, and which page a given block came from. Provenance is baked in, so every element traces back to source coordinates.
From library to service
As Docling matured it split into a conversion core and a serving layer. The core library does the parsing; docling-serve wraps it in a REST API with sync and async endpoints, a Gradio demo UI, and official Docker images (CPU and CUDA). That made it possible to run document conversion as an internal service that other microservices call over HTTP — no ML knowledge required on the consuming side.
Fun facts
Pros & cons
pros
- + Layout awareness — detects headings, tables, figures and reading order, not a flat text dump
- + Multi-format input: PDF, DOCX, XLSX, PPTX, images, HTML, Markdown, and email (.msg)
- + Rich JSON output preserves blocks, tables, provenance and metadata for RAG pipelines
- + Real OCR support (Tesseract, EasyOCR, RapidOCR, or VLM) for scanned archives
- + Self-contained serving: REST API + Gradio UI + official Docker images (CPU and CUDA)
- + Active and fast-moving — MIT-licensed with frequent releases
cons
- − Heavy ML dependencies (PyTorch + model weights) for a full install
- − First boot pre-fetches the model set (~1-2GB) into the artifacts volume — a one-time wait before the API comes up
- − OCR is opt-in; skipping it on scanned PDFs gives empty output
- − CUDA Docker images have no "latest" tag — you must pin :main or an explicit version
Takeaways
- 1Try it fastest: `pip install docling` in a venv, then `docling convert myfile.pdf`
- 2For scanned files always add `--do-ocr` — born-digital PDFs are fine without it
- 3Export JSON (not just Markdown) when you feed a RAG pipeline — you keep structure, tables and provenance
- 4Use `--table-mode accurate` for dense tables you need as real data
- 5Deploy docling-serve via the setup script — it pre-fetches models, works on Linux and macOS, and runs from anywhere
- 6Manage the service with `./setup-docling.sh --down` / `--restart` / `--no-browser`
- 7Persist the model cache (DOCLING_SERVE_ARTIFACTS_PATH or ~/.cache/docling/models) to avoid re-downloading weights