Each of these parsers carries two licenses, one for the code and one for the weights it downloads, and for Marker and MinerU they point in opposite directions. Docling is MIT code with default weights tagged apache-2.0, cdla-permissive-2.0 and mit. Marker v2.0.0 moved its code from GPL-3.0 to Apache 2.0 on 2026-07-20, but every Surya weight it loads, including the layout model used with OCR disabled, sits under a modified OpenRAIL-M with three separate triggers: more than $5,000,000 prior-year revenue, more than $5,000,000 total funding, or offering anything that competes with Datalab. MinerU weights for the 4.0 torch and vLLM paths are tagged apache-2.0, but the code is Apache 2.0 plus a USD 20 million consolidated monthly revenue (or 100 million MAU) threshold, an attribution duty for online services, and automatic termination. MinerU 4.0.7's default llama.cpp weights come from a Hugging Face repo outside the opendatalab org with no license tag at all. This week, stage each parser's models once, list every weight repo and revision that landed in the cache, attach the dated license text, and send that list to counsel.
All three parsers in the "docling vs mineru vs marker" comparison are open source on GitHub, and all three shipped a release in the last ten weeks: Docling v2.130.0 on 22 September 2026, MinerU 4.0.7 on 23 September, Marker v2.0.0 on 20 July. Install any of them with pip and you get a working PDF-to-Markdown pipeline in minutes. The terms a regulated bank, insurer or pharma group operates under are completely different across the three, and the difference does not show up on the GitHub license badge.
The principle that organizes this post: every parser here has two licenses. One covers the code you install from PyPI. The other covers the model weights the code downloads on first run, and those weights are what actually parse your documents. For Docling both are permissive. For Marker and MinerU they point in opposite directions. Marker is permissive code over restricted weights. MinerU is restricted code over weights that are mostly tagged apache-2.0, with exceptions that matter.
If the question you are still answering is whether to self-host a parser at all or buy a managed API, our comparison of Reducto, LlamaParse, Unstructured and Docling covers that decision. This post assumes you are self-hosting inside a regulated perimeter and need to know which of these three you can run, at your size, in your deployment shape.
Docling vs MinerU vs Marker: why the license decides first
Parser accuracy is something you can measure on your own corpus in a week. A license condition is something you discover in procurement, or worse, after the pipeline has indexed the whole archive and legal asks which model produced them. Re-parsing a production corpus is a re-index, and a re-index forced by a license finding runs on legal's schedule rather than yours.
So the order of operations we recommend is inverted from the usual one: read the terms first, eliminate what your organization cannot run, then benchmark what is left.
One statement, once, so it does not have to be hedged in every paragraph: this is not legal advice. The license text and your counsel decide. What follows quotes the text verbatim, dates it as of September 2026, describes what the words say, and names the clause to put in front of a lawyer. Where a question needs a legal reading, we say so and stop.
The license matrix as of September 2026
Start with what each tool declares and what it pulls. Weight licenses below are the Hugging Face license tags on the default repos, fetched September 2026.
Then the conditions attached to each, which is the table your reviewer actually needs.
One more line belongs under the Docling row. Docling 2.130.0 ships VLM presets that load third-party checkpoints, including datalab-to/chandra-ocr-2 (tagged openrail) and opendatalab/MinerU2.5-Pro-2604-1.2B (apache-2.0). The weight license travels with the checkpoint you select, not with the MIT wrapper that loads it. A Docling deployment is only as permissive as its pipeline configuration.
| Code license (as the package declares it) | Default weight repos | Weight license tags | |
|---|---|---|---|
| Docling 2.130.0 | MIT | docling-layout-heron, docling-models (TableFormer, revision v2.3.0), CodeFormulaV2, DocumentFigureClassifier-v2.5; granite-docling-258M for the VLM pipeline | apache-2.0; cdla-permissive-2.0 + apache-2.0; cdla-permissive-2.0; mit; apache-2.0 |
| MinerU 4.0.7 | LicenseRef-MinerU-Open-Source-License (Apache 2.0 plus additional terms) | MinerU-4_models_onnx or MinerU-4_models_torch; MinerU2.5-Pro-2605-1.2B, or a GGUF build of it hosted on an individual Hugging Face account on the llama.cpp path | onnx: no tag; torch: apache-2.0; Pro-2605: apache-2.0; GGUF: no tag |
| Marker 2.0.0 | Apache-2.0 | Loaded through Surya (below) | Modified OpenRAIL-M, tagged openrail |
| Surya 0.22.1 (Marker's model layer) | Apache-2.0 | datalab-to/surya-ocr-2, surya-ocr-2-gguf, surya_layout2 | openrail on all three; LICENSE file is the modified OpenRAIL-M |
| Revenue or funding trigger | Attribution duty | Competitor clause | Termination or remote restriction | What a scanner reports | |
|---|---|---|---|---|---|
| Docling | None in the MIT file or default weight tags | Standard MIT and Apache notice terms | None | None beyond the standard terms | MIT |
| MinerU (code) | Consolidated MAU above 100 million, or total monthly revenue above USD 20 million | Section 2, for online services to third parties | None | Section 3: automatic, no notice | LicenseRef on PyPI, NOASSERTION on GitHub |
| Marker and Surya (weights) | Prior-year gross revenue above $5,000,000, or total equity or debt funding above $5,000,000 | Clause 7: credit and attribution to Licensor in connection with any Output | Attachment A 2(c) | Clause 9: licensor may restrict usage remotely | Apache-2.0 for the package; openrail for weights |
Marker: Apache code, OpenRAIL-M weights, and three separate triggers
Marker's README summarizes its weight license as "free for research, personal use, and startups under $5M funding/revenue". That summary is shorter than the clause it describes. Here is Attachment A, section 2 of MODEL_LICENSE, headed "Commercial:", verbatim as of September 2026:
(a) for any purpose if You (your employer, or the entity you are affiliated with) generated more than five million US Dollars ($5,000,000) in gross revenue in the prior year, except where Your Use is limited to personal use or research purposes; > (b) for any purpose if You (your employer, or the entity you are affiliated with) has raised more than five million US dollars ($5,000,000) in total equity or debt funding from any source, except where Your Use is limited to personal use or research purposes; or > (c) for any purpose if You (your employer, or the entity you are affiliated with) provides or otherwise makes available any product or service that competes with any product or service offered by or made available by Licensor or any of its affiliates.
Read textually, three things stand out, and none of them is a legal conclusion:
Three more clauses belong on counsel's desk. Clause 7, "Attribution", requires you, "In connection with any Output", to "give appropriate credit and attribution to Licensor". Clause 8, "Share-a-Like", requires you to apply the license "to any and all copies of the Model, Derivatives of the Model ... and to the Output and any derivatives, changes or improvements to or of the Output." Clause 9, "Updates and Runtime Restrictions", states that the "Licensor reserves the right to restrict (remotely or otherwise) usage of the Model in violation of this License". What "Output" covers when the output is Markdown extracted from your own documents, and what remote restriction could mean for an air-gapped install, are questions for a lawyer, not for this post.
A requirements file pinned to marker-pdf 1.10.2 is GPL-3.0 code over weights with $2M triggers. A pin on 2.0.0 is Apache-2.0 code over weights with $5M triggers. The weights in your cache, not the version you think you run, decide which text applies. For use beyond those terms, Datalab sells a commercial license, including an on-prem option, listed on Datalab's pricing page; we have not reviewed its terms.
The history, dated
Older comparisons that quote a $2M threshold were right at the time. The dates:
| Date | Commit or release | Change |
|---|---|---|
| 2025-10-21 | 3c12b302 | MODEL_LICENSE added with "two million US Dollars ($2,000,000)" in 2(a) and 2(b) |
| 2026-01-31 | v1.10.2 released | Code GPL-3.0; README: "under $2M funding/revenue) and our code is GPL" |
| 2026-07-17 | 65f73c99 | Code LICENSE changed to Apache 2.0 |
| 2026-07-20 | 619377f1 | Both thresholds raised to $5,000,000 |
| 2026-07-20 | v2.0.0 | First release with Apache-2.0 code and $5M thresholds |
MinerU: Apache 2.0 plus thresholds, attribution and automatic termination
MinerU's LICENSE.md opens: "MinerU is licensed under Apache License 2.0 and is subject to the additional terms below. Except to the extent expressly modified or supplemented by these additional terms, your other rights and obligations are governed by Apache License 2.0." The additional terms, verbatim as of September 2026:
1. MinerU may be used for commercial purposes without a separate commercial license. However, if you and your Affiliates, on a consolidated basis, meet either of the following thresholds, you must obtain a separate commercial license from [MinerU Team] before continuing such use: a. monthly active users (MAU) exceed 100 million; or b. total monthly revenue exceeds USD 20 million. > 2. If you provide online services to third parties based on MinerU, you must clearly and prominently indicate, in the relevant product or service interface or in publicly available documentation, that MinerU is used. > 3. Where a separate commercial license is required under Section 1 but is not obtained before continuing such use, or where the attribution obligation under Section 2 is not complied with, this License and all rights granted under this License will terminate automatically, and no further notice from the Licensor is required. > 4. In these additional terms, "Affiliates" means any legal entity that directly or indirectly controls, is controlled by, or is under common control with you. "Control" means the power to direct the management and operating decisions of an entity, whether through equity ownership, voting rights, contractual arrangements, or otherwise.
The arithmetic of section 1 is simple and easy to misread. USD 20 million a month is USD 240 million a year, and it is total revenue of you and your Affiliates on a consolidated basis, not revenue attributable to MinerU. A large banking or insurance group clears that figure. Many community banks, credit unions and regional insurers do not, and a subsidiary's own revenue is not the number the clause names. Which side of the line your group sits on is a question for counsel with your consolidated accounts in front of them.
Section 2 is scoped to "online services to third parties". Whether an internal RAG system serving your own employees is such a service is exactly the kind of reading we will not make here. Section 3 is the clause to flag in bold for a reviewer: termination is automatic and requires no notice.
Commits are not what your lockfile installs, so check the PyPI metadata instead. Every release up to and including 3.0.9 (2026-04-07) declares AGPL-3.0 in its License field, although 3.0.0 to 3.0.9 also carry an "OSI Approved :: Apache Software License" classifier, so a scanner that reads classifiers sees a conflict. Every release from 3.1.0 (2026-04-17) onward, including the whole 4.0 line, declares LicenseRef-MinerU-Open-Source-License. No PyPI release was cut during either plain-Apache window, so there is no MinerU version whose License field says plain Apache-2.0.
One dependency note for a PDF-stack license review: MinerU 4.x depends on pypdfium2 (BSD-3-Clause and Apache-2.0), not PyMuPDF, which is dual-licensed under AGPL-3.0 or a commercial license.
Five license commits in four weeks
The LICENSE.md commit history, from the GitHub commits API:
| Date | Commit | Message |
|---|---|---|
| 2024-03-04 | 9fe81795 | Create LICENSE.md (AGPL era) |
| 2026-03-20 | 7409e645 | update license from AGPL-3.0 to Apache License 2.0 |
| 2026-03-28 | 18d26061 | update license from Apache 2.0 to AGPLv3 |
| 2026-04-14 | e148afa9 | update license from AGPL-3.0 to Apache-2.0 |
| 2026-04-17 | 2de34115 | update license information to include MinerU Open Source License with additional conditions |
| 2026-04-17 | 2f078fc6 | update LICENSE.md to clarify commercial use terms and attribution obligations |
The weights you pinned may not match the license you read
Package pins are not weight pins. Every parser here resolves model repos at runtime, and the repos carry their own license tags, which do not always agree with the README you read. For MinerU and Marker, the repos that matter, fetched September 2026:
Three consequences follow.
MinerU's default CPU install pulls the two untagged repos. The 4.0 README says the default install runs small models on ONNX CPU inference and the VLM on llama.cpp in Vulkan mode. That combination fetches MinerU-4_models_onnx, whose card lists PaddlePaddle-derived components (PP-DocLayoutV2, PP-OCRv6, PP-FormulaNet plus-M) and says their original licenses remain applicable, plus the GGUF files from a MinerU2.5-Pro-2605-1.2B-GGUF repo on an individual account outside the opendatalab org, containing the two Q8_0 files and a .gitattributes, with no license tag and no README. The GPU path pulls repos tagged apache-2.0. Same package, same version, different license posture depending on the backend flags.
Marker has no weight-free mode. Even with --disable_ocr in fast mode, the 20M-parameter layout model comes from datalab-to/surya_layout2, which carries the same modified OpenRAIL-M text as the OCR weights, including the $5,000,000 triggers and clause 2(c).
Tags hide modifications. All three Datalab weight repos are tagged just "openrail". A scanner that reads Hugging Face metadata records a standard OpenRAIL. The dollar thresholds and the competitor clause live in the LICENSE file inside each repo, which we verified is the same text as Marker's MODEL_LICENSE apart from whitespace.
The fix is procedural: inventory weights by repo id and revision from the cache you actually staged, not by package name from the requirements file.
| Weight repo | Where it is used | HF license tag | Note |
|---|---|---|---|
| opendatalab/MinerU2.5-2509-1.2B | Earlier MinerU VLM | agpl-3.0 | Card front matter still reads "license: agpl-3.0" |
| opendatalab/MinerU2.5-Pro-2604-1.2B | Previous Pro VLM; also a Docling preset | apache-2.0 | |
| opendatalab/MinerU2.5-Pro-2605-1.2B | MinerU 4.0.7 VLM on torch, vLLM, lmdeploy | apache-2.0 | |
| opendatalab/MinerU-4_models_torch | 4.0 small models, GPU path | apache-2.0 | |
| opendatalab/MinerU-4_models_onnx | 4.0 small models, default CPU path | none | Card: "Original model licenses and attribution remain applicable" |
| opendatalab/MinerU2.0-2505-0.9B | Earlier VLM | none | |
| MinerU2.5-Pro-2605-1.2B-GGUF (individual account) | 4.0.7 default llama.cpp VLM | none | Outside the opendatalab org; no model card |
| datalab-to/surya-ocr-2 | Marker VLM via vLLM | openrail | LICENSE file is the modified OpenRAIL-M |
| datalab-to/surya-ocr-2-gguf | Marker VLM via llama.cpp | openrail | Same text |
| datalab-to/surya_layout2 | Marker layout, including OCR-disabled mode | openrail | Same text |
Staging each parser's models for an air-gapped install
Inside an enclave the parser cannot download anything on first run, so staging is where the license inventory gets built as a side effect. The generic Hugging Face offline mechanics and an egress-deny smoke test are covered in our vLLM air-gapped deployment guide; what follows is only what differs per parser.
docling-tools models download -o /opt/models/docling
export DOCLING_ARTIFACTS_PATH=/opt/models/docling # or per invocation: docling --artifacts-path=/opt/models/docling report.pdf
mineru-kit models download --tier standard --small-backend onnx --vlm-engine llama-cpp --source huggingface mineru-kit models verify --tier standard --small-backend onnx --vlm-engine llama-cpp
export MINERU_MODEL_SMALL_BACKEND=onnx export MINERU_MODEL_VLM_ENGINE=llama-cpp export MINERU_MODEL_SOURCE=local mineru-kit parse document.pdf -o document.md --tier standard
export SURYA_INFERENCE_URL=http://127.0.0.1:8000/v1 # or, for the auto-spawned llama.cpp backend: export SURYA_GGUF_LOCAL_MODEL_PATH=/opt/models/surya/surya-2.gguf export SURYA_GGUF_LOCAL_MMPROJ_PATH=/opt/models/surya/surya-2-mmproj.gguf export MODEL_CACHE_DIR=/opt/models/datalab
Docling
pip install docling is a meta-package over docling-slim[standard], which brings RapidOCR, torch, docling-ibm-models, the office, web, LaTeX and email formats, and the CLI. Docling documents a prefetch command whose default set is layout, TableFormer, code formula, picture classifier and RapidOCR, written to ~/.cache/docling/models unless you pass an output directory. Run it on a connected machine: Copy that directory into the perimeter and point conversions at it. Either the environment variable or the flag is enough on its own: The Python equivalent is PdfPipelineOptions(artifacts_path=...). Remote API-backed options stay off: enable_remote_services defaults to False, and without it those options raise OperationNotAllowed.
MinerU 4.0
MinerU 4.0 replaced the old model workflow, so ignore any guide mentioning mineru.json or MINERU_MODEL_STACK (setting the latter now raises an error). On a build machine, download and verify for the backend the enclave will run, not the backend the build machine would pick automatically. For a CPU enclave: For an NVIDIA target the documented flags are --small-backend torch --vlm-engine vllm, which also swaps the untagged ONNX and GGUF repos for the apache-2.0-tagged torch and Pro-2605 repos. Models land under ~/.mineru/models (the configured model.base_dir). Copy the whole directory; the docs say not to create completion markers by hand. Inside the perimeter: With local, a missing file is an error instead of a silent fetch from Hugging Face or ModelScope. Leave model.source on auto and it probes Hugging Face first, then falls back to ModelScope. The CLI reference for mineru-kit models is still marked Draft, eight days after 4.0.0, so re-check the flags at each minor release.
Marker 2.0 and Surya 0.22.1
Marker has no documented pre-download command, and it pulls from three places: Hugging Face (surya-ocr-2 for vLLM, surya-ocr-2-gguf for llama.cpp, surya_layout2 for layout), the models.datalab.to host (the ocr_error_detection checkpoint, plus a font file, GoNotoCurrent-Regular.ttf, written into the package's static/fonts directory if it is missing), and on NVIDIA a vllm/vllm-openai:v0.20.1 Docker image that Surya spawns with your Hugging Face cache mounted. Marker's base converter calls that font download every time it is constructed, so a missing font fails the first conversion inside the perimeter; stage it. The workable procedure is to run a conversion once on a connected staging box with the same backend and mode, then copy the resulting caches, the Docker image and the font. Inside the perimeter, these settings from Surya's settings.py take the network out of the loop: SURYA_INFERENCE_URL attaches to a server you run and skips the auto-spawn. The two GGUF paths replace the Hugging Face download. MODEL_CACHE_DIR relocates the models.datalab.to checkpoints. Keep --use_llm off, or point it at a local OpenAI-compatible endpoint: its default is a cloud Gemini model. MinerU's default VLM engine and Marker's path on non-NVIDIA hardware both run on llama.cpp. Surya launches llama-server bound to 127.0.0.1 by default, the same default llama-server itself ships with; before anyone widens that bind address or exposes a shared inference server, read our notes on hardening a llama.cpp server.
What each parser is good at, dated and sourced
Format coverage stopped being a differentiator this year. Docling's standard install covers PDF, Office formats, web pages, LaTeX and email. MinerU 4.0 lists PDF, images, DOC/DOCX, PPT/PPTX, XLS/XLSX, RTF, OpenDocument, EPUB, OFD, HTML/MHTML and CSV/TSV, with everything except PDF and images handled by its local Flash native parsing. Marker lists PDF, image, PPTX, DOCX, XLSX, HTML and EPUB with marker-pdf[full].
Hardware defaults differ more. MinerU's base install is CPU-first, with roughly 0.8 GB of models for the basic tier and roughly 2 GB for standard on ONNX and llama.cpp (about 3 GB with vLLM), per its README table, and it recommends mineru[full] on NVIDIA. Marker defaults to balanced mode on GPU and fast mode on CPU, and --disable_ocr needs no inference server at all. Docling runs its default models locally through torch.
On accuracy, the only current numbers in hand are Datalab's own, from the Marker v2.0.0 release notes in July 2026: olmOCR-bench 76.0 in balanced mode, 66.6 in fast mode and 43.6 with OCR disabled, and a claim of more than 5x the pages per second of MinerU's pipeline backend. Those are vendor self-reported and we have not reproduced them. We have no neutral head-to-head of all three at current versions, so we do not rank them.
What decides accuracy is your corpus. Before any parser change reaches a production index, run the shadow-diff loop described in our long-document parsing post (its section "The shadow-diff loop before a production index moves"). Score the parser's structured output where it gets consumed, by the chunker, as covered in chunking for context preservation. If tables and charts carry the answers, consider whether page-image retrieval beats OCR for that slice of the corpus.
Decision table: organization size, deployment and document type
This table routes situations to the clause that needs reading. It does not tell anyone they are compliant. Seven situations:
Where our own recommendation lands, for a regulated organization self-hosting a RAG ingestion path: Docling is the one whose default code and weights carry no size trigger, no competitor clause and no termination-without-notice term, and whose offline staging is a documented single command. Use it as the baseline. Evaluate MinerU on a GPU path, where its weights are tagged apache-2.0, once counsel has read sections 1 to 3 against your consolidated figures. Evaluate Marker only with a commercial license in hand or a clear counsel reading of 2(a) through 2(c), because every weight it loads carries those clauses. If your system falls under the EU AI Act high-risk regime, the same repo-plus-revision record belongs in the Annex IV technical file for an open-weight model. For how parsing fits the wider retrieval stack, the rest of our RAG systems guides cover what comes after it.
The action for this week is small. On a connected staging machine, run each candidate parser's model staging for the exact backend you will deploy. List every weight repo and revision that landed in the cache, next to the dated license text from that repo's LICENSE file rather than its tag. Send that list, with the clauses quoted above, to counsel before the first production ingest.
| Situation | Docling | MinerU | Marker |
|---|---|---|---|
| Under $5M prior-year revenue and under $5M total funding, internal use | Default weights tagged permissive; check any VLM preset you enable | Below the section 1 revenue figure; section 2 only if you serve third parties | 2(a) and 2(b) not met on these figures; 2(c) and clauses 7, 8 and 9 still apply |
| Above either $5M test, under USD 20M consolidated monthly revenue, internal RAG only | No threshold in the text | Below the section 1 revenue figure; confirm consolidated Affiliate revenue with counsel | Words of 2(a) or 2(b) met; Datalab commercial license or counsel review |
| Group above USD 20M monthly revenue or 100M MAU | No threshold in the text | Section 1 names a separate commercial license; section 3 terminates automatically | Words of 2(a) met; Datalab commercial license or counsel review |
| Consultancy, document-AI vendor, or any product that could compete with Datalab | No competitor clause | No competitor clause | 2(c) applies regardless of size and, as written, has no research carve-out |
| Parsing exposed as an online service to third parties | MIT and Apache notice terms | Section 2 attribution; section 3 if missed | Clause 1(f) counts hosted API access as Distribution; read with clauses 7 and 8 |
| Fully air-gapped, CPU-only enclave | One prefetch command plus --artifacts-path | ONNX and llama.cpp path pulls the two untagged repos | Fast mode runs on CPU; no prefetch command; stage from three sources |
| Scanned, formula-heavy or VLM-dependent corpus | VLM pipeline default is granite-docling-258M (apache-2.0) | VLM weights: tagged apache-2.0 on GPU, untagged GGUF on llama.cpp | Balanced mode needs a VLM server; all weights modified OpenRAIL-M |
FAQ
Quick answers to the questions this post tends to raise.



