MinerU2.5-Pro is a 1.2B-parameter document parsing vision-language model from OpenDataLab, the team behind the MinerU PDF-to-Markdown toolchain. It converts page images of PDFs and scanned documents into structured output — Markdown text, HTML tables, and LaTeX formulas — and is the current flagship checkpoint of the MinerU 2.5 line (opendatalab/MinerU2.5-Pro-2605-1.2B).
Key Features
MinerU2.5 adopts a coarse-to-fine, two-stage parsing strategy: it first performs global layout analysis on a downsampled view of the page, then runs fine-grained content recognition on native-resolution crops of each detected block. This decoupled design lets a small 1.2B model handle high-resolution documents with low computational overhead while preserving non-body elements such as headers, footers, and page numbers.
The Pro checkpoint was produced purely through data engineering on the same 1.2B architecture: the training corpus was expanded from under 10 million to 65.5 million pages with difficulty-aware sampling, and annotations for complex tables and dense formulas were regenerated with cross-model consistency verification and an iterative judge-and-refine pipeline, followed by three-stage progressive training ending in GRPO format alignment.
Beyond core parsing, MinerU2.5-Pro adds practical capabilities for production document pipelines: image and chart analysis, truncated paragraph merging across blocks, in-table image recognition, and improved handling of rotated, borderless, and partially bordered tables. Mixed-language Chinese-English formulas and long, complex equations are parsed to LaTeX.
Architecture
The model is built on the Qwen2-VL architecture (Qwen2VLForConditionalGeneration) with a vision encoder feeding a compact 1.2B language decoder, and ships in Safetensors format on the Transformers framework. It serves with vLLM out of the box; the publisher recommends vLLM for inference and reports concurrent throughput of about 2 frames per second on a single data-center GPU with the asynchronous engine.
Each parsed page is returned as a list of content blocks with a type (text, table, equation, or image), a normalized bounding box, an optional rotation angle, and the recognized content. The publisher's mineru-vl-utils Python package wraps this protocol: point its http-client backend at the served OpenAI-compatible endpoint and call two-step extract to get structured blocks or converted Markdown for each page image.
Use Cases
MinerU2.5-Pro is designed for document understanding workloads at scale:
- PDF-to-Markdown conversion for LLM training data and RAG ingestion pipelines
- Table extraction from complex, rotated, or borderless layouts into HTML
- Formula recognition to LaTeX for scientific and technical documents
- Layout analysis with reading-order recovery across multi-column pages
- Chart, flowchart, and seal recognition within documents
Performance
On OmniDocBench v1.6, MinerU2.5-Pro achieves a state-of-the-art overall score of 95.7, outperforming both specialized OCR models and much larger general-purpose frontier vision-language models, and improving substantially over the original MinerU2.5 baseline score of 92.98. Text recognition reaches an edit distance of 0.036, dense formula parsing scores 97.15 CDM, and table recognition reaches 93.62 TEDS across the benchmark, with leading results reported across five diverse table benchmarks.