Model Library/MinerU2.5-Pro

MinerU2.5-Pro

Vision Language

Document parsing vision-language model that converts PDFs and page images to structured Markdown

On-Demand Dedicated 1xRTX 4090

Details

Modalities

vision

Recommended Hardware

1xRTX 4090

Estimated Price

Loading...

Provider

OpenDataLab

Family

MinerU

Parameters

1B

Context

8192 tokens

License

apache-2.0

MinerU2.5-Pro is a 1.2B-parameter document parsing vision-language model from OpenDataLab, the team behind the MinerU PDF-to-Markdown toolchain. It converts page images of PDFs and scanned documents into structured output — Markdown text, HTML tables, and LaTeX formulas — and is the current flagship checkpoint of the MinerU 2.5 line (opendatalab/MinerU2.5-Pro-2605-1.2B).

Key Features

MinerU2.5 adopts a coarse-to-fine, two-stage parsing strategy: it first performs global layout analysis on a downsampled view of the page, then runs fine-grained content recognition on native-resolution crops of each detected block. This decoupled design lets a small 1.2B model handle high-resolution documents with low computational overhead while preserving non-body elements such as headers, footers, and page numbers.

The Pro checkpoint was produced purely through data engineering on the same 1.2B architecture: the training corpus was expanded from under 10 million to 65.5 million pages with difficulty-aware sampling, and annotations for complex tables and dense formulas were regenerated with cross-model consistency verification and an iterative judge-and-refine pipeline, followed by three-stage progressive training ending in GRPO format alignment.

Beyond core parsing, MinerU2.5-Pro adds practical capabilities for production document pipelines: image and chart analysis, truncated paragraph merging across blocks, in-table image recognition, and improved handling of rotated, borderless, and partially bordered tables. Mixed-language Chinese-English formulas and long, complex equations are parsed to LaTeX.

Architecture

The model is built on the Qwen2-VL architecture (Qwen2VLForConditionalGeneration) with a vision encoder feeding a compact 1.2B language decoder, and ships in Safetensors format on the Transformers framework. It serves with vLLM out of the box; the publisher recommends vLLM for inference and reports concurrent throughput of about 2 frames per second on a single data-center GPU with the asynchronous engine.

Each parsed page is returned as a list of content blocks with a type (text, table, equation, or image), a normalized bounding box, an optional rotation angle, and the recognized content. The publisher's mineru-vl-utils Python package wraps this protocol: point its http-client backend at the served OpenAI-compatible endpoint and call two-step extract to get structured blocks or converted Markdown for each page image.

Use Cases

MinerU2.5-Pro is designed for document understanding workloads at scale:

  • PDF-to-Markdown conversion for LLM training data and RAG ingestion pipelines
  • Table extraction from complex, rotated, or borderless layouts into HTML
  • Formula recognition to LaTeX for scientific and technical documents
  • Layout analysis with reading-order recovery across multi-column pages
  • Chart, flowchart, and seal recognition within documents

Performance

On OmniDocBench v1.6, MinerU2.5-Pro achieves a state-of-the-art overall score of 95.7, outperforming both specialized OCR models and much larger general-purpose frontier vision-language models, and improving substantially over the original MinerU2.5 baseline score of 92.98. Text recognition reaches an edit distance of 0.036, dense formula parsing scores 97.15 CDM, and table recognition reaches 93.62 TEDS across the benchmark, with leading results reported across five diverse table benchmarks.

Quick Start Guide

Choose a model and click 'Deploy' above to find available GPUs recommended for this model.

Rent your dedicated instance preconfigured with the model you've selected.

Start sending requests to your model instance and getting responses right now.