baidu/Qianfan-OCR
door Baidu
Baidu/Qianfan-OCR is a 4B-parameter end-to-end multimodal model for document intelligence, unifying OCR, layout analysis, and document understanding in a single vision-language architecture. It directly converts images to Markdown and supports tasks like table extraction, chart understanding, key information extraction (KIE), and multilingual OCR (192 languages). Powered by Qianfan-ViT (vision encoder) + Qwen3-4B (language model), it achieves #1 rankings on OmniDocBench v1.5 (93.12), OlmOCR Bench (79.8), and KIE benchmarks (87.9). The model introduces Layout-as-Thought, an optional thinking phase (⟨think⟩ tokens) for structured layout recovery, and delivers high throughput (1.024 PPS on A100 with W8A8 quantization). It is open-source (Apache 2.0 License) and deployable via transformers or vLLM.
Transparantiescore
Gewogen over vier pijlers · bijgewerkt June 30, 2026
Functionaliteiten
Leveranciersinformatie
Volledige informatie over de leverancier/provider van deze AI-applicatie
Leveranciersinformatie volgens de EU AI Act
Krijg inzicht in risico's door beoordelingen uit te voeren op deze AI-applicatie.
Werk je bij Baidu? Claim deze vermelding om de gegevens te corrigeren of aan te vullen.
EU-alternatieven
Ontdek EU-gebaseerde alternatieven voor deze AI-applicatie.
Klaar om AI-applicaties te beheren?
Volg, beoordeel en beheer je AI-applicaties met Anove.