NVIDIA released a specialized vision-language model that combines OCR, layout understanding, table extraction, and reading order in one efficient architecture. It significantly outperforms traditional OCR on complex, real-world documents while preserving structure and semantics.