Mistral AI released a major update to its document intelligence model that helps users comprehend all elements in a document, such as media, text, tables and equations, rather than just extracting those data points.
Launched on June 23, the model can process images and PDFs and extract content as ordered, interleaved text and images. It is designed for retrieval-augmented generation (RAG) systems. The model supports 170 languages across 10 language groups, according to Mistral. It can also process up to 2,000 pages per minute on a single GPU and can be used as a standalone for context extraction.
With OCR 4, Mistral is tackling one of the main challenges enterprises face in OCR (optical character recognition) and document intelligence: instead of just extracting the document, it also aims to understand it.
“What OCR typically does is just pull stuff out,” said Mark Beccue, an analyst at Omdia, a division of Informa TechTarget. “It does not really look at it and understand it. Whereas this is saying we understand it.”
OCR 4 comes 15 months after the Paris-based AI lab introduced its original OCR using its API. It later introduced OCR 3 in December.
With the fourth-generation OCR 4, Mistral provides a model that creates machine-readable, actionable data and gives enterprises data they can work with and prompt, Beccue said.
Unstructured Data
OCR 4 also follows a trend in the AI market: vendors are now focusing on tools and models to help organize unstructured data. Unstructured data is information that cannot be organized into a standard spreadsheet. According to Gartner, 80% of data is unstructured.
“All we’re dealing with before is structured data for all these models,” Beccue said. “That’s leaving a lot on the table. If you cannot index this stuff very easily, it breaks pipelines, data pipelines.”
“It’s a missed opportunity,” he continued. “The large language models struggle with the raw documents. They cannot figure it out.”
He added that while older OCR applications can extract information, it is hard to interpret what they pull out.
Bounding Boxes
And while other vendors, including Google and Microsoft, offer OCR applications such as Google Document AI and Azure AI Document Intelligence, Mistral OCR 4 stands out for its bounding boxes feature. Bounding boxes enables users to localize text, highlight it and draw boxes over it in the document so that it is clear where the information or data was extracted from. When combined with retrieval-augmented generation (RAG), the feature enables AI assistants to provide clickable citations that link directly to the data’s source.
“This is a pretty human-intensive thing for people to do right now,” Beccue said. “When you are looking at unstructured data, there is a lot that is not very automated. So, there is a breakthrough.”
OCR 4 is integrated with the Mistral Search toolkit, an open source, composable search framework, now in public preview. Those who use the model with the API can access it for $4 per 1,000 pages. Teams can use it in Document AI in Mistral Studio, which is priced at $5 per 1,000 pages.

