Chloë Farr with Jessica Jack and Jacob Polay
This submit is a part of a collection on AI and Collaboration.
For many individuals who use archives, the accessibility of archival materials is a notable issue. Archivists and librarians have lengthy been turning to software program to assist with that accessibility. One in all their primary instruments is Optical Character Recognition (OCR). One of these software program turns pictures of textual content into digital characters, which may then be accessed by anybody with an web connection and will be simply processed and analyzed exterior of an archive. However OCR began as a company resolution to company issues, which meant it was not superb at coping with the variations current in archival supplies. Archivists have been attempting to deal with this with bespoke expertise, however the previous few years have introduced a brand new method to cope with the issue.
OCR enabled by AI imaginative and prescient language fashions (VLM-OCR) has lately exploded as a analysis matter in a number of fields together with pc science, digital humanities, library info science, linguistics, and past. This new motion in OCR started as a result of AI builders hit a wall. That they had educated their fashions on all accessible information on the internet so that they tried utilizing AI-generated information as an alternative, however the outcomes have been problematic, inconsistent, and dangerous. To resolve this, they turned to bodily paperwork to broaden their coaching base. Nonetheless, they wanted higher OCR to make this attainable. The paper by AllenAI that opened the sector, “olmOCR: Unlocking Trillions of Tokens in PDFs with Imaginative and prescient Language Fashions,” says as a lot in its opening traces: PDFs maintain monumental volumes of “novel, high-quality” coaching information, however their variety of codecs and layouts makes that content material laborious to extract faithfully.
AllenAI’s assertion highlights a few of the motivations behind this expertise. Many AI firms at the moment are growing OCR fashions and releasing a few of them brazenly with a view to encourage adoption, construct person communities, and set up their instruments as a part of the broader document-processing ecosystem. Open supply fashions—whose supply code is publicly accessible, freely licensed to make use of, modify, and redistribute—can cut back the price of entry to high-quality OCR. They allow galleries, libraries, archives, and museums (GLAM) establishments and researchers to run, consider, and adapt instruments on their very own infrastructure. Nonetheless, this doesn’t remove industrial competitors: firms can nonetheless cost for hosted companies, enterprise assist, specialised fine-tuning, and licences for larger-scale industrial use. For instance, Datalab’s Surya makes its software accessible underneath a semi-open license. The flexibility to fine-tune the software’s coaching is free for analysis, private use, and smaller startups however they require broader industrial licensing for different customers. Whereas helpful in some respects, Surya shouldn’t be made for GLAM and humanities analysis and thus represents a restricted use case of those semi-open company options.
This sort of for-profit mannequin system may also be seen in Transkribus, some of the broadly used OCR companies for arts analysis. This software program shouldn’t be totally open supply, as an alternative working on a credits-based system the place the primary 50 credit are free after which the fee will increase. Prospects can fine-tune or “prepare” their very own fashions for improved transcription for his or her collections. Usually fine-tuning is most useful on homogenous collections with a excessive quantity of paperwork. Nonetheless, the corporate retains the underlying mannequin for themselves, together with the shoppers’ paperwork used for that improvement. This restricts the portability of the mannequin exterior of the Transkribus setting. Alternatively, Adobe PDF reader is a broadly accessible OCR engine requiring little technical information, however is on the market with a subscription price, and performs poorly on historic paperwork. Whereas these are however two examples, the expansive use of those fashions display the demand for OCR, but in addition the simultaneous want for this OCR to turn out to be extra accessible and sustainable. It’s in GLAM’s curiosity to seek out methods to make use of OCR fashions with out counting on for-profit companies as open supply fashions guarantee possession of information whereas minimizing overhead bills for constantly underfunded establishments. This independence will be achieved via capacity- and knowledge-sharing, and collaborating on OCR processing to take away redundant work.
The outputs of many OCR applied sciences have been developed for coaching LLMs, however which are ceaselessly of little use within the humanities. They can be utilized by individuals with particular information mining and information science functions, like researchers doing Named Entity Recognition. However these are nonetheless quite area of interest, and the vast majority of archival researchers and historians who wish to entry these OCR outputs would as an alternative profit from the doc processing of OCR leading to searchable archives. That is one other space the place AI-driven OCR is useful, because it has the capability to simply remodel OCR outputs into particular and bespoke codecs for the wants of the researchers who’re utilizing these outputs.
In my work as a GLAM researcher located in libraries, I give attention to making VLM-OCR accessible and helpful for analysis and archival use. I’m ceaselessly requested “What’s one of the best VLM-OCR mannequin proper now?” Earlier than March 2026, it was often fairly clear. The fashions have been fairly uniform, dealing with the identical sort of paperwork, offering the identical output codecs however with totally different ranges of transcription accuracy primarily based on the supply doc’s language, scan high quality, and textual content structure. Now, every mannequin has their very own distinct strengths. It’s a welcome improvement that fashions are now not competing on accuracy alone. For instance, Hunyuan OCR offers coordinates for every phrase on the web page, which in flip allows individuals to seek for the phrase and see it highlighted proper on the web page. ChandraOCR-2 additionally does an excellent job of transcription, and it may well detect pictures inside a doc (images, artwork, graphics, and many others.) and hold them separate from the encompassing textual content, together with writing quick descriptions of what’s in every one. Surya OCR is a really small mannequin, which means it may well run on lower-quality {hardware}, and transcribes at the next pace whereas often sacrificing accuracy. Differentiating by strengths eases the burden on customers, who now not should chase a 0.1% accuracy edge and may as an alternative decide whichever mannequin fits their functions. Common customers appear to develop an instinct for this. It comes with expertise, gained by testing totally different fashions throughout a spread of doc sorts and matching them to what the person wants from a transcription. On this sense, collaboration between these skilled researchers within the house is essential to serving to everybody entry the fashions that greatest swimsuit their wants.
What emerges from this shift is an ecosystem of more and more complementary AI-enabled OCR fashions, whose actual worth is dependent upon the researchers who know the best way to use them. With these fashions, establishments don’t have to guess the whole lot on one firm’s roadmap or pricing mannequin. They’ll now run the software program on native {hardware} that ensures information possession stays with the establishment. And since the instruments are open, GLAM professionals can pool their experience constructed via hands-on testing, matching fashions to supplies, and sharing their work throughout establishments. The fashions themselves at the moment are adequate that additional good points will likely be small and specialised. What wants bettering is how we use them collectively. The individuals making these paperwork accessible want a shared and rising toolkit, constructed with and for one another to make use of, made simpler by the assist of LLMs. For chronically underfunded archives and libraries, that collaboration is value as a lot as any accuracy acquire. Higher OCR is value having. Constructing the capability of archives and libraries to serve the individuals who depend on them is value extra.
Chloë Farr is a researcher working on the intersection of synthetic intelligence, archives, and digital humanities. Figuring out of the Open Science Lab at TIB – Leibniz Info Centre for Science and Know-how, her analysis focuses on large-scale textual content recognition and evaluation of historic paperwork, together with newspapers, maps, and archival information. Be taught extra about Farr’s work on GitHub.
Jessica Jack is a PhD scholar in Historical past on the College of Saskatchewan, growing purposes for Massive Language Fashions in historic analysis. They’re doing so via learning settler land use in late nineteenth century and early twentieth century Saskatchewan.
Jacob Polay is a PhD scholar in Historical past on the College of Saskatchewan, learning the roles Massive Language Fashions have within the historic methodology. His present analysis entails creating an info retrieval pipeline utilizing synthetic intelligence instruments to unlock the early trendy archive at scale.
Associated



