- Client
- DGLAB, Portugal's Directorate-General for Books, Archives and Libraries
- Sector
- Public administration · Cultural heritage
- Scope
- AI models, data migration, front office and back office, delivered by Caixa Mágica
- AI
- Named-entity recognition and disambiguation, trained on the Portuguese state supercomputing network
- Data model
- Graph database built on the CIDOC-CRM international standard
- Reach
- 20 archives searchable in a single query, including the National Archive of Torre do Tombo and the district archives
- Scale
- 8M+ descriptive records · 63M+ digitised images · 100+ km of physical documentation
- Integration
- CRAV, the online consultation, reproduction and certificate service
- Status
- In production at digitarq.arquivos.pt
Caixa Mágica Software rebuilt DIGITARQ, the platform through which Portugal opens its documentary memory to the world, for DGLAB, the Directorate-General for Books, Archives and Libraries. The new generation of the platform pairs a modern, fully web-based front and back office with artificial intelligence models trained on the Portuguese state's supercomputing network. Those models read archival descriptions, recognise the people, institutions and places named in them, and weave millions of records into a connected knowledge graph built on CIDOC-CRM, the international standard for cultural heritage information.
The result is an archive that can be explored the way historians, genealogists, journalists and citizens actually think: by person, by place and by relationship, rather than only by catalogue reference. It draws on more than two decades of Caixa Mágica experience in public sector platforms, open standards and applied machine learning.
In production today: 20 archives in a single search · 8M+ documents · 10M+ actors (people and collective bodies) · 16M+ events · 180k+ places · 63M+ digitised images. Current production figures from the CIDOC-CRM knowledge graph.
Challenge
Portugal's archives are among the richest in Europe. Through DGLAB and the National Archive of Torre do Tombo, together with the district archives, the country safeguards more than 100 kilometres of documentation. That means over eight million descriptive records and more than sixty-three million digitised images reaching back many centuries. DIGITARQ is the door to that memory. When researchers, students and citizens want to find a baptism record, a notarial deed, a colonial-era document or a family name, DIGITARQ is where the search begins.
The previous generation of the platform had grown over many years around a set of regional, relational databases. Each archive described its holdings in its own silo. The front office was regional too: to search across the country, a user had to visit a separate website for each archive and repeat the same query on every one of them. Search itself matched strings of text. Type a name and you retrieved the records where that exact string appeared, with no understanding that "Torre do Tombo", "ANTT" and "Arquivo Nacional" might refer to the same institution, or that a person named in a will might be the same person baptised in a parish register held in a different archive.
The Core Problem
Archival value lies in connections between people, places, institutions and events. Yet the data was stored as isolated records and searched as plain text. The challenge was to make those connections legible to a machine, and therefore discoverable by a human, across millions of documents and dozens of archives.
DGLAB set out to replace this fragmented landscape with a single platform, fully web-based, more interoperable, easier to use and measurably faster. That ambition carried three hard constraints. Scale came first: any new system had to serve millions of records and tens of millions of images without degrading the search experience. Continuity came second, because decades of carefully catalogued description could not be lost or distorted in the migration. Longevity came third. As a public cultural institution, DGLAB needed a platform built on open, international standards it could sustain and evolve for decades, free from lock-in to any single supplier.
Layered on top was the most demanding requirement of all: to use artificial intelligence responsibly and sovereignly. The entities hidden inside the archives, the individual actors, the collective bodies, the places, had to be surfaced automatically, because no team of archivists could ever tag millions of records by hand. The AI doing that work also had to run on national infrastructure, keeping sensitive cultural heritage data within the control of the Portuguese state.
Solution
Caixa Mágica delivered a new DIGITARQ built as a fully web-based platform on open technologies, structured around three pillars. An AI layer reads and understands archival descriptions. A knowledge graph connects everything it finds. A modern front and back office puts that work in the hands of citizens and archivists alike. The system launched as the new generation of the national platform in 2025.

A single sign-in connects DIGITARQ to CRAV, either with a platform account or through Autenticação.Gov.
Key components of the new DIGITARQ platform:
An AI Layer Trained on National Supercomputing
The foundation of the new experience is a set of language models trained to perform named-entity recognition across Portuguese archival metadata. Given a descriptive record, the models identify and label the meaningful entities inside it: the names of individuals, of collective and institutional actors, and of places. Historical Portuguese is irregular, heavily abbreviated and spelled inconsistently across centuries, so generic off-the-shelf tools are not enough. The models were adapted and evaluated specifically for this domain, comparing modern transformer architectures to select the best performer for archival language.
Recognising a name is only half the task. The same string can refer to many different people, and the same person can be written a dozen ways. Named-entity disambiguation resolves each mention to a single, unambiguous entity, so every reference to a given individual, body or place connects to one canonical node. This is what allows DIGITARQ to answer a question about a person, instead of merely listing the documents where a sequence of characters happens to appear.
Training these models, then applying them across millions of records, is computationally intensive. Caixa Mágica carried out the model training on the Portuguese state's advanced computing and supercomputing network. Beyond the performance this made possible, it kept the processing of the nation's cultural heritage on sovereign, public infrastructure, a deliberate choice aligned with European data residency and digital sovereignty principles. Recognition and disambiguation ran as a batch process over the entire migrated dataset. The entities were extracted and resolved once, across all the data, and now populate the platform. The models are not invoked in real time as users browse.
A CIDOC-CRM Knowledge Graph
In parallel, Caixa Mágica migrated the legacy regional and relational databases into a graph database modelled on CIDOC-CRM, the ISO-standard conceptual reference model for documenting cultural heritage. Where the old model stored rows in tables, the new model stores entities and the relationships between them: this person created that document, that document concerns this place, this institution held that record at a given time. The entities surfaced by the AI layer flow into the graph, turning a collection of separate catalogues into one connected web of Portuguese history.
On top of the graph, a modern indexing and search engine delivers fast, relevant results across the entire holdings. A user can begin with a name, a place or a document and travel outward along real relationships, discovering related actors, associated locations and connected records that a keyword search would never have surfaced. The scale is considerable: the platform indexes millions of descriptive contents and tens of millions of images, and stays responsive throughout.
A Modern Front and Back Office
The public front office is where researchers, genealogists, students and curious citizens meet the archive. The old regional front offices sent users from site to site, repeating the search at each archive. A single search now covers all twenty participating archives at once. It offers clean, fast, fully web-based search and browsing across the national collection, rich entity pages, and direct pathways from a discovered record to the services that let a user act on it. The interface was designed for a broad, non-specialist audience while still rewarding the depth professional researchers demand.
Behind the scenes, a dedicated back office gives DGLAB's archivists the tools to describe, curate and manage holdings, and to maintain the quality of the catalogue and the knowledge graph over time. The archivist remains firmly in control of description and curation, working on the same connected data the batch entity processing produced.
Discovery only delivers its full value when it leads to action. DIGITARQ is integrated with CRAV (Consulta Real em Ambiente Virtual), the archives' online service desk for requesting consultations, reproductions and certificates. A citizen can move directly from finding a document in DIGITARQ to requesting it through CRAV, so search and service form one continuous digital process.
Impact
The new DIGITARQ changes what it means to consult a national archive. What was once a text search over isolated catalogues is now a path through a connected map of people, institutions and places, accessible from anywhere, at any time, to anyone with a question about Portugal's past.
For researchers, the practical gain is reach. Work that once required visiting several archives and cross-referencing catalogues by hand can now begin online, and often finish there. For DGLAB, the gain is structural: a single national platform, built on open standards, that the institution can sustain and extend without depending on one supplier.

Over 100 km of documentation is now one search away, with no queues and no travel.
Some notable outcomes of the new DIGITARQ:
By recognising and disambiguating entities and connecting them in a knowledge graph, the platform lets users find records they would previously have missed. Search is no longer limited to matching the exact words a cataloguer happened to write.
The platform brings more than eight million descriptive contents and over sixty-three million images into a single, responsive, fully web-based experience, giving structured access to a documentary heritage measured in hundreds of kilometres of shelving.
Entity recognition and disambiguation were applied across the whole migrated dataset, enriching millions of records with structured people, institutions and places at a scale manual cataloguing could never reach. The CIDOC-CRM knowledge graph now holds more than 10.8 million actors, over 16 million events, 189,000 places and 8 million documents, all interconnected.
Built on open technologies and the CIDOC-CRM international standard, and trained on national supercomputing infrastructure, the platform gives DGLAB a durable, interoperable foundation under public control, which it can sustain and evolve over the long term.
Integration with CRAV connects finding a document to requesting it, reducing friction for citizens and lightening the load on reading room and reproduction services.
Caixa Mágica Software, working with DGLAB, replaced a fragmented set of regional catalogues with one national platform that reads its own holdings. The project met the scale, continuity and longevity requirements DGLAB set out, and it did so on open standards and on public infrastructure, keeping the processing of Portuguese cultural heritage inside the country. Twenty archives, eight million records and sixty-three million images now answer a single query. The archivist keeps control of description and curation, while the AI handles the volume no team could ever cover by hand.
Looking Ahead
An archive is never finished, and neither is the platform that opens it. Already live in production, DIGITARQ keeps growing, with thousands of new images and a steady stream of new records added over time. The knowledge graph will expand as more holdings are described and more entities are recognised and linked. The AI models can be retrained and refined as the collections, and the language used to describe them, are better understood.
Semantic search, entity linking to external authorities and closer integration between discovery and services all point to a future in which Portugal's documentary memory becomes progressively easier to explore. Built on open standards and a modular architecture, the platform is designed to evolve with the needs of its users and the pace of technological change, an evolution DGLAB can steer with the partners it selects over time.
Frequently Asked Questions
What is DIGITARQ?
DIGITARQ is the online platform of DGLAB, Portugal's Directorate-General for Books, Archives and Libraries. It lets anyone search and consult the catalogues and digitised holdings of the national and district archives, spanning millions of descriptive records and tens of millions of images. A single search now queries all 20 archives at once, replacing the old regional front offices that forced users to visit each archive's own website separately.
How many archives can I search at once?
DIGITARQ covers 20 archives in a single query, including the National Archive of Torre do Tombo, the district archives and other national archives. Previously each archive had to be searched on its own website.
How does DIGITARQ use artificial intelligence?
Caixa Mágica trained AI models for named-entity recognition (NER) and named-entity disambiguation (NED) on the Portuguese state supercomputing network. The models identify people, organisations and places inside archival descriptions and link each mention to a single, unambiguous entity, which supports richer search and discovery.
What is the CIDOC-CRM knowledge graph used for?
Regional and relational databases were migrated into a graph database modelled on CIDOC-CRM, the international standard for cultural heritage information. Isolated records became a connected web of people, institutions, places and documents that can be explored by relationship rather than by keyword alone. The graph holds tens of millions of interconnected nodes: over 10.8 million actors, 16 million events, 189,000 places and 8 million documents.
More About DIGITARQ
How does DIGITARQ relate to CRAV?
DIGITARQ is the discovery layer where users find and explore records. CRAV (Consulta Real em Ambiente Virtual) is the online service desk for requesting consultations, reproductions and certificates. The two platforms are integrated, so a citizen can move from finding a document to requesting it without leaving the archives' digital services.
How large is the collection behind DIGITARQ?
The platform provides access to more than 8 million descriptive contents and over 63 million associated images, representing more than 100 kilometres of physical documentation held across Portugal's archives network.
Who developed DIGITARQ and for whom?
The platform was developed by Caixa Mágica Software for DGLAB, the Portuguese Directorate-General for Books, Archives and Libraries, which coordinates the National Archive of Torre do Tombo and the district archives.
Is the platform built on open standards?
Yes. DIGITARQ is a fully web-based system built on open technologies and on international archival and cultural heritage standards, including CIDOC-CRM. That choice supports interoperability, long-term sustainability and independence from any single vendor.
