Blog · Archives and Retrieval · October 2, 2026

The Archive Search That Misses What Was Written

A newspaper page can survive, be scanned, and become searchable while the name a reader needs remains invisible. Evaluating an archive requires testing what searches recover, alongside how accurately software transcribes its characters.

The missing surname

Consider a hypothetical newspaper notice whose only mention of a local resident spells her surname “Marden.” The scan is legible, but optical character recognition records “Mardcn.” A literal search for the correct spelling misses that occurrence. The surrounding report might be almost perfectly transcribed. Its most consequential error occupies a single letter.

The example exposes a mismatch between counting errors and finding evidence. For the reader following this person, the surname carries more value than many correctly recognized words around it. An archive can improve its average transcription quality while leaving that reader's question unanswered. Our argument is that historical search should be evaluated through representative questions as well as representative text.

Search settings matter too. The Library of Congress's January 2025 guidance distinguishes newspaper-title searches from full-text searches and recommends page-level results for the latter. It also recommends proximity searches in many cases where exact phrases miss variations in names or OCR errors. This is practical evidence that the query and interface help determine visibility; transcription quality alone does not describe the whole search process.

Different denominators

Character error rate measures the edits needed to turn recognized text into reference text, divided by the reference character count. Substitutions, insertions, and deletions count. The Thomas, Gaizauskas, and Lu correction study uses this measure. At word level, an error measure instead treats words as the units being compared.

Neither denominator is a set of historical questions. For our proposed search evaluation, the relevant denominator is a set of independently established occurrences: names, dates, or unusual expressions known to appear in sampled images. How many can the specified search recover? A separate check asks how many returned results actually concern the intended person or event. Broadening a query may recover missing material while increasing the time needed to reject irrelevant matches.

Suppose, hypothetically, that two transcriptions contain the same number of substituted letters. One damages common words while preserving the sole surname; the other damages that surname. Their character scores can match while their performance on the name query differs. Conversely, an occurrence missed by exact matching might be found through a neighboring place name. There is no universal conversion from a character score to the chance of answering a question.

A 2022 Finnish newspaper user study adds another distinction. Participants rated clippings displayed with different OCR quality, while retrieval and ranking used the improved index throughout. Better displayed text received higher average usefulness ratings. That supports the value of readability, but it does not measure how many additional documents a better index retrieves. The experiment concerned one newspaper and recruited students and teachers rather than a historian sample.

This separation matters when an archive commissions improvements. A clearer result, a newly discoverable result, and a more faithful transcription are different achievements. All may be valuable. Reporting them separately tells readers what has actually improved.

What correction can establish

There is positive evidence for targeted correction. In their 2024 study, Thomas and colleagues report a 54.51% average character-error-rate reduction for instruction-tuned Llama 2 13B on their test set. The underlying BLN600 corpus contains nineteenth-century British newspaper material, largely London crime reporting. Their examples also include an incorrect replacement of a personal name. The reported gain concerns transcription errors, not name-search recall.

A different 2024 study by Boros and colleagues tested fourteen foundation language models across varied historical transcription benchmarks. In its zero- and few-shot prompting settings, outputs mostly degraded the input against reference transcriptions. The authors documented paraphrasing and hallucination among the problems. Their limitations include single-pass generation and possible imperfections in reference texts and alignment. These are results for those experimental settings, not a verdict on every later correction system.

Our interpretation is that correction deserves evaluation at the point where it will be used. These studies differ in training, material, and setup; they are not a controlled contest between two universal approaches. A model tuned for a particular collection may help substantially. An instruction to clean up old text does not by itself establish fidelity.

For historical work, fluency can even make inspection harder. A visibly broken name invites doubt; a plausible replacement may pass unnoticed. This is a reason to test whether corrections preserve already accurate names and dates, alongside whether they repair broken ones. A readable reconstruction should not silently acquire the authority of the printed page.

Start the sample with images

The following is our proposed evaluation method, not a protocol validated by the cited studies. Start with pages selected independently of keyword results. Sample across the collection's relevant differences: titles, periods, languages, print conditions, and page layouts. Searching first and checking only successful hits would leave the missing occurrences outside the test.

Read the sampled images and record query-bearing details with their locations. Include personal and place names, printed dates, uncommon vocabulary, and some ordinary terms for comparison. Mark illegible material as uncertain rather than forcing a transcription. Keep publication-date metadata separate from dates printed inside articles; a filter on the former does not test recognition of the latter.

For each established occurrence, try a documented query sequence: the literal term, a reasonable spelling variant, and a contextual search using neighboring information. Record whether the correct page appears and whether it falls within a predefined amount of result inspection. A page buried beyond what readers can examine is technically retrieved but practically difficult to find.

Compare the original and corrected indexes under the same search settings. Record newly recovered occurrences, previously findable occurrences that disappear, and misleading new matches. Keeping those outcomes separate prevents a net improvement from concealing damage to a small but important group of queries.

This sample cannot establish the complete contents of an archive or its overall recall. It can reveal where a specific search strategy fails against evidence already verified in images. Publish the sampling boundaries with the result. A test concentrated on clear front pages should not become a claim about damaged advertisements or unfamiliar typefaces elsewhere.

Keep the route to the facsimile

We recommend retaining the original OCR, the corrected version, and a stable link to the page image, with the passage location wherever possible. Record which text version was indexed and when a search was run. These are practical proposals for making a finding inspectable and for diagnosing why a later search differs.

The image link serves discovery as well as verification. A reader who arrives through a related term can inspect the surrounding notice and discover a name the transcription missed. Correction should expand these routes into the collection while preserving the evidence needed to challenge a confident-looking result.

A failed query therefore supports a bounded statement: this search, with these settings, found no matching result in the indexed material. Moving from that observation to a claim about historical absence requires evidence about coverage and retrieval. The archive's responsibility is to help readers investigate that gap; the researcher's responsibility is to keep it visible in the claim.

Sources

Sources consulted October 2, 2026. Historical experiments are cited within their tested settings.

Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.


Return to Blog