Skip to content
Help articles
  1. 1. Getting Started
  2. 2. Workspaces & Teams
  3. 3. Uploading Documents
  4. 4. Processing & Outputs
  5. 5. Reviewing in the Workbench
  6. 6. Curating the Batch Roll-up
  7. 7. Delivering & the ArchivesSpace Integration
  8. 8. Exporting Your Results
  9. 9. Credits & Billing
  10. 10. Privacy & Data Handling
  11. 11. Trying Acervo without an account

Processing workflow

Uploading Documents

Prepare files, choose page order and models, and avoid common upload problems.

On this page

Section links appear after the article loads.

Uploading Documents

Uploading is where every Acervo job begins. What you upload, which model you choose, and how you group a large collection all shape what comes back. Better inputs lead to cleaner OCR, richer entities, and less time spent reprocessing.

This section covers what to upload and how to set yourself up for the cleanest possible output.

While your workspace still holds the sample folder, the Upload page links to it beneath Browse documents, so you can read a finished batch before preparing your own.

Supported file types

Acervo accepts three formats. The same per-page rate applies to all three. 1 credit per page on Standard, 5 credits per page on Premium. A single-page JPG, a single-page TIFF, and one page of a PDF all count as one page for credit purposes.

  • JPG · the most common format for digitized archival material. Good for typical photographs of documents, where file size matters.
  • TIFF · preferred for high-resolution archival masters. Larger files, but no compression artifacts. If your institution archives in TIFF, upload TIFF.
  • PDF · multi-page or single-page scans. A 12-page PDF counts as 12 pages of processing.

Per-file size limit. Each file you upload must be under 300 MB. Most JPG scans and typical PDFs are well below this; very high-resolution multi-page TIFFs or large multi-page PDFs are the most likely to bump against the cap. If a file is too large, downsample or split it before upload.

Image quality tips

Acervo's models perform much better when the source image is good. A few rules of thumb

  • If your scans are low resolution, plan to use Premium. Standard performs best on clear, well-resolved images. For low-DPI scans, faded ink, or otherwise difficult source material, Premium is the better fit and will save you reprocessing time.
  • Color or grayscale is better than thresholded black-and-white. Bitonal scans lose tonal information that helps OCR distinguish faded ink from background.
  • Avoid skewed or rotated scans. Acervo handles slight skew, but heavily rotated pages will hurt OCR accuracy. Straighten before uploading where possible.
  • Crop tightly if you can. Wide margins of empty page slow processing and don't add value. Tight crops also reduce file size.
  • Don't upload thumbnails. Low-resolution preview images won't produce useful output. If it's hard for a human to read, it'll be hard for Acervo too.

If you're unsure whether an image is good enough, run a single page through first. If the OCR confidence is low or the entities are sparse, the source image is likely the bottleneck.

How many pages per job

Acervo runs jobs of up to 100 pages each. There's no hard cap on how many jobs you can run; only the per-job limit.

For larger collections

  • Split logically, not arbitrarily. Group pages so each job represents an intellectual unit (a single letter, a series of related correspondence, a section of a report). This keeps the entity extraction and scope and content notes coherent within each job.
  • Avoid cross-cutting jobs. Don't put the last 30 pages of one folder and the first 20 of another in the same job. The aggregated output will mix them in ways that are hard to untangle on review.
  • Order matters. Acervo processes pages in upload order and uses cross-page context for some outputs (entity deduplication, multi-page scope and content). Reorder before processing so the sequence reflects the document's natural reading order.

Choosing between Standard and Premium

Two models, two prices, two strengths

  • Standard (1 credit per page) · typed and printed text, legible handwriting, multilingual material across most major scripts, multi-column layouts when they belong to one intellectual unit (a printed report or pamphlet. Not a newspaper page where each column is its own article).
  • Premium (5 credits per page) · a stronger model that reads with more nuance. Faded or heavily degraded handwriting, unusual or archaic scripts, low-contrast source material, and subtler judgment calls on ambiguous text.
  • Optional Transkribus OCR (+1 credit per page) · runs the public Transkribus model you explicitly select before Acervo continues with the independent Standard or Premium processing pipeline.

If you're unsure, start with Standard. Acervo will flag pages it couldn't read confidently, so you can decide whether to rerun those pages at Premium.

What Acervo handles well

The Getting Started page introduced this at a high level. Here's the detail.

Standard handles cleanly

  • Printed books, pamphlets, reports, ephemera · typeface OCR is the model's strongest case
  • Modern correspondence · typed letters, memoranda, business correspondence
  • Legible handwriting · neat, unhurried script in modern hands; copybooks; carefully prepared manuscripts
  • Multilingual material · Latin, Cyrillic, Greek, and most major scripts; the model is comfortable with code-switching within a single page
  • Multi-column layouts that are one intellectual unit · a two-column report, a three-column pamphlet, an annotated index. The model preserves reading order within the unit.

Premium handles harder cases

  • Faded or degraded ink · historical iron-gall ink that has bled, oxidized, or faded into the paper
  • Unusual or archaic scripts · Gothic Fraktur, blackletter, secretary hand, abbreviated medieval scripts
  • Low-contrast material · photocopies, carbon copies, faint pencil annotations, water-damaged sheets
  • Difficult correspondence · heavily marked-up drafts, telegrams (often degraded), scribbled marginalia
  • Articles and contracts with dense legal or technical language · where the model needs more capacity to handle dense or specialized text

What Acervo doesn't do well

  • Photographs and other primarily visual material. Acervo will produce a basic title and scope and content note from any visible text or recognizable elements, but entity extraction, authority matches, and subject headings will be limited or absent. For photograph-heavy collections, treat Acervo's output as a starting point and expect more cataloger work per item.
  • Newspapers where columns are independent articles. Acervo treats a page as one intellectual unit. If your page is actually six unrelated stories, the output will conflate them. Process clipped articles individually instead.
  • Heavily redacted or partially destroyed pages. OCR can only read what's there. Acervo won't fabricate redacted content or guess at missing pieces.

Institution profile and batch facts

The institution profile carries your repository identity. Publisher, repository code, country code, and language of description. It rides every export and send, so your records always say who they came from, and it rarely changes once set.

Everything that can change from batch to batch is declared with the batch instead, on the upload card and the About-this-batch strip. A profile that hard-codes a collection goes stale the moment you process a different one, so the batch carries its own facts.

Profile fields, your identity

  • Publisher · the institution or organization responsible for the records
  • Repository Code · your repository's standard short code
  • Country Code · two-letter country code (e.g., US, MX, GB)
  • Language of description · the language your finding aids are written in, one or several. Left empty, exports and sends record English. This declares the language of the description, not of the materials; the documents' own languages are read from the documents.

Batch facts, declared at upload

  • Collection · the collection this batch belongs to, chosen from your workspace or created inline
  • Rights · a controlled rights status plus an optional statement, applied to every document in the batch
  • Donor · the donor or source of the materials
  • About these materials · a scope note in your words that grounds narrative drafts and rides the export manifest
  • Geographic coverage · what the material is about, never where it was sent or where your repository sits. Left blank it derives from each document's own evidence, and the manifest notes whether coverage came from the documents or from you

Batch facts apply to every document in the batch with no per-item override, which is also guidance for composing batches. Materials with different rights or donors belong in separate batches.

If you're working in a team workspace, institution profiles and collections are shared across the team. See §2 Workspaces & Teams.

Preview and reorder before processing

Before you click Start processing, you can

  • Preview thumbnails of each uploaded page
  • Reorder pages by drag-and-drop so the sequence reflects the document's intended reading order
  • Remove pages that don't belong (a stray cover sheet, a duplicate scan, a blank back side)
  • Confirm your model choice (Standard vs Premium) before any credits are charged

Once processing starts, the page set and order are locked for that job. To change them, you'd need to start a new job.

Where to go next