Skip to content
Help articles
  1. 1. Getting Started
  2. 2. Workspaces & Teams
  3. 3. Uploading Documents
  4. 4. Processing & Outputs
  5. 5. Reviewing in the Workbench
  6. 6. Curating the Batch Roll-up
  7. 7. Delivering & the ArchivesSpace Integration
  8. 8. Exporting Your Results
  9. 9. Credits & Billing
  10. 10. Privacy & Data Handling
  11. 11. Trying Acervo without an account

Start here

Getting Started

Understand Acervo's workflow, model fit, first job steps, and review expectations.

On this page

Section links appear after the article loads.

Getting Started with Acervo

Acervo turns digitized archival images into item-level records ready for your catalog. For every page or item you upload, Acervo produces a title, a scope and content note, named entities (people, organizations, places, works, events), authority matches against Library of Congress, VIAF, GeoNames, and Wikidata, FAST/OCLC and LCSH subject headings, and searchable OCR text. With searchable OCR text of every page.

Acervo doesn't replace cataloger judgment; it accelerates the descriptive work and surfaces what needs human review. Confidence scores on entities and subject headings tell you where to focus, and named entities link out to their authority files for direct verification.

Who Acervo is for

Acervo is built for catalogers, archivists, and the broader teams that support them at digitization-heavy institutions. Digital Library of the Caribbean (dLOC) partners, university libraries, special collections, community archives, and historical societies. You don't need to be a cataloger to use it. But the output is meant to be reviewed by someone who understands DACS, ISAD(G), or your institution's own descriptive standards before it lands in a public catalog.

What a job looks like, end to end

Every job moves through four steps

  1. Upload · You upload up to 100 pages per job (JPG, TIFF, or PDF). Larger collections are split across multiple jobs. You can preview and reorder pages before processing starts.
  2. Process · Acervo extracts OCR text, identifies named entities (people, organizations, places, works, events), matches them against authority files, generates titles and scope and content, applies FAST/OCLC and LCSH subject headings, and assigns confidence scores to every output. You can leave the page; jobs continue running on the server and you'll see progress when you return.
  3. Review & curate · Open the workbench to review and curate the output. Triage low-confidence claims, edit titles and descriptions in place, select access points in the batch roll-up, and resolve sensitivity flags. What you choose there is exactly what delivers.
  4. Deliver · Download in the format your catalog or repository accepts. CSV, Dublin Core, EAD3, EAD2002, MODS, or raw OCR text. Individually, or bundled as a configurable ZIP. If you run ArchivesSpace, send curated batches directly through its standard API instead. Nothing to install, and re-sends update records in place. For dLOC partners, Acervo generates the dLOC-ready CSV; you submit it to dLOC for ingestion.

Past jobs and their outputs stay accessible in your job history. Reprocessing is always user-initiated. Acervo never reruns a job on its own. If you want to reprocess (for example, after a vocabulary update or model improvement), you'll re-upload the source documents and start a fresh job, since Acervo deletes the originals after each run.

What Acervo handles well

Acervo's models are tuned for the descriptive material you'd typically find in a digitized archive. Text-bearing documents that need to become discoverable item-level records. Knowing what each model is built for will save you credits and reprocessing time.

The default (Standard) model handles

  • Printed text
  • Multi-column layouts when they belong to a single intellectual unit (e.g., a printed report or pamphlet. Not a newspaper page where each column is its own article)
  • Multilingual material across most major scripts
  • Legible handwriting

The Premium model is built for harder cases

  • Faded, heavily degraded, or hard-to-read handwriting
  • Unusual or archaic scripts
  • Low-contrast source material
  • Subtler judgment calls on ambiguous text in letters, telegrams, contracts, and articles

What Acervo doesn't do well

  • Photographs and other primarily visual material. Acervo will produce a basic title and scope and content note from any visible text or recognizable elements, but entity extraction, authority matches, and subject headings will be limited or absent. For photograph-heavy collections, treat Acervo's output as a starting point rather than a complete record. And expect more cataloger work per item.

Personal vs. team workspaces

Acervo supports two kinds of workspaces

  • Personal · Owned by you alone. Your credits, your jobs, your institution profiles.
  • Team · Shared across your institution or working group. Team credits are pooled. Jobs and exports belong to the team. Roles (owner, admin, billing, member) control who can invite, purchase credits, manage billing, or process documents.

If you're working alone or evaluating Acervo, your personal workspace is the right starting point. If your institution will use Acervo across multiple staff members, set up a team workspace early so credit accounting and history are shared from day one. See §2 Workspaces & Teams for invites, roles, and switching.

The sample folder

Every new personal workspace opens with a sample folder already on the bench. Two documents from the Olga Iglesias Project, read by Acervo and ready to review, so you can walk the Review and Documents tabs before you spend a credit. Open it from the Upload page. The sample folder never leaves Acervo. Exports and sends are reserved for your own documents, and Remove sample folder deletes it, and any curation you saved on it, when you are done.

Your first job, step by step

  1. Sign up at acervo.org.
  2. Get credits. Personal accounts and new teams alike can choose
    • Claim 25 free starter credits. A one-time offer redeemed at $0, valid for 30 days (about 25 standard pages or 5 difficult pages).
    • Or buy a credit bundle outright and skip the trial.
  3. Go to Upload. Drag and drop your files, or click to select. Acervo supports JPG, TIFF, and PDF.
  4. Preview and reorder the pages so they appear in the order they should be described.
  5. Choose a model. Standard handles typed/printed text, legible handwriting, and most multilingual material at 1 credit per page. Premium handles harder cases, faded or degraded handwriting, unusual scripts, low-contrast source material, at 5 credits per page. If you're not sure, start with Standard; Acervo will flag pages it can't read confidently and you can rerun those at Premium.
  6. Start processing. A single document can finish in as little as one minute. Processing time increases with batch size and depends on page complexity, source quality, and whether you've chosen Standard or Premium. You can leave the page and come back. Jobs continue running on the server.
  7. Open the workbench when the job finishes. You'll see titles and scope and content, named entities (each with its confidence score and a hyperlink to its authority file for sanity-check), and subject headings (each with a confidence score and rationale you can read on hover), all in one view. Review, correct, and curate right there. Edits and selections are saved with the batch.
  8. Export. Download individually in CSV, Dublin Core, EAD3, EAD2002, MODS, or raw OCR text. Or bundle them together as a configurable ZIP where you pick which formats to include.
  9. Edit and import. Make any corrections to the exported file in your usual tools, or wait and edit after importing into your catalog or repository. Either path works. For dLOC partners, submit the dLOC-ready CSV to dLOC for ingestion.

A note on review discipline

Acervo is fast, but it is not a replacement for cataloger judgment. Two principles to keep in mind

  • Always review in the workbench before exporting. Confidence scores tell you where to focus; they don't tell you a record is ready. Selection is opt-in on purpose. Nothing reaches your catalog unless a person chose it.
  • Edit where the data lives. Titles, descriptions, and wrong authority identifiers are corrected in the workbench itself, and the fixes ripple into every export and send. Identity spellings, roles, and selection are curated in the batch roll-up. You can still refine exported files in your own tools afterwards, but you no longer have to.

AI-assisted records are a starting point for your judgment, not a substitute for it.

Where to go next