How to digitize a client archive without creating another messy folder
A workflow for agencies and freelancers to turn photographs, documents, and physical material into a searchable, reusable archive.
Founder of Polimake, YouTuber.

A client delivers four boxes of old photographs, catalogues, press cuttings, and contracts. They want “everything digitized” for a brand anniversary.
Scanning the boxes creates thousands of files. Digitizing the archive means deciding what to preserve, how people will find it, and what can legally be published.
Define the output before capture
Ask whether the material will support internal reference, brand research, social content, large-format print, legal records, an exhibition, or a documentary.
The use determines capture quality, format, and metadata. An invoice for reference and a photograph for a large display do not need the same treatment.
Produce a twenty-item sample before processing every box. The client can correct names, categories, and quality while changes are still inexpensive.
Inventory in batches
Do not begin with “scan0001.pdf”. Give every box, folder, or album an identifier and record its source, approximate dates, material type, physical condition, estimated quantity, and known restrictions.
For example: BOX03_Catalogues_1998-2004.
The identifier maintains a connection to the physical source if questions appear later.
Match capture to the material
- Flat documents: scan and apply OCR so text is searchable.
- Photographs: use sufficient resolution and avoid destructive corrections.
- Books or fragile items: use a fixed camera, even lighting, and safe support.
- Small cuttings: include a scale reference when size matters.
Keep a high-quality master and generate a lighter reference copy. Never compress the only digital original.
Name files using known facts
A useful pattern is:
client_approx-date_type_batch_sequence
For example: acme_2002_summer-catalogue_B03_014.tif.
Do not invent dates or identities. Use “approximate date” or “person to identify” when facts remain uncertain.
Capture context while people still know it
The best time to identify a photograph is while the client and people familiar with the history are reviewing the batch.
Record the date or period, people, location, event, product, campaign, creator, publication rights, likely search terms, and physical source.
OCR can read words. It cannot know that the person on the left founded the company unless somebody records it.
Check quality by batch
Review a sample from every batch for focus, legibility, complete cropping, orientation, color, OCR quality, naming, master and reference copies, and connection to the physical source.
Stop when a sample fails. Repeating one box is annoying; repeating four thousand captures is expensive.
Deliver something another person can use
Do not hand over only a drive full of folders. Include a short guide explaining the structure, fields, formats, exclusions, unidentified material, and how to find an item.
Load the result into an asset library where people can search by person, campaign, date, or visual content. The archive no longer depends on whoever ran the scanner.
Good digitization is not measured by pages captured. It is measured by whether somebody can find the right photograph six months later and know whether they have permission to publish it.