Documentation

How to Find One Photo Among 100,000 Images

Past a certain volume, scrolling stops working. The four handles that make a large collection genuinely searchable.

Organization

The point where browsing collapses

Up to a few thousand images, you find everything by scrolling: visual memory plus a date-based folder tree is enough. Around 20,000, scrolling becomes a cost. Past 50,000 it stops working entirely — not because the tool is slow, but because you can no longer hold a mental map of the collection.

Beyond that threshold, finding an image is no longer a display problem. It’s a description problem.

Four handles to catch an image by

A large collection must be searchable through several independent entry points, because you never remember the same detail twice:

  1. When. Capture date, read from EXIF, is the one piece of metadata you always have and never have to type. It’s your safety net.
  2. What. Keywords drawn from a controlled vocabulary — not free text, which produces twenty spellings of one subject.
  3. Where. A normalised place, as an entity rather than free text, so “Marseille” is always the same Marseille.
  4. Why. Membership: project, collection, series, commission. Often the most effective, because you remember the working context better than the content.

An image described along all four axes can be caught by any of them. An image described along one is lost the moment that particular memory fails.

What actually speeds things up

  • Facets rather than a single query: filtering cumulatively (year, then project, then keyword) narrows 100,000 to a few dozen in three clicks.
  • Curation ratings. Searching among picks rather than everything divides the volume by ten.
  • Local previews. A search that has to read RAW files from an external drive is slow for physical reasons; lightweight local previews change the experience.
  • Working offline. A collection searchable without a network avoids the wait that discourages use.

The mistake to avoid

Trying to describe everything in detail, all at once. A 100,000-image collection can’t be caught up in one pass, and an exhaustive data-entry campaign is always abandoned halfway.

Describe coarsely but completely first: project and year across the whole collection. Then refine in batches, starting with what is most likely to be searched.

In practice

  1. Make sure every image carries at least a project and a usable date.
  2. Build a controlled vocabulary of twenty to fifty terms, not three hundred.
  3. Normalise places and people into reusable entities.
  4. Use bulk editing to work in batches, never record by record.
  5. Measure: if a typical search takes more than thirty seconds, description is what’s missing, not hardware.

Obscura Flow indexes locally, offers cumulative faceted search, supports synonyms and aliases, and provides bulk editing to catch up an existing collection without redoing it one image at a time.