Skip to content

videre faces

Detects faces, then groups them so you can name a person once instead of tagging each photo.

Terminal window
videre faces # detect, group, and store (resumable)
videre faces --limit 500 # only process 500 new images, then stop
videre faces --recluster # regroup unassigned faces without re-detecting
videre faces --reset # wipe all face state and start over (asks first)
videre faces --dry-run # detect but write nothing
videre faces --profile # print per-stage timing when finished
videre faces --silent # no per-image progress
videre --library ~/Photos faces # select a different library
videre faces --ext heic # only HEIC photos
videre faces --date 2024-07 # only that month

The first run downloads about 180 MB, separate from the search model.

Detection and naming are separate steps. The first is slow and automatic, the second is fast and manual.

Terminal window
videre faces # 1. find faces and group them (slow, resumable)
videre gallery # 2. name the groups in your browser
videre search --person "Alice"

Step 2 opens localhost:7878; naming happens on its People tab (/people), which has three sections: People you have named, Unassigned Clusters (groups it is confident about but has no name for), and Singletons (faces it could not group). Drag a cluster onto a person to assign it, or create a new person from it.

Clicking a cluster or a person opens its own page, at /people/cluster/<id> and /people/person/<name>.

The payoff is the ratio: one drag can name forty photos.

Detection on tens of thousands of photos takes hours. --limit lets you do it in sittings:

Terminal window
videre faces --limit 2000 # a chunk, then stop
videre faces --limit 2000 # continue where it left off
videre faces --recluster # once, after the last chunk

A limited run skips the grouping step, because grouping is a whole-library pass that is not worth repeating after every chunk. Run --recluster once at the end, or grouping never happens at all.

Plain videre faces with no --limit does group at the end, so if you are happy to let it run to completion you never need --recluster.

Grouping is the part that needs judgement, and --recluster makes experimenting cheap: it reuses existing detections, so it takes minutes rather than hours.

Terminal window
videre faces --recluster --eps 0.55 # stricter: fewer, tighter groups
videre faces --recluster --merge-sim 0.30 # more willing to merge groups
videre faces --recluster --min-cluster-size 2 # allow smaller groups

Labeled people are never moved. Reclustering regroups only the faces you have not named. A face you assigned to a person keeps its name and its person page exactly as it is, whatever tuning flags you pass, so experimenting is safe.

If you have named some faces, --evaluate can measure how the current grouping parameters reproduce those confirmed labels:

Terminal window
videre faces --evaluate
videre faces --evaluate --json
videre faces --evaluate --eps 0.55 --min-cluster-size 3 --json

Confirmed labels act as ground truth only for this report. Evaluation replays the selected grouping parameters over every stored face embedding in memory, including labeled faces, then compares the temporary groups with the labels. It does not detect faces, load or download models, migrate the database, or write group assignments. The groups and labels stored in your library stay unchanged.

The report includes:

  • Pair precision: among pairs placed in the same group, the share with the same confirmed identity.
  • Pair recall: among pairs with the same confirmed identity, the share placed in the same group.
  • Mixed clusters: groups containing more than one confirmed identity.
  • Fragmented identities: confirmed identities spread across more than one group or the unassigned outcome.
  • Unassigned rate: the share of labeled faces left outside a group.

When a denominator does not exist, such as pair metrics for one labeled face, the human report says n/a and JSON uses null. Current videre does not suggest person labels, so suggestion coverage is zero and the JSON suggestions field is null rather than a fabricated score.

Evaluation always covers the complete labeled library, so it rejects selection flags such as --date, --path, and --tag. JSON output contains aggregate measurements and selected parameters only. It intentionally excludes paths, hashes, person names, face ids, and embeddings.

Assigning a face to a person freezes it. Clustering, --recluster, watch, and every other automatic process leave labeled faces exactly as they are. You remove a label in the gallery’s People tab, or wipe everything at once with videre faces --reset. Removing a person returns their faces to the unassigned pool and reopens the grouping pass, so the next --recluster (or watch’s maintenance pass) regroups them fresh.

Symptom Try
One person split across several groups Lower --merge-sim toward 0.30
Two people merged into one group Raise --merge-sim, or lower --eps
Too many singletons, few groups Lower --min-cluster-size to 2
A group full of unrelated faces Raise --min-face-size, lower --max-generic-sim

Change one value at a time and look at the result. These interact, and a combination that fixes one library often overfits it.

For a single wrongly-merged group, the labeling UI’s Dissolve cluster button is better than retuning: it breaks that group back into singletons without touching anything else. Faces are not deleted.

Terminal window
videre faces --eps 0.6 # how alike faces must be to group (default 0.6)
videre faces --min-cluster-size 3 # fewest faces that can form a group (default 3)
videre faces --merge-sim 0.35 # how readily two groups merge (default 0.35)
videre faces --min-face-size 80 # ignore faces smaller than this in pixels (default 80)
videre faces --min-blur 80 # ignore faces too soft to read (default 80)
videre faces --max-landmark-error 7 # ignore faces the detector mislocated (default 7)
videre faces --max-generic-sim 0.4 # legacy fallback, see below (default 0.4)
videre faces --attach-sim 0.4 # stray faces join their nearest group (default 0.4)
videre faces --batch 8 # images per batch (default 8)
videre faces --workers 8 # parallel workers (default: 2x your CPU cores)
videre faces --qlmanage-concurrency 6 # simultaneous HEIC conversions, shared by every videre process (default 6)

The eight grouping values above (--eps through --attach-sim) can also be tuned from the gallery: the People page’s Recluster control previews and applies them without leaving the page, and saves what differs from the defaults to the library’s .videre/gallery.json. From then on each value comes from, in order:

run first then then
videre faces (a detection run, --recluster, or videre pipeline) its flag, when given gallery.json the default
videre watch’s grouping pass gallery.json the default

Per value: a saved eps changes only eps. When gallery.json supplies anything, the run says so in one line, and names a flag that overrode it:

Clustering with gallery settings: eps 0.7, min_cluster_size 2 (from .videre/gallery.json)
Clustering flags override gallery settings: eps 0.65 (gallery.json has 0.7)

--silent keeps the line out of the terminal (it still reaches the log at the info level). A flag never changes gallery.json; to keep a value, apply it from the gallery. A gallery.json that cannot be read, or a value of the wrong type or out of range, is reported and the default is used instead, so a typo there never stops a faces run.

Before grouping, low-quality faces are held out. They come back as unassigned singletons, still visible and still nameable by hand, just not grouped.

This exists because low-quality faces produce near-identical featureless fingerprints regardless of who they are, so grouping them piles unrelated people into one large mixed cluster. A face is held out if it fails any check:

  • Size (--min-face-size, default 80px). Tiny crops, such as distant faces in group shots, are mostly blur once scaled up.
  • Sharpness (--min-blur, default 80). The variance of the Laplacian of the aligned crop, which is the image the model is actually given. Measured on a 1,063-face library: faces that grouped had a median of 894, faces left ungrouped 182.
  • Alignment (--max-landmark-error, default 7). The model never sees your photo, only a small square built by warping it so the detected eyes, nose and mouth land on fixed positions. When the detector mislocates those points the square is a mangled image, and the fingerprint describes the mangling rather than the person, so such faces resemble each other. This measures how far the five points are from being a face shape at all.

Each is disabled by setting it to 0, except --max-landmark-error, which is disabled by setting it high.

:warning: --max-generic-sim (default 0.4) is a fallback, and only applies to faces recorded before sharpness was measured. It compares a face against the average of every face in the library, which sounds reasonable and is not: on a personal library that average largely is the most photographed person, so the check discards the very faces most worth grouping. Measured on a labelled library, it blocked 15 photos of the owner, several of which matched a face already in their group almost exactly. Run videre faces --reset to record sharpness and leave it behind.

When videre prune removes the last indexed copy of a photo, its detected faces disappear from People, but the person’s name remains. If the last face of a person was on that photo, the person can be labeled again later without recovering the deleted photo. The historical learning journal remains, though events that relied on the removed face no longer train new profiles.

videre faces --reset is the start-over button. It deletes every face row, every person you have named, the detection markers, and the decode-failure records. It also deletes face-learning events, identity questions, and learned profiles, then immediately re-runs the full detection and grouping pipeline, so the library ends up exactly as it would after a first-ever videre faces.

Because it deletes your labels and learning history, it asks first. The prompt shows the counts of face rows, labeled faces, people, learning events, identity questions, and learned profiles that it would delete:

Terminal window
videre faces --reset # asks: "deletes 400 face row(s) (312 labeled
# face(s) across 9 people), ... Continue? [y/N]"
videre faces --reset --yes # skip the prompt (for scripts)

In a non-interactive session (a script without a terminal, a cron job), reset refuses rather than wipe unseen: rerun with --yes after reading the counts. --dry-run shows the counts and deletes nothing.

A reset is always whole-library. It cannot be combined with --recluster or --limit, and selection flags are refused: wiping everything and rebuilding only part of it is never what those combinations would deliver.

If faces has never run on the library there is nothing to reset, and reset says so and exits. Libraries whose face rows predate the detection markers reset normally.

--attach-sim (default 0.4) runs a second pass after grouping: an ungrouped face joins the group holding its nearest face, if that face is at least this similar. Grouping normally asks a face to resemble the average of a whole group, which someone photographed over many years can fail even when three members match them almost exactly. Lower it to recover more, raise it (or set it to 1) to turn the pass off.

It runs after grouping and can never merge two groups, which is what keeps it safe.

Labeling actions in the gallery’s People page write durable, inspectable teaching evidence when there is a comparison to record: naming a cluster, moving a face, removing a face, or dissolving a cluster each record what changed and the generic face-derived factors behind it. An action with nothing to compare against (naming the very first face of a new person, for example) records none. In the background, the gallery trains small interpretable scorers from that evidence and, once a scorer passes the shipped quality gates, asks bounded yes/no questions such as “Is this Elena?” on the People page. Answering Yes names that cluster; No only teaches; Skip does neither.

What this is and is not:

  • The scoring uses generic face measurements (similarities, sizes, quality), never filenames, paths, locations, or dates. ArcFace itself stays frozen; nothing about detection or grouping changes.
  • Teaching evidence is stored locally, is viewable per person on the People page, and carries no images or embeddings when served; it is excluded from exported learned profiles.
  • A failed training run keeps the previous profile. A passed gate promotes a profile for suggestions and questions only; production grouping still comes from the deterministic pipeline until a future recluster integration.
  • videre faces --reset deletes all of it along with your labels.

Average-linkage grouping alone fragments one person into several groups, because one person’s photos legitimately spread wide (pose, lighting, age). A second pass then merges any two established groups whose averaged fingerprints are at least --merge-sim alike.

Deciding on averages rather than individual faces is what makes this safe: on real data, confirmed different people never exceeded about 0.29, while one person’s fragments ran 0.37 to 0.76. Only established groups take part, never lone singletons: a single bad crop can resemble a different person, whereas a whole group’s average cannot.

Warm the HEIC cache first if your library is HEIC-heavy. Detection reads already-decoded images when they exist, about 108 ms against 7.6 s:

Terminal window
videre --library ~/Photos watch --heic # then Ctrl-C once it settles
videre faces

See the thumbnail cache for what that stores and how much space it takes.

Names live on faces, not on groups. Group numbers are reassigned on every --recluster, but the names you assigned are kept, because they are stored per face. Retuning does not lose your work.

An interrupt can cost more than one image. Each worker batches, so a Ctrl-C can lose up to workers x batch images of progress, which is 160 with the defaults on a 10-core machine. Everything already committed is safe and the rerun continues correctly; it just redoes a little.

A file that cannot be decoded is skipped after it fails twice, so a broken or unreadable file stops costing a timeout on every run (and every watch cycle) instead of being re-attempted forever. A single failure never skips a file, so a one-off timeout from contention still retries; two do. videre faces --reset clears those records along with everything else.

A failing library volume stops the run. An unreadable file is skipped as above, and a drive whose files cannot be read at all keeps skipping them on every run. What stops videre faces at the first failed write, with one message instead of one error per photo, is a disk-level write failure: a full disk, or a database gone read-only. Everything already committed is safe; re-run when the drive is healthy and it continues.

Detection is not perfect. Faces in profile, heavily shadowed, or very small are often missed entirely, and no amount of retuning brings them back, since tuning only affects grouping of faces that were already found. --reset re-detects from scratch, which only helps after a videre upgrade changes detection itself.

Nothing is uploaded. Detection, grouping, and the naming UI all run locally. The server on localhost:7878 is reachable only from your machine.

Flag Does
--workers detection threads (default: twice your core count)
--profile print per-stage timings: load, detect, align, embed, db write

--profile is the quickest way to see whether a slow run is bound by decoding or by inference. Why the default oversubscribes your cores, and the measured 4.48x it is worth, are in tuning.

Every flag below narrows an existing set, never widens it, and they combine: each condition must hold.

Flag Selects
--type image or video. Repeatable, or comma-separated
--ext file extension, e.g. mov. Repeatable, or comma-separated
--mime exact type, e.g. video/quicktime. Repeatable, or comma-separated
--after date on or after this (inclusive)
--before date before this (exclusive)
--date a whole year, month or day: YYYY, YYYY-MM, YYYY-MM-DD
--location within --radius km of a place, e.g. "Berlin, Germany"
--radius radius in km for --location (default 20)
--path only files under this directory. Repeatable
--has only files with this metadata. Supported fields: gps, date
--missing only files missing this metadata. Supported fields: gps, date
--rating only photos rated at least this many stars (0-5)
--pick only photos with this pick state: keep or reject
--label only photos with this colour label
--like only liked photos
--tag only files carrying this tag. Repeatable; all must be present
--query only files matching a query, filters only, for OR and NOT: '(tag:deniz OR tag:plaj) -tag:ekran'

--person and --category are deliberately absent: both are derived from data this command produces, so selecting its input by one would be circular.

A scoped run prints N of M, so a filter that matches nothing is distinguishable from an empty library. Full detail, including how missing data excludes a file, is in scoping a run.

If a photo has an .xmp sidecar (or embedded XMP) with named face regions, for example one digiKam or Lightroom wrote, videre faces reads those regions and assigns the names to the faces it detects, matching each region to a face by where it sits in the frame. This is the read side of videre export --xmp: a name you gave a face in another tool imports into videre.

Names are imported while faces are being detected, so the region has a detected face to attach to. To import into a library whose faces were detected earlier, re-run with --reset.

The --xmp flag decides who wins when both carry a name:

--xmp Effect
db (default) The database wins: imported names fill only faces you have not already named
file The sidecar wins: an imported name replaces an existing one
newest Reserved; currently behaves as db

The default comes from the xmp_precedence config key (see videre config), the same setting scan uses for ratings.

Photos stored rotated on disk (an EXIF orientation tag, common in exports and edits) are decoded the way you see them before detection and grouping, so sideways-stored files cluster exactly like upright ones. Files processed by older videre versions were detected on the stored pixel canvas, which could leave a person’s own photos as unassigned singletons; re-run videre faces --reset to bring a library current, and see troubleshooting for the recovery recipes and their costs.