videre faces
Detects faces, then groups them so you can name a person once instead of tagging each photo.
videre faces # detect, group, and store (resumable)videre faces --limit 500 # only process 500 new images, then stopvidere faces --recluster # regroup unassigned faces without re-detectingvidere faces --reset # wipe all face state and start over (asks first)videre faces --dry-run # detect but write nothingvidere faces --profile # print per-stage timing when finishedvidere faces --silent # no per-image progressvidere --library ~/Photos faces # select a different libraryvidere faces --ext heic # only HEIC photosvidere faces --date 2024-07 # only that monthThe first run downloads about 180 MB, separate from the search model.
The whole workflow
Section titled “The whole workflow”Detection and naming are separate steps. The first is slow and automatic, the second is fast and manual.
videre faces # 1. find faces and group them (slow, resumable)videre gallery # 2. name the groups in your browservidere search --person "Alice"Step 2 opens localhost:7878; naming happens on its People tab
(/people), which has three sections: People you have named, Unassigned
Clusters (groups it is confident about but has no name for), and Singletons
(faces it could not group). Drag a cluster onto a person to assign it, or create
a new person from it.
Clicking a cluster or a person opens its own page, at /people/cluster/<id> and
/people/person/<name>.
The payoff is the ratio: one drag can name forty photos.
Working through a large library
Section titled “Working through a large library”Detection on tens of thousands of photos takes hours. --limit lets you do it
in sittings:
videre faces --limit 2000 # a chunk, then stopvidere faces --limit 2000 # continue where it left offvidere faces --recluster # once, after the last chunkA limited run skips the grouping step, because grouping is a whole-library
pass that is not worth repeating after every chunk. Run --recluster once at
the end, or grouping never happens at all.
Plain videre faces with no --limit does group at the end, so if you are
happy to let it run to completion you never need --recluster.
Fixing bad grouping
Section titled “Fixing bad grouping”Grouping is the part that needs judgement, and --recluster makes experimenting
cheap: it reuses existing detections, so it takes minutes rather than hours.
videre faces --recluster --eps 0.55 # stricter: fewer, tighter groupsvidere faces --recluster --merge-sim 0.30 # more willing to merge groupsvidere faces --recluster --min-cluster-size 2 # allow smaller groupsLabeled people are never moved. Reclustering regroups only the faces you have not named. A face you assigned to a person keeps its name and its person page exactly as it is, whatever tuning flags you pass, so experimenting is safe.
Evaluate grouping quality
Section titled “Evaluate grouping quality”If you have named some faces, --evaluate can measure how the current grouping
parameters reproduce those confirmed labels:
videre faces --evaluatevidere faces --evaluate --jsonvidere faces --evaluate --eps 0.55 --min-cluster-size 3 --jsonConfirmed labels act as ground truth only for this report. Evaluation replays the selected grouping parameters over every stored face embedding in memory, including labeled faces, then compares the temporary groups with the labels. It does not detect faces, load or download models, migrate the database, or write group assignments. The groups and labels stored in your library stay unchanged.
The report includes:
- Pair precision: among pairs placed in the same group, the share with the same confirmed identity.
- Pair recall: among pairs with the same confirmed identity, the share placed in the same group.
- Mixed clusters: groups containing more than one confirmed identity.
- Fragmented identities: confirmed identities spread across more than one group or the unassigned outcome.
- Unassigned rate: the share of labeled faces left outside a group.
When a denominator does not exist, such as pair metrics for one labeled face,
the human report says n/a and JSON uses null. Current videre does not
suggest person labels, so suggestion coverage is zero and the JSON
suggestions field is null rather than a fabricated score.
Evaluation always covers the complete labeled library, so it rejects selection
flags such as --date, --path, and --tag. JSON output contains aggregate
measurements and selected parameters only. It intentionally excludes paths,
hashes, person names, face ids, and embeddings.
Your labels are frozen
Section titled “Your labels are frozen”Assigning a face to a person freezes it. Clustering, --recluster, watch, and
every other automatic process leave labeled faces exactly as they are. You
remove a label in the gallery’s People tab, or wipe everything at once with
videre faces --reset. Removing a person returns their faces to the
unassigned pool and reopens the grouping pass, so the next --recluster (or
watch’s maintenance pass) regroups them fresh.
| Symptom | Try |
|---|---|
| One person split across several groups | Lower --merge-sim toward 0.30 |
| Two people merged into one group | Raise --merge-sim, or lower --eps |
| Too many singletons, few groups | Lower --min-cluster-size to 2 |
| A group full of unrelated faces | Raise --min-face-size, lower --max-generic-sim |
Change one value at a time and look at the result. These interact, and a combination that fixes one library often overfits it.
For a single wrongly-merged group, the labeling UI’s Dissolve cluster button is better than retuning: it breaks that group back into singletons without touching anything else. Faces are not deleted.
All tuning options
Section titled “All tuning options”videre faces --eps 0.6 # how alike faces must be to group (default 0.6)videre faces --min-cluster-size 3 # fewest faces that can form a group (default 3)videre faces --merge-sim 0.35 # how readily two groups merge (default 0.35)videre faces --min-face-size 80 # ignore faces smaller than this in pixels (default 80)videre faces --min-blur 80 # ignore faces too soft to read (default 80)videre faces --max-landmark-error 7 # ignore faces the detector mislocated (default 7)videre faces --max-generic-sim 0.4 # legacy fallback, see below (default 0.4)videre faces --attach-sim 0.4 # stray faces join their nearest group (default 0.4)videre faces --batch 8 # images per batch (default 8)videre faces --workers 8 # parallel workers (default: 2x your CPU cores)videre faces --qlmanage-concurrency 6 # simultaneous HEIC conversions, shared by every videre process (default 6)Clustering parameters
Section titled “Clustering parameters”The eight grouping values above (--eps through --attach-sim) can also be
tuned from the gallery: the People page’s Recluster control previews and
applies them without leaving the page, and saves what differs from the
defaults to the library’s .videre/gallery.json. From then on each value
comes from, in order:
| run | first | then | then |
|---|---|---|---|
videre faces (a detection run, --recluster, or videre pipeline) |
its flag, when given | gallery.json |
the default |
videre watch’s grouping pass |
gallery.json |
the default |
Per value: a saved eps changes only eps. When gallery.json supplies
anything, the run says so in one line, and names a flag that overrode it:
Clustering with gallery settings: eps 0.7, min_cluster_size 2 (from .videre/gallery.json)Clustering flags override gallery settings: eps 0.65 (gallery.json has 0.7)--silent keeps the line out of the terminal (it still reaches the log at the
info level). A flag never changes gallery.json; to keep a value, apply it
from the gallery. A gallery.json that cannot be read, or a value of the
wrong type or out of range, is reported and the default is used instead, so a
typo there never stops a faces run.
Why some faces are ignored
Section titled “Why some faces are ignored”Before grouping, low-quality faces are held out. They come back as unassigned singletons, still visible and still nameable by hand, just not grouped.
This exists because low-quality faces produce near-identical featureless fingerprints regardless of who they are, so grouping them piles unrelated people into one large mixed cluster. A face is held out if it fails any check:
- Size (
--min-face-size, default 80px). Tiny crops, such as distant faces in group shots, are mostly blur once scaled up. - Sharpness (
--min-blur, default 80). The variance of the Laplacian of the aligned crop, which is the image the model is actually given. Measured on a 1,063-face library: faces that grouped had a median of 894, faces left ungrouped 182. - Alignment (
--max-landmark-error, default 7). The model never sees your photo, only a small square built by warping it so the detected eyes, nose and mouth land on fixed positions. When the detector mislocates those points the square is a mangled image, and the fingerprint describes the mangling rather than the person, so such faces resemble each other. This measures how far the five points are from being a face shape at all.
Each is disabled by setting it to 0, except --max-landmark-error, which is
disabled by setting it high.
:warning: --max-generic-sim (default 0.4) is a fallback, and only applies to
faces recorded before sharpness was measured. It compares a face against the
average of every face in the library, which sounds reasonable and is not: on a
personal library that average largely is the most photographed person, so the
check discards the very faces most worth grouping. Measured on a labelled
library, it blocked 15 photos of the owner, several of which matched a face
already in their group almost exactly. Run videre faces --reset to record
sharpness and leave it behind.
When a photo is deleted
Section titled “When a photo is deleted”When videre prune removes the last indexed copy of a photo, its detected
faces disappear from People, but the person’s name remains. If the last face
of a person was on that photo, the person can be labeled again later without
recovering the deleted photo. The historical learning journal remains, though
events that relied on the removed face no longer train new profiles.
videre faces --reset is the start-over button. It deletes every face row,
every person you have named, the detection markers, and the decode-failure
records. It also deletes face-learning events, identity questions, and learned
profiles, then immediately re-runs the full detection and grouping pipeline, so
the library ends up exactly as it would after a first-ever videre faces.
Because it deletes your labels and learning history, it asks first. The prompt shows the counts of face rows, labeled faces, people, learning events, identity questions, and learned profiles that it would delete:
videre faces --reset # asks: "deletes 400 face row(s) (312 labeled # face(s) across 9 people), ... Continue? [y/N]"videre faces --reset --yes # skip the prompt (for scripts)In a non-interactive session (a script without a terminal, a cron job), reset
refuses rather than wipe unseen: rerun with --yes after reading the counts.
--dry-run shows the counts and deletes nothing.
A reset is always whole-library. It cannot be combined with --recluster or
--limit, and selection flags are refused: wiping everything and rebuilding
only part of it is never what those combinations would deliver.
If faces has never run on the library there is nothing to reset, and reset says so and exits. Libraries whose face rows predate the detection markers reset normally.
Recovering more of one person
Section titled “Recovering more of one person”--attach-sim (default 0.4) runs a second pass after grouping: an ungrouped
face joins the group holding its nearest face, if that face is at least this
similar. Grouping normally asks a face to resemble the average of a whole
group, which someone photographed over many years can fail even when three
members match them almost exactly. Lower it to recover more, raise it (or set
it to 1) to turn the pass off.
It runs after grouping and can never merge two groups, which is what keeps it safe.
Teaching the gallery (experimental)
Section titled “Teaching the gallery (experimental)”Labeling actions in the gallery’s People page write durable, inspectable teaching evidence when there is a comparison to record: naming a cluster, moving a face, removing a face, or dissolving a cluster each record what changed and the generic face-derived factors behind it. An action with nothing to compare against (naming the very first face of a new person, for example) records none. In the background, the gallery trains small interpretable scorers from that evidence and, once a scorer passes the shipped quality gates, asks bounded yes/no questions such as “Is this Elena?” on the People page. Answering Yes names that cluster; No only teaches; Skip does neither.
What this is and is not:
- The scoring uses generic face measurements (similarities, sizes, quality), never filenames, paths, locations, or dates. ArcFace itself stays frozen; nothing about detection or grouping changes.
- Teaching evidence is stored locally, is viewable per person on the People page, and carries no images or embeddings when served; it is excluded from exported learned profiles.
- A failed training run keeps the previous profile. A passed gate promotes a profile for suggestions and questions only; production grouping still comes from the deterministic pipeline until a future recluster integration.
videre faces --resetdeletes all of it along with your labels.
Why grouping runs in two stages
Section titled “Why grouping runs in two stages”Average-linkage grouping alone fragments one person into several groups, because
one person’s photos legitimately spread wide (pose, lighting, age). A second
pass then merges any two established groups whose averaged fingerprints are at
least --merge-sim alike.
Deciding on averages rather than individual faces is what makes this safe: on real data, confirmed different people never exceeded about 0.29, while one person’s fragments ran 0.37 to 0.76. Only established groups take part, never lone singletons: a single bad crop can resemble a different person, whereas a whole group’s average cannot.
Caveats
Section titled “Caveats”Warm the HEIC cache first if your library is HEIC-heavy. Detection reads already-decoded images when they exist, about 108 ms against 7.6 s:
videre --library ~/Photos watch --heic # then Ctrl-C once it settlesvidere facesSee the thumbnail cache for what that stores and how much space it takes.
Names live on faces, not on groups. Group numbers are reassigned on every
--recluster, but the names you assigned are kept, because they are stored per
face. Retuning does not lose your work.
An interrupt can cost more than one image. Each worker batches, so a Ctrl-C
can lose up to workers x batch images of progress, which is 160 with the
defaults on a 10-core machine. Everything already committed is safe and the
rerun continues correctly; it just redoes a little.
A file that cannot be decoded is skipped after it fails twice, so a broken
or unreadable file stops costing a timeout on every run (and every watch
cycle) instead of being re-attempted forever. A single failure never skips a
file, so a one-off timeout from contention still retries; two do.
videre faces --reset clears those records along with everything else.
A failing library volume stops the run. An unreadable file is skipped as
above, and a drive whose files cannot be read at all keeps skipping them on
every run. What stops videre faces at the first failed write, with one
message instead of one error per photo, is a disk-level write failure: a
full disk, or a database gone read-only. Everything already committed is
safe; re-run when the drive is healthy and it continues.
Detection is not perfect. Faces in profile, heavily shadowed, or very small
are often missed entirely, and no amount of retuning brings them back, since
tuning only affects grouping of faces that were already found. --reset
re-detects from scratch, which only helps after a videre upgrade changes
detection itself.
Nothing is uploaded. Detection, grouping, and the naming UI all run locally.
The server on localhost:7878 is reachable only from your machine.
Performance
Section titled “Performance”| Flag | Does |
|---|---|
--workers |
detection threads (default: twice your core count) |
--profile |
print per-stage timings: load, detect, align, embed, db write |
--profile is the quickest way to see whether a slow run is bound by decoding
or by inference. Why the default oversubscribes your cores, and the measured
4.48x it is worth, are in tuning.
Scoping the run
Section titled “Scoping the run”Every flag below narrows an existing set, never widens it, and they combine: each condition must hold.
| Flag | Selects |
|---|---|
--type |
image or video. Repeatable, or comma-separated |
--ext |
file extension, e.g. mov. Repeatable, or comma-separated |
--mime |
exact type, e.g. video/quicktime. Repeatable, or comma-separated |
--after |
date on or after this (inclusive) |
--before |
date before this (exclusive) |
--date |
a whole year, month or day: YYYY, YYYY-MM, YYYY-MM-DD |
--location |
within --radius km of a place, e.g. "Berlin, Germany" |
--radius |
radius in km for --location (default 20) |
--path |
only files under this directory. Repeatable |
--has |
only files with this metadata. Supported fields: gps, date |
--missing |
only files missing this metadata. Supported fields: gps, date |
--rating |
only photos rated at least this many stars (0-5) |
--pick |
only photos with this pick state: keep or reject |
--label |
only photos with this colour label |
--like |
only liked photos |
--tag |
only files carrying this tag. Repeatable; all must be present |
--query |
only files matching a query, filters only, for OR and NOT: '(tag:deniz OR tag:plaj) -tag:ekran' |
--person and --category are deliberately absent: both are derived from data
this command produces, so selecting its input by one would be circular.
A scoped run prints N of M, so a filter that matches nothing is
distinguishable from an empty library. Full detail, including how missing data
excludes a file, is in scoping a run.
Importing names from XMP
Section titled “Importing names from XMP”If a photo has an .xmp sidecar (or embedded XMP) with named face regions, for
example one digiKam or Lightroom wrote, videre faces reads those regions and
assigns the names to the faces it detects, matching each region to a face by
where it sits in the frame. This is the read side of
videre export --xmp: a name you gave a face in another
tool imports into videre.
Names are imported while faces are being detected, so the region has a detected
face to attach to. To import into a library whose faces were detected earlier,
re-run with --reset.
The --xmp flag decides who wins when both carry a name:
--xmp |
Effect |
|---|---|
db (default) |
The database wins: imported names fill only faces you have not already named |
file |
The sidecar wins: an imported name replaces an existing one |
newest |
Reserved; currently behaves as db |
The default comes from the xmp_precedence config key (see
videre config), the same setting scan uses for ratings.
Rotated photos
Section titled “Rotated photos”Photos stored rotated on disk (an EXIF orientation tag, common in exports and
edits) are decoded the way you see them before detection and grouping, so
sideways-stored files cluster exactly like upright ones. Files processed by
older videre versions were detected on the stored pixel canvas, which could
leave a person’s own photos as unassigned singletons; re-run
videre faces --reset to bring a library current, and see
troubleshooting for the recovery recipes and
their costs.
More detail
Section titled “More detail”- Long-running jobs covers running this alongside other commands, and resuming.
- Caches and disk use covers the decode cache this reads from.
- Backing up covers why the names you assign are the one thing that cannot be recomputed.