Skip to content

videre embed

Prepares photos so they can be searched by description. One-time per photo, and resumable.

Terminal window
videre embed # process everything not done yet
videre --library ~/Photos embed # select a different library
videre embed --model <model-id> # prepare with a specific model, kept separately
videre embed --batch 64 # images per inference batch (default 32, max 96)
videre embed --chunk 1000 # rows saved per transaction (default 500)
videre embed --silent # no per-image progress
videre embed --type video # only videos
videre embed --after 2024-01-01 # only files from this year on
videre embed --reprocess # re-embed, including already-embedded files

The first run downloads about 1.5 GB of model data, then works through every image. On a large library this takes hours, so plan to leave it running.

Terminal window
videre embed # start; Ctrl-C whenever you like
videre stats # how far it got
videre embed # continue from there

videre stats reports a row count per model, so the gap between that and your photo count is what remains:

embeddings
google/siglip2-base-patch16-224 12,481 rows 768 dims 28.4 MB

Work is committed every --chunk rows (500 by default), so an interrupt loses at most that much. There is no separate resume flag: rerunning is resuming, because the command only ever looks for hashes that have no vector yet.

A file that cannot be decoded (an unreadable file, or one that repeatedly times out) is skipped after it has failed twice, so a broken file stops costing a timeout on every run instead of staying pending forever. It is logged when it is skipped.

--reprocess lifts both: it re-embeds every eligible image, including ones that already have a vector under the model, and clears the record of skipped undecodable files so they are attempted again. It exists for the rare case where the pixels a model sees have changed for reasons the files themselves did not, such as a decode fix in videre. See troubleshooting.

Adding photos later works the same way. Run videre scan to pick them up, then videre embed again to cover only the new ones.

Embed also stores each file’s near-duplicate fingerprint, which videre dedupe --similar reads. Files embedded before this was added get theirs on the next run, decoded and hashed without loading the model; the summary line counts them as fingerprints added.

Type Embedded?
jpg, png, tiff, webp, bmp, gif yes
heic macOS only
mov, mp4 macOS only, from one frame
dng never

.dng is excluded up front rather than attempted and failed, so a library full of raw files does not waste a decode attempt on each of them every run. Their EXIF is still recorded by scan.

Video embedding uses a single representative frame, not the motion, so video search is noticeably weaker than photo search. A clip whose subject appears later than its opening frame will not match well.

On Linux, HEIC and video are skipped entirely. See platform support.

Each model writes to its own database, so they never overwrite each other and you can hold several at once:

Terminal window
videre embed # the default model
videre embed --model google/siglip2-base-patch16-384 # a second, higher-resolution one
videre stats # row counts per model
videre search "sunset" --model google/siglip2-base-patch16-384

Preparing a second model does not disturb the first, and switching between them invalidates nothing. That makes it practical to try a larger model on a real library and keep the old vectors until you are convinced.

The per-model database is created only when embedding actually starts: a model that fails to download or load, or a run with nothing left to embed, leaves no file behind. Anything stats lists as a model therefore has a real embedding pass behind it, and one that failed mid-flight is cleaned up by the next prune.

Be aware of the cost before starting: a second model means a second full pass over every image, plus its own download (1.5 GB for siglip2-base-patch16-384, 4.5 GB for siglip2-so400m-patch14-384) and its own 130 MB to 190 MB of vectors per 70,000 photos.

If you settle on the new one, make it the default and the old vectors can be deleted by removing that model’s file:

Terminal window
videre config set model google/siglip2-base-patch16-384

See search models for where the files live.

Do not run embed and faces at the same time. Both drive HEIC conversion through the same macOS service, and each limits itself independently, so together they overwhelm it. Measured with both running: a HEIC file took over 16 seconds against about 7.6 normally, and one exceeded the timeout entirely. Nothing is lost, since skipped files retry next run, but it is slower than doing one after the other. The same applies to videre watch if it is running with its faces or HEIC stages.

Vectors are keyed by content, not path. Two copies of one photo cost a single embedding, and moving a file keeps its vector as soon as it is re-scanned.

--chunk controls how often work is committed. Larger values are slightly faster and lose more on an interrupt.

VIDERE_EMBED_DTYPE=f16 switches inference to half precision: about 11% faster on pure JPEG/PNG, 7% on a realistic mix, with no meaningful quality change. It is opt-in because 7% did not justify perturbing an existing library, and it does not affect vectors already written.

Every flag below narrows an existing set, never widens it, and they combine: each condition must hold.

Flag Selects
--type image or video. Repeatable, or comma-separated
--ext file extension, e.g. mov. Repeatable, or comma-separated
--mime exact type, e.g. video/quicktime. Repeatable, or comma-separated
--after date on or after this (inclusive)
--before date before this (exclusive)
--date a whole year, month or day: YYYY, YYYY-MM, YYYY-MM-DD
--location within --radius km of a place, e.g. "Berlin, Germany"
--radius radius in km for --location (default 20)
--path only files under this directory. Repeatable
--has only files with this metadata. Supported fields: gps, date
--missing only files missing this metadata. Supported fields: gps, date
--rating only photos rated at least this many stars (0-5)
--pick only photos with this pick state: keep or reject
--label only photos with this colour label
--like only liked photos
--tag only files carrying this tag. Repeatable; all must be present
--query only files matching a query, filters only, for OR and NOT: '(tag:deniz OR tag:plaj) -tag:ekran'

--person and --category are deliberately absent: both are derived from data this command produces, so selecting its input by one would be circular.

A scoped run prints N of M, so a filter that matches nothing is distinguishable from an empty library. Full detail, including how missing data excludes a file, is in scoping a run.