Tuning
Every default here was measured, not guessed. You should not need this page: the defaults are the right answer for almost every library, and the commands that expose these flags work without them. Read on when a run is slow enough that you want to know why, or when your hardware is unusual.
Each flag is also listed on its own command’s page. This is where the reasoning lives.
Rule of thumb
Section titled “Rule of thumb”| Symptom | Look at |
|---|---|
| scan skips files because reads stop progressing | Slow drives and large files |
| face detection is slower than expected | --workers, --profile |
| an interrupted run loses too much work | --chunk |
| embedding a big library takes too long | --batch, VIDERE_EMBED_DTYPE |
Slow drives and large files
Section titled “Slow drives and large files”Hashing has a 20-second no-progress window, not a total deadline. A file can take hours if bytes keep arriving; a read that stops returning bytes for 20 seconds is skipped and can be retried. Even a very slow trickle counts as progress. Before hashing, file metadata is checked under a separate five-second limit, so a stalled stat cannot start a long read.
read-rate no longer controls hashing. It still scales the total timeout for
other whole-file operations, such as XMP and log reads. Lower it if those
operations time out on a healthy slow mount:
videre config set read-rate 5io-workers is a separate safety limit on helper threads. If a kernel read
does not return after the caller times out, its thread keeps a slot until the
read finishes. Increase the limit only if a healthy workload is exhausting it;
the default is four times the CPU count, clamped to 32 through 128.
Face detection
Section titled “Face detection”--workers defaults to twice your core count. The oversubscription is
deliberate: HEIC decoding waits on an external subprocess rather than the CPU,
so other workers use the cores meanwhile.
Measured on a 10-core machine: about 3.23x faster than a single worker, and raising the HEIC conversion limit from 3 to 6 added another 1.23x, for roughly 4.48x overall. Past 6 the gain is a few percent while per-image detect time creeps up, so 6 is the default.
--profile prints per-stage timings (load, detect, align, embed, db write) and
is the quickest way to see whether a slow run is bound by decoding or inference.
--eps and --min-cluster-size control how faces are grouped into people, not how
fast detection runs. They are covered on the
faces page, since changing them changes
results rather than speed.
Embedding
Section titled “Embedding”--batch is how many images go through the model at once. The default is 32 and
the maximum is 96.
Larger batches buy little anyway: 31.0 ms/image at 96 against 39.1 ms at 768.
VIDERE_EMBED_DTYPE=f16 switches inference to half precision: about 11% faster
on pure JPEG/PNG, 7% on a realistic mix, with no meaningful quality change. It
is opt-in because 7% did not justify perturbing an existing library, and it does
not affect vectors already written.
How often work is committed
Section titled “How often work is committed”--chunk sets how many rows are written per transaction. Larger values are
slightly faster and lose more if the run is interrupted. Both embed and
classify accept it; the default is 500.
Resumability does not depend on this. Every command records what it has already tried, not just what produced a result, so an interrupted run picks up where it stopped either way. See long-running jobs.
Concurrency across commands
Section titled “Concurrency across commands”The HEIC and video conversion limit (--qlmanage-concurrency, default 6) is
shared by every videre process on the machine, so running faces next to
embed, a gallery or watch does not multiply the load on QuickLook: the
commands take turns within one limit. See
running things at once.