videre scan
Scans the selected library recursively and records every
supported media file in .videre/hashes.db inside
that library. Run this first. Other commands read the database it creates.
cd ~/Photosvidere scan # scan the current directory (only new and changed files)videre scan --force # re-read and re-hash every filevidere scan --silent # suppress progress outputvidere scan --json # print one JSON summary objectUse --library to select a different library without changing directory:
videre --library ~/Photos scanvidere scan --library ~/Photos--library may appear before or after the command, but only once. A relative
value is resolved from the directory where videre was invoked.
Local state
Section titled “Local state”The first scan creates this state inside the selected library:
.videre/├── config.toml├── hashes.db└── locks/The database location is fixed at .videre/hashes.db inside the library. scan
takes no directory operand, no path filter, and no database or output-location
selector: the library is chosen the same way for every command. To scan another
collection, run the command from that directory or select it with --library.
Re-running is safe. Existing rows are updated by path while annotations and location fields owned by later processing are preserved.
What it records
Section titled “What it records”For every file, scan stores its content hash, size, timestamps, extension, detected type, and available media metadata. Photo metadata includes capture date, GPS coordinates, and dimensions. Video metadata can also include capture date, GPS coordinates, dimensions, duration, and codec.
Scan does not prepare semantic search, near-duplicate fingerprints or faces.
Run videre embed and videre faces
separately for those features.
Incremental by default
Section titled “Incremental by default”Scan skips a file whose recorded row is already current: the same size and
modification time as the last scan, with its type already identified. New files, and files
that changed since the last scan, are processed; everything else costs a cheap
stat, not a re-read. So re-scanning a large library that has barely changed is
fast, and re-running is always safe.
videre scan # only new and changed filesvidere scan --force # re-read and re-hash every file--force ignores the skip and re-reads every file. Reach for it
after restoring from a backup, or to re-examine files whose contents changed
without their modification time changing (rare, but some tools preserve mtime).
A file whose bytes were read but whose type could not be identified gets an explicit sentinel, so it counts as complete and is not re-read on every scan.
Reading marks from XMP
Section titled “Reading marks from XMP”Scan reads ratings, colour labels, and keywords from an adjacent XMP sidecar and
from the XMP packet embedded in the photo. They are combined field by field: the
sidecar’s value wins where it has one, the photo’s fills whatever the sidecar
lacks, and keywords from both are kept. So a rating another app writes into the
photo is read even when the photo also has a sidecar, for example one
videre export wrote with face regions only. The --xmp
option controls how ratings and labels are reconciled with database values:
| Value | Behaviour |
|---|---|
db |
Database values win. XMP fills missing values |
file |
XMP values win |
newest |
Reserved for timestamp comparison; currently behaves as db with a warning |
The local default is configured with:
videre config set xmp dbKeywords are imported as additive tags, independently of the rating and label precedence.
Reading is incremental. A file’s XMP is read the first time it is scanned and
again whenever it changes: either the media file itself changes, or its sidecar
changes on its own (you re-rate a photo in another tool while the media file is
untouched). A sidecar change is detected from the sidecar’s modification time,
so re-rating in another tool is picked up on the next scan without re-reading
the media. A file whose media and sidecar are both unchanged is skipped
entirely, so a repeat scan reads no metadata at all. --xmp file and --xmp newest still reconcile every file, unchanged or not, because they exist to
overwrite the database from the file.
Nested libraries
Section titled “Nested libraries”A parent scan includes media in nested directories, even when one of those
directories is also used as a separate library. It skips every .videre
directory before descending, so nested databases, config files, caches, and
other state are never scanned as media or merged into the parent library.
For example, scanning ~/Photos/x and later scanning ~/Photos creates two
independent databases. The parent sees media under x; it does not adopt or
modify x/.videre.
Caveats
Section titled “Caveats”Rows are keyed by path. Moving a file creates a new path on the next scan
and leaves the old row until videre prune removes it.
Faces, embeddings, marks, and tags are keyed by content hash.
Unreadable files are skipped. Permission failures and files that time out
are reported without stopping the rest of the scan. A file with no recorded row
is always retried, so just run videre scan again after fixing the underlying
problem.
Reading a file’s bytes is the expensive part. A first scan, a changed
file, or a --force pass reads every byte; on external drives and network
shares disk speed usually dominates. Hashing has no total time limit. It
continues as long as bytes arrive, even for a very large or slow file, and
skips a file after 20 seconds without read progress. The initial file stat
has a separate five-second limit. read-rate does not control hashing;
it still applies to other size-bounded file reads. See
tuning for details.
While a file of 1 GB or more is being hashed, its bytes read so far and the
current rate show beside the progress bar (for several at once, their
totals), so one large video does not leave the counter sitting still for
minutes. A falling rate is the early sign of a stalled drive. Without a
terminal the same line is logged every 30 seconds; --silent shows none of
it.