Background jobs
Lalabase processes work like sending email, notifications, and AI computations asynchronously via Solid Queue — the native job system of Rails 8. Jobs live in the same PostgreSQL database; a separate service like Sidekiq is not required.
The worker process
In the container installation, the worker runs as its own service next to the web application. Check its status with:
docker compose ps
docker compose logs -f worker
Spotting stuck jobs
If jobs pile up, the worker process is usually stopped, or an external dependency (e.g. the AI provider) is unreachable. The worker logs show the cause.
Restart the worker after every update so it loads the current code — otherwise it keeps processing jobs with the old version.
Recurring jobs
Some tasks run on a fixed schedule (Solid Queue recurring tasks in
config/recurring.yml), such as clearing finished jobs, scheduled account deletions,
and documentation sync.
Purging unattached files (GDPR)
An upload creates a file in storage before its form is submitted. If the form is never submitted — or an image is dropped into a text, removed again, and the text is saved — the file stays behind with nothing pointing at it. Rails does not sweep those.
Daily at 05:15, UnattachedBlobPurgeJob deletes every file that has been unattached
for more than seven days. Only from that deletion onwards do such files enter the same
backup retention windows as everything else — without the job they enter none at all.
Between upload and form submission, a not-yet-attached file is technically indistinguishable from an abandoned one. The waiting period is what keeps the job from deleting a file out from under someone with an open form. That is also why this job has no enforcement variable: it deletes from its first run, and the safety lives in the grace period and a per-run cap.
Preview what would be deleted (deletes nothing):
bin/rails attachments:unattached_report
What the job did is in the application log under [unattached_blob_purge] — one line per
deleted file with name, type, size and date, never with content.
Register of deleted attachments (GDPR)
When a file attachment is deleted, Lalabase writes an entry into a register: organisation, filename, type, size, carrier and the moment of deletion. The file's content is never recorded. The register is what proves to the controller that a deletion happened and when (accountability, Art. 5(2) GDPR).
Two recurring jobs maintain it:
| Job | Time | Purpose |
|---|---|---|
AttachmentBackupConfirmationJob |
05:45 | carries over, from the backup script's own log, when the copy left mirror and remote store |
AttachmentDeletionRecordPurgeJob |
06:00 | deletes entries that are three years old |
The three years run from the deletion in production, not from the backup expiry carried over later. That carry-over may never arrive (a failed backup run, a lost log line, a file that never made it into a backup at all). Anchored there, such an entry would have no end date at all — precisely the defect the register exists to document.
The carry-over requires the worker to be able to read the backup script's log. The
script creates it root-owned; releasing it through a read group is described in
ops/backup/README.md. Without that release the job logs a warning under
[deletion_register] instead of silently doing nothing.
Answering a request. When a controller asks for the proof concerning their organisation, two routes give the same answer: the Deletion register page in the admin area, or the command-line report.
bin/rails deletion_register:report # everything
ORG=<organisation-uuid> bin/rails deletion_register:report
Both name the date the register started and count entries without a confirmed backup expiry separately. Both belong in the answer: anything deleted before that date is not recorded, and an entry without a carry-over is one whose expiry nobody has confirmed. The organisation list is read from the register itself rather than from the organisations table, because an entry outlives the organisation it concerns.
Embedding cleanup (GDPR)
For semantic search and the AI assistant, Lalabase stores derived embedding vectors. When the underlying content is deleted, the deletion paths normally remove their vectors immediately and synchronously — especially for personal data such as meeting transcripts and analyses.
As an additional safety net, a cleanup job (Embeddings::OrphanGcJob) runs daily at
04:30 and reclaims orphaned vectors whose source content was removed through a
technical edge path without synchronous cleanup. As a result, derived vectors of deleted
content persist for at most ~24 hours.
For safety the daily job starts in dry-run: it only logs how many
orphaned vectors it would reclaim, deleting nothing. Check the log counts for
the first few nights (or DRY_RUN=1 bin/rails "embeddings:gc"). Once the
numbers look right, set the environment variable EMBEDDINGS_GC_ENFORCE=1
on the worker process — then the job actually deletes. The manual task below (without
DRY_RUN) always deletes for real, regardless of this variable.
Run it manually (e.g. after a large deletion):
bin/rails "embeddings:gc" # one global sweep across all organisations
DRY_RUN=1 bin/rails "embeddings:gc" # count only, delete nothing
Reclaiming old vector spaces (retention)
When an organisation switches its embedding model, the old vector space stays behind as a
roll-back net — you can return to it at any time. How many spaces an organisation keeps
is governed by the embedding_retention_limit setting (default: 2, i.e. the active one plus
one older). The same nightly job reclaims everything beyond that.
Two spaces it will never touch: the one currently serving search, and one that is being built right now. The space under construction also does not count against the limit — otherwise the very space you are comparing against would disappear while you evaluate two models against each other.
This cleanup has its own switch, deliberately separate from
EMBEDDINGS_GC_ENFORCE. The orphan sweep deletes vectors whose source content
is provably gone — those are worthless. This one deletes intact vectors: the
roll-back net. Throwing one switch must not silently arm the other. Set
EMBEDDINGS_RETENTION_ENFORCE=1 on the worker process once the dry-run numbers
look right. Careful: the manual task above (bin/rails "embeddings:gc" without
DRY_RUN) also reclaims spaces beyond the limit immediately — regardless of
either variable. DRY_RUN=1 is the only guard there.
Timing note for rollbacks: While rake embeddings:rollback runs (reactivation +
delta reindex — this takes as long as the corpus needs), the target space is temporarily
neither active nor configured, and therefore a regular candidate for the retention sweep.
Keep rollbacks out of the nightly GC window (04:30) while
EMBEDDINGS_RETENTION_ENFORCE=1 is set.
Vectors from a model switch that predates space management belong to no registered
space. They are deliberately not deleted: the same signature also
describes a freshly configured target whose build has not started yet, so deleting on that
suspicion would hit a build in progress. The job logs them as UNREGISTERED
instead. Seeing rows there is a cleanup decision for you, not a malfunction.