Skip to content

Background jobs

Lalabase processes work like sending email, notifications, and AI computations asynchronously via Solid Queue — the native job system of Rails 8. Jobs live in the same PostgreSQL database; a separate service like Sidekiq is not required.

The worker process

In the container installation, the worker runs as its own service next to the web application. Check its status with:

docker compose ps
docker compose logs -f worker

Spotting stuck jobs

If jobs pile up, the worker process is usually stopped, or an external dependency (e.g. the AI provider) is unreachable. The worker logs show the cause.

After an update

Restart the worker after every update so it loads the current code — otherwise it keeps processing jobs with the old version.

Recurring jobs

Some tasks run on a fixed schedule (Solid Queue recurring tasks in config/recurring.yml), such as clearing finished jobs, scheduled account deletions, and documentation sync.

Purging unattached files (GDPR)

An upload creates a file in storage before its form is submitted. If the form is never submitted — or an image is dropped into a text, removed again, and the text is saved — the file stays behind with nothing pointing at it. Rails does not sweep those.

Daily at 05:15, UnattachedBlobPurgeJob deletes every file that has been unattached for more than seven days. Only from that deletion onwards do such files enter the same backup retention windows as everything else — without the job they enter none at all.

Why wait seven days?

Between upload and form submission, a not-yet-attached file is technically indistinguishable from an abandoned one. The waiting period is what keeps the job from deleting a file out from under someone with an open form. That is also why this job has no enforcement variable: it deletes from its first run, and the safety lives in the grace period and a per-run cap.

Preview what would be deleted (deletes nothing):

bin/rails attachments:unattached_report

What the job did is in the application log under [unattached_blob_purge] — one line per deleted file with name, type, size and date, never with content.

Register of deleted attachments (GDPR)

When a file attachment is deleted, Lalabase writes an entry into a register: organisation, filename, type, size, carrier and the moment of deletion. The file's content is never recorded. The register is what proves to the controller that a deletion happened and when (accountability, Art. 5(2) GDPR).

Two recurring jobs maintain it:

Job Time Purpose
AttachmentBackupConfirmationJob 05:45 carries over, from the backup script's own log, when the copy left mirror and remote store
AttachmentDeletionRecordPurgeJob 06:00 deletes entries that are three years old
Why the deadline hangs on the deletion

The three years run from the deletion in production, not from the backup expiry carried over later. That carry-over may never arrive (a failed backup run, a lost log line, a file that never made it into a backup at all). Anchored there, such an entry would have no end date at all — precisely the defect the register exists to document.

The carry-over requires the worker to be able to read the backup script's log. The script creates it root-owned; releasing it through a read group is described in ops/backup/README.md. Without that release the job logs a warning under [deletion_register] instead of silently doing nothing.

Answering a request. When a controller asks for the proof concerning their organisation, two routes give the same answer: the Deletion register page in the admin area, or the command-line report.

bin/rails deletion_register:report                        # everything
ORG=<organisation-uuid> bin/rails deletion_register:report

Both name the date the register started and count entries without a confirmed backup expiry separately. Both belong in the answer: anything deleted before that date is not recorded, and an entry without a carry-over is one whose expiry nobody has confirmed. The organisation list is read from the register itself rather than from the organisations table, because an entry outlives the organisation it concerns.

Embedding cleanup (GDPR)

For semantic search and the AI assistant, Lalabase stores derived embedding vectors. When the underlying content is deleted, the deletion paths normally remove their vectors immediately and synchronously — especially for personal data such as meeting transcripts and analyses.

As an additional safety net, a cleanup job (Embeddings::OrphanGcJob) runs daily at 04:30 and reclaims orphaned vectors whose source content was removed through a technical edge path without synchronous cleanup. As a result, derived vectors of deleted content persist for at most ~24 hours.

Arming deletion (EMBEDDINGS_GC_ENFORCE)

For safety the daily job starts in dry-run: it only logs how many orphaned vectors it would reclaim, deleting nothing. Check the log counts for the first few nights (or DRY_RUN=1 bin/rails "embeddings:gc"). Once the numbers look right, set the environment variable EMBEDDINGS_GC_ENFORCE=1 on the worker process — then the job actually deletes. The manual task below (without DRY_RUN) always deletes for real, regardless of this variable.

Run it manually (e.g. after a large deletion):

bin/rails "embeddings:gc"          # one global sweep across all organisations
DRY_RUN=1 bin/rails "embeddings:gc" # count only, delete nothing

Reclaiming old vector spaces (retention)

When an organisation switches its embedding model, the old vector space stays behind as a roll-back net — you can return to it at any time. How many spaces an organisation keeps is governed by the embedding_retention_limit setting (default: 2, i.e. the active one plus one older). The same nightly job reclaims everything beyond that.

Two spaces it will never touch: the one currently serving search, and one that is being built right now. The space under construction also does not count against the limit — otherwise the very space you are comparing against would disappear while you evaluate two models against each other.

Its own switch (EMBEDDINGS_RETENTION_ENFORCE)

This cleanup has its own switch, deliberately separate from EMBEDDINGS_GC_ENFORCE. The orphan sweep deletes vectors whose source content is provably gone — those are worthless. This one deletes intact vectors: the roll-back net. Throwing one switch must not silently arm the other. Set EMBEDDINGS_RETENTION_ENFORCE=1 on the worker process once the dry-run numbers look right. Careful: the manual task above (bin/rails "embeddings:gc" without DRY_RUN) also reclaims spaces beyond the limit immediately — regardless of either variable. DRY_RUN=1 is the only guard there.

Timing note for rollbacks: While rake embeddings:rollback runs (reactivation + delta reindex — this takes as long as the corpus needs), the target space is temporarily neither active nor configured, and therefore a regular candidate for the retention sweep. Keep rollbacks out of the nightly GC window (04:30) while EMBEDDINGS_RETENTION_ENFORCE=1 is set.

What is NOT reclaimed

Vectors from a model switch that predates space management belong to no registered space. They are deliberately not deleted: the same signature also describes a freshly configured target whose build has not started yet, so deleting on that suspicion would hit a build in progress. The job logs them as UNREGISTERED instead. Seeing rows there is a cleanup decision for you, not a malfunction.