AI keys (BYOK)
Lalabase follows the bring-your-own-key principle: AI features run through your own provider key. No data flows through shared third-party accounts, and you keep full control and billing.
Setting the key
Add your provider's key as an environment variable:
OPENAI_API_KEY=sk-...
Without a key, the AI features stay disabled — the rest of Lalabase keeps working without any limitation.
What is sent to the provider
AI functions only send the context they need — for example the text of a ticket to be summarized. Full database contents are never transmitted.
With your own key, the contractual relationship stays between you and your AI provider. Lalabase is only the tool in between.
EU providers (GDPR)
Besides OpenAI and Anthropic, you can connect EU providers with data processing in the EU — via the OpenAI-compatible provider in the admin backend. When creating a new credential, an EU preset prefills the endpoint URL and model:
- Mistral (France)
- IONOS AI Model Hub (Germany)
- OVHcloud AI Endpoints (EU)
The presets deliberately pre-fill the chat model only. For embeddings, 768- and
1024-dimensional EU models (e.g. mistral-embed, nomic, bge-m3) are configurable
by now — but per organisation via the console rather than through a preset, because
switching an existing corpus triggers a full re-index. Details in "Choosing the
embedding dimension" further down.
With an EU provider and your own key, both the contractual relationship and the data processing stay entirely within the EU.
Google Gemini via Vertex AI (EU)
Google Gemini can only be connected GDPR-compliantly through Vertex AI with an EU region — not through Google AI Studio (the simple API-key route ends up on a global endpoint without guaranteed EU data residency). That is why this path is deliberately built as an operator setup: you configure it once, system-wide, in the admin backend, after which "Google (Gemini)" is available to all organisations — nobody has to touch a Google Cloud account themselves.
Note up front: Vertex uses no API key — it uses a service account with a JSON
key. That is the only way to pin the regional EU endpoint. In the commands below,
replace YOUR_PROJECT with your GCP project ID.
Step 1 — Enable the Vertex AI API
Create a GCP project and enable the Vertex AI API. In the Cloud Console this
entry is now sometimes labelled "Agent Platform API" — the service name underneath is
still aiplatform.googleapis.com, which is the one you want. Fastest via Cloud Shell:
gcloud services enable aiplatform.googleapis.com --project=YOUR_PROJECT
A billing account must be linked to the project.
Step 2 — Create a service account with the role
Create a service account and grant it the Vertex AI User role
(roles/aiplatform.user) — as narrow as possible, not "Owner":
gcloud iam service-accounts create lalabase-vertex --project=YOUR_PROJECT
gcloud projects add-iam-policy-binding YOUR_PROJECT \
--member="serviceAccount:lalabase-vertex@YOUR_PROJECT.iam.gserviceaccount.com" \
--role="roles/aiplatform.user"
Step 3 — Generate the JSON key
Service account → Keys → Add key → JSON. The file downloads once (it cannot be retrieved again) — treat it like a password.
New Google Cloud organisations enforce the policy
iam.disableServiceAccountKeyCreation by default and block the key download.
To lift it you need the Organization Policy Administrator role
(roles/orgpolicy.policyAdmin) — "Owner" and "Organization Administrator"
are not enough. After granting it, reset the policy for the project:
gcloud org-policies reset iam.disableServiceAccountKeyCreation --project=YOUR_PROJECT
Step 4 — Choose region and model
Model availability differs per EU region and changes over time — verify it before
going live. Enumerate the models in your region directly (status code 200 = available,
404 = not):
TOKEN=$(gcloud auth print-access-token); PROJECT=YOUR_PROJECT; LOC=europe-west3
for MODEL in gemini-2.5-flash gemini-2.5-pro gemini-embedding-001; do
CODE=$(curl -s -o /dev/null -w "%{http_code}" -X POST \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
"https://${LOC}-aiplatform.googleapis.com/v1/projects/${PROJECT}/locations/${LOC}/publishers/google/models/${MODEL}:generateContent" \
-d '{"contents":[{"role":"user","parts":[{"text":"ping"}]}]}')
echo "${MODEL} -> ${CODE}"
done
Proven combination (as of 2026-07): europe-west3 (Frankfurt) + gemini-2.5-flash
for chat. Older models such as gemini-2.0-flash have been retired; the newest
generations (3.x) are currently served in the EU only via the multi-region endpoint,
which Lalabase deliberately does not use for residency reasons.
Step 5 — Configure in the admin backend
New AI credential, provider "Google (Gemini)":
- Service account JSON: paste the complete key (stored encrypted; not an API key).
- GCP project ID: your project's ID.
- Vertex AI region: the verified EU region, e.g.
europe-west3.globalis rejected because it would defeat EU data residency. - Model: the model confirmed available in step 4, e.g.
gemini-2.5-flash.
Then set the credential as default (star button — it then serves every organisation using the provider without its own keys) and switch the organisation to the "Google" provider.
Step 6 — Verify
Send a chat message. If it fails, the server log names the cause
(grep "ChatResponseJob: provider" log/production.log):
| Log line | Cause | Fix |
|---|---|---|
provider unauthorized: … HTTP 403 |
role missing on the service account | grant roles/aiplatform.user (step 2) |
provider unauthorized: google token mint failed |
service account JSON corrupted | paste the JSON again (private_key with intact line breaks) |
provider error: … HTTP 404 |
model not available in the region | correct region/model per step 4 |
Which models are actually reachable in which EU region is shown by the
"Model availability (Vertex AI)" card on the credential's page in the admin
backend: one click checks every region/model combination free of charge via
countTokens — availability differs per EU region and changes over time.
Gemini is an EU-capable path for embeddings too: its embedding model truncates (Matryoshka) cleanly to 1536 dimensions. Since multidim support it is no longer the only one — 768- and 1024-dimensional EU models from other providers can be configured as well (see below). Switching the embedding space on an existing corpus always requires a full re-index of the organisation.
The Vertex access is configured once by the instance operator. End users do not select it themselves and need no Google Cloud account.
Embeddings
For semantic search and the AI assistant, Lalabase generates embeddings stored in PostgreSQL's
vector column. These also run through your key.
Choosing the embedding dimension
An embedding space is the triple of provider, model and dimension. Supported are
768, 1024, 1536 and 3072 dimensions, which makes EU models such as
mistral-embed (1024), nomic (768) and bge-m3 (1024) usable, as well as the large
text-embedding-3-large (3072) and gemini-embedding-001 (3072).
3072 requires pgvector 0.7+. For that size the search index uses a halfvec cast,
because pgvector only allows regular vector indexes up to 2000 dimensions. The stored
vector keeps full precision; only the index precision is halved, which in practice does
not change result ordering. If the extension is older than 0.7 the schema cannot be
loaded in the first place — and should an organisation still end up on 3072, search
fails with an error rather than getting slow: the query uses the same halfvec cast
as the index, and that type does not exist there.
There is deliberately no UI for this setting: it takes effect immediately and without an intermediate stage, so it belongs on the console and in the operator's hands. An organisation without any configuration stays on 1536, unchanged.
Lalabase separates configuration (where new embeddings are written) from serving (which space answers search). Reconfiguring alone does not move serving: the previous space keeps answering while the new one is built in the background. Only the switch (embeddings:flip) moves serving — in a single step.
The procedure for a switch, using Mistral at 1024 dimensions as the example:
# 0. Check the starting point — must be green
bin/rails "embeddings:check_alignment[ORG_ID]"
# 1. Reconfigure (Rails console)
# org.update!(embedding_provider: 'openai_compatible',
# embedding_api_base: 'https://api.mistral.ai/v1',
# embedding_access_token: '…',
# embedding_model: 'mistral-embed',
# embedding_dims: 1024)
# → search continues unchanged, in the PREVIOUS space
# 2. Build the target space in the background (no switch yet).
# Idempotent and restartable; runs through the AI quota.
bin/rails "embeddings:build_target[ORG_ID]"
# 3. Check progress — shows the active space, the target space and the remaining gap
bin/rails "embeddings:status[ORG_ID]"
# 4. Switch. Refused while the target space is incomplete.
bin/rails "embeddings:flip[ORG_ID]"
# 5. Verify — must be green
bin/rails "embeddings:check_alignment[ORG_ID]"
After the switch the old space's vectors are removed from the search index (in the
background) but stay around as a way back. In check_alignment they show up as stale
— that is the normal state and not a finding.
Rolling back, if the new model does not convince:
# SPACE_ID from `embeddings:status`. Reactivates the old space, catches up on the
# changes made since the switch, and switches back.
bin/rails "embeddings:rollback[ORG_ID,SPACE_ID]"
The way back works with your own credentials too: endpoint and token are stored per retained space, so they never have to be re-entered — not even when the organisation's configuration meanwhile points at a different endpoint.
AI must be enabled: if AI is disabled for the organisation, the build aborts. Enable it first, then build.
New content during the build: it is already embedded with the target configuration and therefore only becomes findable after the switch. A short freshness window for the newest content, not a gap in the corpus.
Unlike documents and notes, code symbols are deleted and recreated on every full index. The retained vectors of the old space lose their reference in the process: for code, rolling back after the next push is effectively a full re-index. For all other content it stays cheap.
AI packages: the same mechanism, triggered by the customer
Everything above describes the console route. Alongside it there are curated AI packages — bundles maintained by the operator in the admin area, each pairing an embedding provider with a ladder of chat tiers and carrying a jurisdiction classification. Once an organisation picks a package, that package's embedding slot decides its vector space — no longer the organisation's own configuration.
What matters operationally: when a package switch moves the vector space, exactly the
staged switch described above runs — build the target, keep serving the old space, flip
automatically once the new one is complete. The only difference is who starts it: since
KI-Pakete P4 that is the organisation itself, through a consequence dialog, with no console
and no involvement from you. A watcher performs the flip; running embeddings:flip by hand
is not needed.
Changing the provider, model or dimension in a package's embedding slot moves the vector space of every organisation sitting on that package. Each of them then rebuilds its entire corpus — and since re-indexes count against the monthly allowance (P4a), that is a spending decision on someone else's account, not merely a configuration change. To offer a different embedding model, create a new package and withdraw the old one instead of rebuilding the existing one.
A withdrawn package (active: false) disappears from the catalogue but keeps running
unchanged for organisations on it. They still see their card, but can no longer switch tiers
within it — the way out is another package or "own configuration". So publish a successor
before withdrawing a package.
Who switched when is recorded in two forms: the organisation records the timestamp and
person of its latest choice in its own columns, and every choice additionally writes an
Activity carrying both jurisdiction classifications (before/after) and whether a search
rebuild was started. These entries deliberately trigger no notifications.
The same applies to the on/off switch since the AI area rebuild: every activation and every
deactivation writes its own Activity (ai_activated / ai_deactivated) with timestamp and
person. Here too the organisation's columns hold only the latest activation — the chain exists
nowhere but in these entries.
Since the AI area rebuild that history has a user interface: in the AI area of the organisation management, the Protocol page lists all four verbs in chronological order, with timestamp, person and, where a package moved the jurisdiction, both classifications. It is restricted to organisation administrators and is read-only.
The console route remains useful next to it when you search across organisations or for raw values from the payload:
Activity.where(organisation_id: ORG_ID,
action: %w[ai_activated ai_deactivated ai_package_selected ai_tier_selected])
.order(:created_at)
.pluck(:created_at, :action, :actor_id, :parameters)
Two limits, so neither the page nor the query promises more than it delivers. First, the "own configuration" form writes nothing yet: changing the provider or the stored token does not appear in the trail, even though it changes the processor. Second, writing the switch entries is deliberately fault-tolerant: if it fails, the activation still goes through and the failure is only logged, because the switch must not fail on account of an audit row. The package choice behaves the other way round and does not commit at all without its entry.