# AI keys (BYOK)

Lalabase follows the **bring-your-own-key** principle: AI features run through your own
provider key. No data flows through shared third-party accounts, and you keep full control
and billing.

## Setting the key

Add your provider's key as an environment variable:

```bash
OPENAI_API_KEY=sk-...
```

Without a key, the AI features stay disabled — the rest of Lalabase keeps working without
any limitation.

## What is sent to the provider

AI functions only send the context they need — for example the text of a ticket to be
summarized. Full database contents are never transmitted.

<div class="docs-callout docs-callout--tip">
  <div class="docs-callout__title">Data sovereignty</div>
  <p>With your own key, the contractual relationship stays between you and your AI provider. Lalabase is only the tool in between.</p>
</div>

## EU providers (GDPR)

Besides OpenAI and Anthropic, you can connect EU providers with data processing in the
EU — via the **OpenAI-compatible** provider in the admin backend. When creating a new
credential, an **EU preset** prefills the endpoint URL and model:

- **Mistral** (France)
- **IONOS AI Model Hub** (Germany)
- **OVHcloud AI Endpoints** (EU)

The presets deliberately pre-fill the **chat** model only. For embeddings, 768- and
1024-dimensional EU models (e.g. `mistral-embed`, `nomic`, `bge-m3`) are configurable
by now — but per organisation via the console rather than through a preset, because
switching an existing corpus triggers a full re-index. Details in "Choosing the
embedding dimension" further down.

<div class="docs-callout docs-callout--tip">
  <div class="docs-callout__title">Data sovereignty</div>
  <p>With an EU provider and your own key, both the contractual relationship and the data processing stay entirely within the EU.</p>
</div>

## Google Gemini via Vertex AI (EU)

Google Gemini can only be connected **GDPR-compliantly through Vertex AI** with an EU
region — not through Google AI Studio (the simple API-key route ends up on a global
endpoint without guaranteed EU data residency). That is why this path is deliberately
built as an **operator setup**: you configure it once, system-wide, in the admin
backend, after which "Google (Gemini)" is available to all organisations — nobody has
to touch a Google Cloud account themselves.

Note up front: Vertex uses **no API key** — it uses a **service account** with a JSON
key. That is the only way to pin the regional EU endpoint. In the commands below,
replace `YOUR_PROJECT` with your GCP project ID.

### Step 1 — Enable the Vertex AI API

Create a **GCP project** and enable the **Vertex AI API**. In the Cloud Console this
entry is now sometimes labelled **"Agent Platform API"** — the service name underneath is
still `aiplatform.googleapis.com`, which is the one you want. Fastest via Cloud Shell:

```bash
gcloud services enable aiplatform.googleapis.com --project=YOUR_PROJECT
```

A billing account must be linked to the project.

### Step 2 — Create a service account with the role

Create a **service account** and grant it the **Vertex AI User** role
(`roles/aiplatform.user`) — as narrow as possible, not "Owner":

```bash
gcloud iam service-accounts create lalabase-vertex --project=YOUR_PROJECT
gcloud projects add-iam-policy-binding YOUR_PROJECT \
  --member="serviceAccount:lalabase-vertex@YOUR_PROJECT.iam.gserviceaccount.com" \
  --role="roles/aiplatform.user"
```

### Step 3 — Generate the JSON key

Service account → **Keys → Add key → JSON**. The file downloads once (it cannot be
retrieved again) — treat it like a password.

<div class="docs-callout docs-callout--warning">
  <div class="docs-callout__title">Pitfall: "Service account key creation is disabled"</div>
  <p>New Google Cloud organisations enforce the policy
  <code>iam.disableServiceAccountKeyCreation</code> by default and block the key download.
  To lift it you need the <strong>Organization Policy Administrator</strong> role
  (<code>roles/orgpolicy.policyAdmin</code>) — "Owner" and "Organization Administrator"
  are <em>not</em> enough. After granting it, reset the policy for the project:</p>
</div>

```bash
gcloud org-policies reset iam.disableServiceAccountKeyCreation --project=YOUR_PROJECT
```

### Step 4 — Choose region and model

Model availability **differs per EU region and changes over time** — verify it before
going live. Enumerate the models in your region directly (status code `200` = available,
`404` = not):

```bash
TOKEN=$(gcloud auth print-access-token); PROJECT=YOUR_PROJECT; LOC=europe-west3
for MODEL in gemini-2.5-flash gemini-2.5-pro gemini-embedding-001; do
  CODE=$(curl -s -o /dev/null -w "%{http_code}" -X POST \
    -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
    "https://${LOC}-aiplatform.googleapis.com/v1/projects/${PROJECT}/locations/${LOC}/publishers/google/models/${MODEL}:generateContent" \
    -d '{"contents":[{"role":"user","parts":[{"text":"ping"}]}]}')
  echo "${MODEL} -> ${CODE}"
done
```

Proven combination (as of 2026-07): **`europe-west3` (Frankfurt) + `gemini-2.5-flash`**
for chat. Older models such as `gemini-2.0-flash` have been retired; the newest
generations (3.x) are currently served in the EU only via the multi-region endpoint,
which Lalabase deliberately does not use for residency reasons.

### Step 5 — Configure in the admin backend

New AI credential, provider **"Google (Gemini)"**:

- **Service account JSON**: paste the complete key (stored encrypted; not an API key).
- **GCP project ID**: your project's ID.
- **Vertex AI region**: the verified EU region, e.g. `europe-west3`. `global` is rejected because it would defeat EU data residency.
- **Model**: the model confirmed available in step 4, e.g. `gemini-2.5-flash`.

Then **set the credential as default** (star button — it then serves every organisation using the provider without its own keys) and switch the organisation to the "Google" provider.

### Step 6 — Verify

Send a chat message. If it fails, the server log names the cause
(`grep "ChatResponseJob: provider" log/production.log`):

| Log line | Cause | Fix |
|---|---|---|
| `provider unauthorized: … HTTP 403` | role missing on the service account | grant `roles/aiplatform.user` (step 2) |
| `provider unauthorized: google token mint failed` | service account JSON corrupted | paste the JSON again (`private_key` with intact line breaks) |
| `provider error: … HTTP 404` | model not available in the region | correct region/model per step 4 |

Which models are actually reachable in which EU region is shown by the
**"Model availability (Vertex AI)"** card on the credential's page in the admin
backend: one click checks every region/model combination free of charge via
`countTokens` — availability differs per EU region and changes over time.

Gemini is an EU-capable path for **embeddings** too: its embedding model truncates
(Matryoshka) cleanly to 1536 dimensions. Since multidim support it is no longer the only
one — 768- and 1024-dimensional EU models from other providers can be configured as well
(see below). Switching the embedding space on an existing corpus always requires a full
re-index of the organisation.

<div class="docs-callout docs-callout--tip">
  <div class="docs-callout__title">Operator setup, not an end-user step</div>
  <p>The Vertex access is configured once by the instance operator. End users do not select it themselves and need no Google Cloud account.</p>
</div>

## Embeddings

For semantic search and the AI assistant, Lalabase generates embeddings stored in PostgreSQL's
`vector` column. These also run through your key.

### Choosing the embedding dimension

An embedding space is the triple of **provider, model and dimension**. Supported are
**768**, **1024**, **1536** and **3072** dimensions, which makes EU models such as
`mistral-embed` (1024), `nomic` (768) and `bge-m3` (1024) usable, as well as the large
`text-embedding-3-large` (3072) and `gemini-embedding-001` (3072).

**3072 requires pgvector 0.7+.** For that size the search index uses a `halfvec` cast,
because pgvector only allows regular vector indexes up to 2000 dimensions. The stored
vector keeps full precision; only the index precision is halved, which in practice does
not change result ordering. If the extension is older than 0.7 the schema cannot be
loaded in the first place — and should an organisation still end up on 3072, search
**fails with an error** rather than getting slow: the query uses the same `halfvec` cast
as the index, and that type does not exist there.

There is deliberately **no UI** for this setting: it takes effect immediately and without
an intermediate stage, so it belongs on the console and in the operator's hands. An
organisation without any configuration stays on 1536, unchanged.

<div class="docs-callout docs-callout--info">
  <div class="docs-callout__title">A switch runs in stages — search stays available throughout</div>
  <p>Lalabase separates <strong>configuration</strong> (where new embeddings are written) from <strong>serving</strong> (which space answers search). Reconfiguring alone does <em>not</em> move serving: the previous space keeps answering while the new one is built in the background. Only the switch (<code>embeddings:flip</code>) moves serving — in a single step.</p>
</div>

The procedure for a switch, using Mistral at 1024 dimensions as the example:

```bash
# 0. Check the starting point — must be green
bin/rails "embeddings:check_alignment[ORG_ID]"

# 1. Reconfigure (Rails console)
#    org.update!(embedding_provider: 'openai_compatible',
#                embedding_api_base: 'https://api.mistral.ai/v1',
#                embedding_access_token: '…',
#                embedding_model: 'mistral-embed',
#                embedding_dims: 1024)
#    → search continues unchanged, in the PREVIOUS space

# 2. Build the target space in the background (no switch yet).
#    Idempotent and restartable; runs through the AI quota.
bin/rails "embeddings:build_target[ORG_ID]"

# 3. Check progress — shows the active space, the target space and the remaining gap
bin/rails "embeddings:status[ORG_ID]"

# 4. Switch. Refused while the target space is incomplete.
bin/rails "embeddings:flip[ORG_ID]"

# 5. Verify — must be green
bin/rails "embeddings:check_alignment[ORG_ID]"
```

After the switch the old space's vectors are removed from the search index (in the
background) but stay around as a way back. In `check_alignment` they show up as `stale`
— that is the normal state and not a finding.

**Rolling back**, if the new model does not convince:

```bash
# SPACE_ID from `embeddings:status`. Reactivates the old space, catches up on the
# changes made since the switch, and switches back.
bin/rails "embeddings:rollback[ORG_ID,SPACE_ID]"
```

The way back works with your own credentials too: endpoint and token are stored per
retained space, so they never have to be re-entered — not even when the organisation's
configuration meanwhile points at a different endpoint.

<div class="docs-callout docs-callout--warning">
  <div class="docs-callout__title">Two pitfalls</div>
  <p><strong>AI must be enabled:</strong> if AI is disabled for the organisation, the build aborts. Enable it first, then build.</p>
  <p><strong>New content during the build:</strong> it is already embedded with the target configuration and therefore only becomes findable after the switch. A short freshness window for the newest content, not a gap in the corpus.</p>
</div>

<div class="docs-callout docs-callout--warning">
  <div class="docs-callout__title">The way back for code symbols only lasts until the next push</div>
  <p>Unlike documents and notes, code symbols are deleted and recreated on every full index. The retained vectors of the old space lose their reference in the process: for code, rolling back after the next push is effectively a full re-index. For all other content it stays cheap.</p>
</div>

## AI packages: the same mechanism, triggered by the customer

Everything above describes the console route. Alongside it there are **curated AI
packages** — bundles maintained by the operator in the admin area, each pairing an embedding
provider with a ladder of chat tiers and carrying a jurisdiction classification. Once an
organisation picks a package, that package's embedding slot decides its vector space —
**no longer the organisation's own configuration**.

What matters operationally: when a package switch moves the vector space, **exactly the
staged switch described above** runs — build the target, keep serving the old space, flip
automatically once the new one is complete. The only difference is who starts it: since
KI-Pakete P4 that is the organisation itself, through a consequence dialog, with no console
and no involvement from you. A watcher performs the flip; running `embeddings:flip` by hand
is not needed.

<div class="docs-callout docs-callout--warning">
  <div class="docs-callout__title">Editing a published package's embedding slot forces re-indexes</div>
  <p>Changing the provider, model or dimension in a package's embedding slot moves the vector space of <em>every</em> organisation sitting on that package. Each of them then rebuilds its entire corpus — and since re-indexes count against the monthly allowance (P4a), that is a spending decision on someone else's account, not merely a configuration change. To offer a different embedding model, create a <strong>new package</strong> and withdraw the old one instead of rebuilding the existing one.</p>
</div>

A withdrawn package (`active: false`) disappears from the catalogue but keeps running
unchanged for organisations on it. They still see their card, but can no longer switch tiers
within it — the way out is another package or "own configuration". So publish a successor
before withdrawing a package.

**Who switched when** is recorded in two forms: the organisation records the timestamp and
person of its *latest* choice in its own columns, and every choice additionally writes an
`Activity` carrying both jurisdiction classifications (before/after) and whether a search
rebuild was started. These entries deliberately trigger **no** notifications.

The same applies to the **on/off switch** since the AI area rebuild: every activation and every
deactivation writes its own `Activity` (`ai_activated` / `ai_deactivated`) with timestamp and
person. Here too the organisation's columns hold only the *latest* activation — the chain exists
nowhere but in these entries.

Since the AI area rebuild that history has a **user interface**: in the **AI** area of the
organisation management, the **Protocol** page lists all four verbs in chronological order,
with timestamp, person and, where a package moved the jurisdiction, both classifications.
It is restricted to organisation administrators and is read-only.

The console route remains useful next to it when you search across organisations or for raw
values from the payload:

```ruby
Activity.where(organisation_id: ORG_ID,
               action: %w[ai_activated ai_deactivated ai_package_selected ai_tier_selected])
        .order(:created_at)
        .pluck(:created_at, :action, :actor_id, :parameters)
```

<div class="docs-callout docs-callout--warning">
  <div class="docs-callout__title">What the trail does not (yet) contain</div>
  <p>Two limits, so neither the page nor the query promises more than it delivers. <strong>First</strong>, the "own configuration" form writes nothing yet: changing the provider or the stored token does not appear in the trail, even though it changes the processor. <strong>Second</strong>, writing the switch entries is deliberately fault-tolerant: if it fails, the activation still goes through and the failure is only logged, because the switch must not fail on account of an audit row. The package choice behaves the other way round and does not commit at all without its entry.</p>
</div>

