WordloopWordloop
Decisions (ADRs)

Voice embeddings are locked to 192 dimensions; audio playback links are per-user signed capabilities

Speaker voice embeddings are ECAPA-TDNN 192-dimension vectors locked to pgvector vector(192), and audio playback proxy links are per-user HMAC capability tokens signed with a key dedicated to that purpose.

0008 — Voice embeddings are locked to 192 dimensions; audio playback links are per-user signed capabilities

Status: Accepted Date: 2026-09-09 Deciders: core platform, ml platform Supersedes: — Superseded by: —

This ADR records two decisions from the same delivery wave that both concern coupling a cross-service contract to a specific numeric value: the dimensionality of a stored vector, and the fields covered by an HMAC signature. Both are the kind of decision that is cheap to get right now and expensive to change later, so both are recorded here rather than left to be re-derived from db/schema.sql and a signer implementation.

Part A — Voice embedding dimension locking

Context

Wordloop matches diarized speaker labels to known people using voice embeddings: ML's ECAPA-TDNN model (onnx_model_dir) produces a fixed-length vector per speaker sample, and Core stores it in pgvector columns (people.voice_vector, transcript_segments.feature_vector) with an HNSW cosine-similarity index for matching. db/schema.sql previously declared these columns as vector(512), left over from an earlier embedding model. The ECAPA-TDNN model ML now runs emits 192-dimension vectors. A pgvector column's dimension is part of its type — vector(512) and vector(192) are not interchangeable, and inserting a 192-d vector into a vector(512) column fails outright rather than silently truncating or padding.

Migrating the column type in place on a database that already holds vector(512) rows is not a simple ALTER TYPE: pg-schema-diff turns the schema change into an ACCESS EXCLUSIVE full-table rewrite, and that rewrite fails with pq: expected 192 dimensions, not 512 on any row still holding the old shape, because there is no meaningful conversion between embeddings from two different models — a 512-d SpeechBrain-era vector and a 192-d ECAPA-TDNN vector do not represent the same feature space.

Decision

db/schema.sql declares both people.voice_vector and transcript_segments.feature_vector as vector(192), and ML's ml_segment_feature_dimensions setting (default 192, sourced from VOICE_EMBEDDING_DIMENSIONS) is the single place that dimension is defined on the ML side. The two are locked together deliberately: Core's column width is not independently tunable, and an embedding model change on the ML side that alters output dimensionality must ship with a matching Core schema migration in the same release — the two services cannot drift on this number without one of them rejecting every embedding the other sends.

Migrating an existing database that still holds vector(512) rows requires an explicit pre-migration runbook rather than a bare schema-diff apply: db/manual/2026-09-09-vector-192.sql, run via cmd/migrate -pre-migration db/manual/2026-09-09-vector-192.sql -allow-hazards up. The runbook drops the two HNSW indexes (rebuilt concurrently by the plan that follows), NULLs any vector whose dimensionality is not 192 (dead data — not comparable across models, and every affected row is re-embeddable from source audio), and is idempotent. See the migration guide for the general workflow and the database reference for current column types.

Consequences

A model swap is a two-service, one-release change. Changing the embedding model's output size requires coordinating a Core schema migration (with the pre-migration runbook pattern established here) alongside the ML model swap — they cannot land independently without one side rejecting the other's data.

Existing databases need an explicit operator step. A database provisioned before this change must run the pre-migration runbook before cmd/migrate up will succeed; a bare up fails with a clear dimension-mismatch error rather than silently corrupting data, and up additionally refuses the resulting ACCESS EXCLUSIVE hazards without -allow-hazards (see the migration guide).

Old embeddings are deliberately discarded, not migrated. Rows with a stale dimension lose their vector and revert to voice_model_status = 'untrained', requiring re-enrollment from source audio. This is accepted because there is no cross-model conversion; keeping mismatched vectors around would only produce silent match failures at query time.

Context

GET /meetings/{id}/audio-url returns a link the App's player fetches audio from. That link must work without shipping the user's Clerk session into a <audio> tag's src, and it must expire. A prior implementation signed the link using SERVICE_AUTH_TOKEN — the same secret ML and CI hold for service-to-service auth — over the request path and an expiry. That reuse meant every holder of the service token (every service, every CI job) could mint a playback link for any user's audio, not just their own: the signature carried no notion of which user it was issued to.

Decision

Audio playback proxy links are signed with a dedicated key, AUDIO_URL_SIGNING_KEY, never SERVICE_AUTH_TOKEN. The signature covers {meeting_id}:{user_id}:{exp}, and the link carries the signed user_id as a uid query parameter — making it a genuine per-user capability rather than a per-resource one. The proxy verifies the signature, then loads the meeting scoped to the signed user (not GetByIDUnscoped), so a correctly signed link naming a user who does not own the meeting streams nothing (404). AUDIO_URL_SIGNING_KEY is required at startup outside APP_ENV=test; Core fails closed (refuses to start) rather than falling back to an empty or shared key. In APP_ENV=test only, an unset key is derived from SERVICE_AUTH_TOKEN via HMAC — a one-way derivation, never the token itself — with a startup warning, so compose/CI stacks need no extra configuration. See Configuration for AUDIO_URL_SIGNING_KEY and RECORDING_AUDIO_URL_TTL_SECONDS.

Consequences

Blast radius of a leaked service token no longer includes user audio. Holding SERVICE_AUTH_TOKEN (every backend service, CI) no longer implies the ability to mint a playback link for arbitrary users' recordings.

A forged or reused link cannot cross users. Signing over user_id and re-checking meeting ownership against the signed user closes the path where a validly-signed link for one user's meeting could be replayed against another user's session.

Operators must provision one more secret. Every environment where Core runs outside APP_ENV=test must set AUDIO_URL_SIGNING_KEY (openssl rand -hex 32), distinct from SERVICE_AUTH_TOKEN. Core will not start without it.

Alternatives considered (Part A)

  • Keep vector(512) and pad/truncate ECAPA-TDNN's 192-d output. Rejected. Padding or truncating a learned embedding does not preserve its similarity properties; matches would be meaningless.
  • A generic vector column with no fixed dimension. Rejected. pgvector allows this, but it removes the type-level guarantee that every stored embedding came from the same model, and the HNSW index requires a fixed dimension to build efficiently.

Alternatives considered (Part B)

  • Keep signing with SERVICE_AUTH_TOKEN but add user_id to the signed payload. Rejected. The blast radius problem is the shared secret itself, not just the missing field — anyone holding the token could still mint a link for any user_id they choose to sign.
  • Short-lived Clerk-issued token passed as the query param instead of an HMAC. Rejected for this delivery as more machinery than the playback use case needs; revisit if audio links need to carry richer, revocable claims.

Debt annotation

Principal: Low for both parts. The dimension lock is one constant shared across two services' config; the signing scheme is a single HMAC helper reused by every link-minting and link-verifying call site.

Interest: Low. Both are covered by regression tests (dimension-mismatch migration probe on a live container; forged/replay/expiry/cross-user signature cases in Core's service tests).

Multiplier: Number of future embedding-consuming services (Part A) and number of future signed-capability link types (Part B). If Wordloop grows a second ML-produced vector type or a second kind of signed capability link, both patterns established here — dimension-locked config plus a migration runbook, and a dedicated signing key plus a subject-scoped payload — should be reused rather than re-invented.

Verification

  • services/wordloop-core/db/schema.sql declares voice_vector vector(192) and feature_vector vector(192).
  • services/wordloop-ml/src/wordloop/config/settings.py — ml_segment_feature_dimensions: int = VOICE_EMBEDDING_DIMENSIONS (192).
  • services/wordloop-core/db/manual/2026-09-09-vector-192.sql — idempotent pre-migration runbook; verified against a live pgvector/pgvector:pg15 container seeded with 512-d rows (fails without it, succeeds with it, no-ops on re-run).
  • Core's fix-wave regression suite (lane core, review 2026-09-09) covers forged, replayed, expired, and cross-user signed-link cases, plus a startup test asserting Core refuses to boot without AUDIO_URL_SIGNING_KEY outside APP_ENV=test.

On this page