WordloopWordloop
WorkMeeting RecordingTechnical Design Doc

Data Flow

Meeting Recording — system context, per-flow sequence diagrams, and boundary inventory.

Data Flow

This document sits between UI Design (which defines what the user sees) and Contracts (which formalise the API shapes). For each screen and interaction, it draws what calls what: which service initiates, which responds, what data crosses each boundary. Read each arrow two ways: it is a contract boundary (what shape the data takes) and a sequencing constraint (downstream cannot build until the upstream contract is published).

System Context

Rendering architecture map...

Part 1: Live Session

Flows that run automatically during an active recording session — audio capture, ML insights, and system resilience.

Flow 1: Start Recording

The user opens the New Meeting ▾ dropdown and selects Start Live Recording. After the browser grants microphone access, the app creates a meeting and initiates the recording session across all three services.

Pre-conditions: The client checks for an active recording session before enabling the button. If one exists, "Start Live Recording" is disabled with a tooltip. Core enforces this server-side — if StartRecordingCommand arrives while a session is already running, it responds with RecordingErrorEvent (session_conflict). If the browser denies microphone access, the app shows a blocking modal with a link to browser audio settings — no data flow occurs.

Rendering architecture map...

Flow 2: Live Audio → Transcription (Lowest Latency Path)

Audio flows from the browser microphone through Core and ML to AssemblyAI. Transcript segments return via the ML WebSocket — the same long-lived connection Core uses to send audio. Audio chunks flow upstream as binary frames, segments and insights flow downstream as CloudEvents text frames.

OPFS shadow buffer: Every audio chunk is simultaneously written to an always-on shadow buffer maintained by a dedicated Web Worker using the Origin Private File System (OPFS) createSyncAccessHandle() API. Each chunk carries a monotonically incrementing sequence number assigned in the browser. This buffer runs unconditionally — it captures audio regardless of Core or GCS connectivity. It is cleared only after Core confirms all chunks are safely in GCS (see Flow 16 and Flow 9).

Aggregated GCS writes: Browser audio is still sequenced at the 100ms frame level for low-latency ML forwarding, but Core does not create one GCS object per 100ms frame. Core aggregates contiguous frames into storage segments (default target: 2 seconds, tunable 1-5 seconds) and stores each segment as a separate GCS object keyed by sequence range: meetings/{id}/segments/{start_seq:08d}-{end_seq:08d}.webm. The segment manifest records the frame sequences, byte counts, CRC32C checksums, object generation, and time offsets contained in each object. This preserves gap recovery by frame sequence while reducing object count, compose operations, lifecycle cleanup, and diagnostic noise. At session end, Core composes storage segments into the final audio.webm (see Flow 9).

Core streams segments directly to the client via WebSocket for minimum latency, and persists them to the database asynchronously in the background. The app distinguishes interim segments from final segments and replaces them in-place when the final version arrives — no layout shift.

Rendering architecture map...

Flow 3: Live Insights Pipeline

Talking points and tasks are extracted by the same LLM query, batched together as a single structured output call. This avoids redundant token spending — the transcript context is loaded once into the prompt cache and both extraction tasks run against it. All insights stream back through the ML WebSocket as CloudEvents text frames, following the dual-write pattern: Core fans out to the browser via WebSocket for latency, and persists to DB asynchronously for durability.

Context management: ML maintains a rolling transcript buffer in memory, appending each finalised segment as it arrives. The full buffer is included in the LLM prompt as cached context. The prompt is always ordered as: [system instructions] [schema] [anchor segments] [recent window] [latest segment]. When the context budget is exceeded, segments are dropped from the oldest non-anchored position — the boundary between the anchor and the recent window — never from the front. Removing from the front would change the content immediately after the static cached prefix, invalidating every transcript token in the cache. Because the anchor only ever grows, the cached region ([system] + [schema] + [anchor]) expands over the course of the meeting and is never invalidated by trimming.

Flow 3a: Live Talking Points & Tasks (Batched — Per Finalised Segment)

On every finalised transcript segment, ML sends the full rolling transcript buffer to the LLM requesting both the latest talking point and any new tasks in a single structured output call. The LLM returns both in one response. Talking points update immediately, and tasks are extracted opportunistically from the same call.

LLM-native task deduplication: The current list of extracted tasks is appended to the dynamic suffix of each prompt. The LLM is instructed to return only tasks that are genuinely new — not already represented in the existing list. This delegates deduplication to the model, which handles paraphrase and semantic overlap naturally without a separate post-processing step.

The prompt is structured for OpenAI prompt caching: [system instructions] [output schema] [anchor segments] forms the stable cached prefix that grows as the session progresses. [recent window] [latest segment] [existing system task list] is the dynamic suffix appended on each call. Placing only system-generated tasks in the dynamic suffix (not user-created tasks and not the cached prefix) keeps the cache hit rate high and avoids making ML depend on every user task mutation. Some overlap with user-created tasks is acceptable in v1; post-meeting task reconciliation removes duplicate or stale system tasks from the final transcript.

Rendering architecture map...

Flow 3b: Live Speaker Identification (Per Diarised Speaker)

AssemblyAI's transcript segments arrive pre-diarised — each segment carries a speaker_label (e.g. speaker_1, speaker_2). ML's job is to resolve each speaker_label to a known Person by matching voice embeddings against enrolled profiles.

Every segment gets a voice embedding. Regardless of whether the speaker has been identified, ML extracts a voice embedding from the segment's audio and stores it on the segment. This happens unconditionally — embeddings are required for post-meeting RAG and future retrieval, not only for speaker matching.

Matching strategy: Speaker matching runs separately, gated on the per-session map speaker_label → { status, person_id?, attempts }. This map lives in ML's memory for the hot path but is mirrored to a meeting_speaker_states table in Postgres on meaningful transitions. On session start — and on reconnect after a pod restart — Core pushes the current speaker states and voice profiles to ML via StreamStartEvent, so ML reconstructs its in-memory map without needing a pull endpoint.

StateBehaviourPersisted?
unmatchedCompare this segment's embedding against all enrolled voice profiles. If confidence exceeds the match threshold → transition to matched. Otherwise, increment attempts and retry on the next segment from this speaker.Attempts tracked in-memory only — an unmatched speaker restarting at 0 on recovery is acceptable.
matchedThe speaker label is locked to a person. All future segments from this speaker are tagged immediately — no further voice comparison needed.Yes — persisted to meeting_speaker_states (status + person reference) on transition.
exhaustedAfter N failed attempts (configurable, e.g. 5 segments), stop comparing for this speaker. The raw speaker_label is preserved. The user can manually resolve it via Flow 7 (speaker labelling).Yes — persisted to meeting_speaker_states on transition.
manualSet by Flow 7 when the user labels a speaker. Takes precedence over voice matching — ML will not attempt to match this speaker regardless of voice similarity.Yes — written synchronously by Core (Flows 7/8) so it is immediately visible on any subsequent pod recovery.
Rendering architecture map...

Part 2: User Mutations

Flows initiated by the user during or after a recording session. All follow the Optimistic Mutation with Echo-Suppressed Streaming pattern: the client updates local state immediately, sends the mutation via REST, and suppresses the returning WebSocket echo.

Flow 4: Notes Auto-Save

The Private Notes scratchpad is the primary surface of the live recording view — it occupies the left column. Notes auto-save continuously with no explicit save button. The app debounces keystrokes and patches the meeting's notes field.

Rendering architecture map...

Flow 5: User Creates Task

The user's task is written via REST (not the streaming path) since it's a user-initiated mutation. Tasks have a description (required), assignee (optional), and due date (optional).

Rendering architecture map...

Flow 6: Task Mutations (Full CRUD)

Flow 5 covers task creation. The UI design specifies a full set of task mutations: edit, delete, toggle completion, nest under other tasks, assign a person, and set a due date. Editing a system-generated task converts it to user-owned.

Rendering architecture map...

Flow 7: User Labels Speaker as Person

When a user identifies "Speaker A" as a known Person (by clicking the speaker label on any transcript segment), the system reassigns all segments from that speaker and records the mapping in meeting_speaker_states as a manual override so that ML respects it immediately on any pod recovery. Voice profile enrichment from the session's embeddings happens during post-meeting processing, not here.

Rendering architecture map...

Flow 8: Create New Person During Speaker Labelling

Flow 7 assumes the person already exists. The UI design says the user can "reassign to a known person or add a new one." When creating a new person, the UI handles this as two sequential operations: first create the person, then assign them to the speaker label using the same endpoint as Flow 7. The speaker-labels endpoint always receives an existing person reference — it has no knowledge of whether that person was just created or long-established. Voice profile enrichment happens during post-meeting processing.

Rendering architecture map...

Part 3: Session End & Post-Meeting

Flows triggered when a recording stops (user-initiated or auto) and the subsequent background processing that upgrades all artefacts to final quality.

Flow 9: Stop Recording

The user presses Stop Recording (or the system auto-stops at the duration limit). The stop sequence is driven by a persisted state machine: active → stopping → draining_ml → awaiting_gap_upload → composing_audio → post_processing → completed|failed. Capture stops first, then ML drains final live segments, then Core collects any remaining OPFS gaps, then Core composes the final audio file, then Core triggers post-meeting processing. Each transition is written to recording_event_history and important transitions publish through the outbox so Core can recover if it crashes mid-stop.

Rendering architecture map...

Flow 10: Duration Warning & Auto-Stop

A configurable maximum recording duration (default: 4 hours) is enforced server-side. Core sends a warning at T-10 minutes and auto-stops at the limit. The auto-stop triggers the same post-meeting pipeline as a user-initiated stop.

Rendering architecture map...

Flow 11: Post-Meeting Processing (Automatic, via Pub/Sub)

Post-meeting processing runs automatically via the shared TranscriptionJob Pub/Sub worker. For live recordings, the job regenerates system tasks from the final transcript, replaces only unedited system-generated tasks from the live session, and preserves all user-created or user-edited tasks. This brings tasks in line with transcript and synthesis: live output is useful immediately, but the final transcript is the source for final artefacts.

The worker:

  1. Batch-transcribes the full audio from GCS (higher accuracy)
  2. Replaces transcript segments with the improved results
  3. Generates headline, summary, topics, and finalises talking points
  4. Reconciles tasks from the final transcript (preserve user-owned tasks; replace unedited system tasks)

The Meeting Summary page shows a subtle progress indicator during re-processing and updates each artefact in real time as it completes.

Rendering architecture map...

Flow 12: Transcription Processing Lifecycle

The transcriptions table tracks processing status through a defined state machine: pending → transcribing → synthesizing → completed (or failed). The client uses this to show re-processing progress on the Meeting Summary page.

Rendering architecture map...

Flow 13: Audio Playback (Signed URL Direct to GCS)

Core generates a short-lived signed URL. The client streams audio directly from Cloud Storage, with standard HTTP range requests for seeking. The audio player appears on the Transcript tab of the Meeting Summary page. If the audio file is still being processed, the endpoint returns 404 and the client retries with exponential backoff.

Rendering architecture map...

Flow 14: Degraded Mode — Layered Resilience

The system has backend and browser-local failure domains. Backend failures degrade gracefully; the OPFS shadow buffer protects audio when local buffering is available and healthy. Core reports recoverable health transitions with RecordingHealthEvent and fatal conditions with RecordingErrorEvent. Recovery is automatic where possible: when a broken backend link restores, Core sends a health transition and the client clears the warning. Gaps in GCS are filled via the gap upload sequence (Flow 16). Browser-local buffer failures are surfaced immediately because they weaken the durability guarantee.

FailureWhat breaksWhat still worksError code
App → Core (WS drops)All commands, events, audio streaming to CoreOPFS shadow buffer captures all audio locallyClient-side onclose
Core → GCS (storage fails)Segment writes — audio gap accumulates in GCSAudio still flows via WS to Core; OPFS captures all audio locally if healthystorage_unavailable
Core → ML (stream fails)Transcription, talking points, tasks, speaker IDAudio→GCS (or OPFS on WS drop), notes auto-saveml_unavailable
ML → AssemblyAITranscript segmentsAudio→GCS, notes, voice embeddingstranscoder_error
ML → OpenAITalking points, live task extraction, summariesTranscript, speaker ID, audio→GCSinsight_warning
Browser OPFS/quotaLocal gap safety netDirect streaming continues while network is healthylocal_buffer_unavailable, local_buffer_full, local_buffer_corrupt_chunk
Rendering architecture map...

The diagram above shows the ml_unavailable path in detail. The remaining failure domains follow the same notification pattern:

App → Core WS drop: The browser's WebSocket onclose event fires. The OPFS shadow buffer captures all audio produced during the outage. On reconnect, the client sends ResumeRecordingCommand and Flow 16 backfills any missing GCS chunks before audio forwarding to ML resumes.

Core → GCS failure: Segment writes fail — Core sends RecordingHealthEvent (degraded: storage_unavailable). Audio continues flowing through the WebSocket; the OPFS shadow buffer captures the gap locally if healthy. On GCS recovery, Core sends RecordingHealthEvent (recovered: storage) with a gap plan and gap chunks are uploaded via Flow 16.

ML → AssemblyAI failure: Transcript segments stop arriving. Core sends RecordingHealthEvent (degraded: transcoder_error). Voice embeddings may be delayed depending on provider state. Audio and notes continue. Missing transcript is rebuilt during post-meeting processing from the full audio in GCS.

ML → OpenAI failure: Talking points and live task extraction stop. Core sends RecordingHealthEvent (degraded: insight_warning). Transcript and speaker ID are unaffected. Missing insights and system tasks are rebuilt during post-meeting processing.

Flow 15: Audio Silence Detection

Two layers detect audio problems: the browser catches microphone issues locally, and Core catches broken streams server-side.

Client-side (primary): The browser monitors the MediaStream via Web Audio API AnalyserNode. If the RMS level falls below a threshold for 10 consecutive seconds, the client shows an inline notice. No server round-trip needed — this is purely a UX signal. Clears automatically when audio levels recover or the first transcript segment arrives.

Server-side (secondary): Core tracks time since the last audio chunk was received on the WebSocket. If no chunks arrive for 10 seconds while a session is active, Core sends a RecordingErrorEvent. This catches the case where the browser believes it's sending audio but the WebSocket stream is silently broken.

Rendering architecture map...

Flow 16: Audio Gap Recovery

When a connectivity gap occurs — either the App→Core WebSocket drops or Core cannot write to GCS — the OPFS shadow buffer accumulates audio produced during the outage. On recovery, Core returns a range-based gap plan: the highest contiguous audio frame sequence plus missing sequence ranges. The client reads those ranges from OPFS and uploads missing chunks via REST. Core verifies checksums, stores or merges them into segment objects, and deduplicates by (meeting_id, sequence, checksum) — chunks already stored are skipped. This same flow runs at session stop time if any gaps remain (see Flow 9).

Rendering architecture map...

Flow 17: Server-Side Inactivity Timeout — New

If no audio chunks arrive for a configurable period (default: 5 minutes) while a recording session is active, Core treats the session as abandoned and triggers the same stop sequence as Flow 9. This covers the case where the user closes their laptop lid, loses power, or otherwise disappears without explicitly stopping — the WebSocket heartbeat timeout (~60 seconds) transitions the connection to closed, but the recording resource would remain in active state indefinitely without this secondary timeout. The inactivity timeout prevents abandoned sessions from blocking the concurrent-session guard.

Rendering architecture map...

Flow 18: Background Tab Audio Continuity — New

Chrome and other browsers aggressively throttle background tabs — JavaScript timers are capped at 1 execution per minute, and some WebSocket activity may be delayed. However, MediaRecorder itself runs on a browser-internal thread and is not throttled when the tab is backgrounded. The critical design choice: all audio chunk processing (sequence numbering, OPFS writes, and WebSocket sends) runs in a dedicated Web Worker, which is exempt from background tab throttling. The main thread only receives notifications for UI updates.

This means audio capture and transmission continue uninterrupted when the user switches to another tab. The page title changes to "● Recording…" so the user can find the tab.

Flow 19: Batched Gap Upload — New

When a large connectivity gap occurs, the client uploads gap chunks in batches rather than one-at-a-time. Core returns missing sequence ranges rather than large per-sequence arrays. The client reads chunks from OPFS in batches of 50, uploads each batch as a single multipart request to POST /meetings/{id}/recording/chunks, and uses remaining_missing_ranges in the response to drive the next batch. A determinate progress indicator shows upload progress on the Meeting Summary page. If the browser closes mid-upload after capture has stopped, the upload can resume from where it left off on next page load until gap_upload_deadline_at. After that deadline, Core composes a degraded audio version from the chunks it has and rejects late chunks for that sealed audio_version.

Flow 20: Soft-Deleted Meeting During Post-Processing — New

Meetings are soft-deleted (flagged with deleted_at, not removed from the database). If a user soft-deletes a meeting while post-meeting processing is running, ML's write-back calls to Core REST will encounter a soft-deleted resource. Core handles this gracefully: write-back endpoints (PUT /transcriptions/{id}/segments, PUT /meetings/{id}/synthesis, PATCH /transcriptions/{id}/status, task reconciliation endpoints) check deleted_at and return 204 No Content without writing new derived artefacts. ML treats this as success (no retry). The post-meeting processing completes silently. This avoids 404 errors, unnecessary retries, DLQ noise, and post-delete derived-content creation.

Flow 21: File Upload & Audio Standardization

When a user uploads an existing audio file (e.g. voice memo, MP3, WAV), Core securely sniffs the magic numbers to validate the format. The file is then passed to an Audio Standardization worker which uses ffmpeg to transcode the audio into a canonical 16kHz mono audio/webm;codecs=opus format. This ensures that the ML pipeline and UI playback always consume a predictable, optimized format without adding latency to the user's upload request.

Rendering architecture map...

Design Decisions

Key architectural choices and their rationale. These are the "why" behind the flows above.

DecisionRationale
Dual-write (WS + async DB)Stream to client via WebSocket for minimum latency (~200ms). Persist to DB asynchronously so a DB hiccup doesn't block the live experience.
Echo-suppressed optimistic mutationsClient updates local state immediately (optimistic), sends via REST, then suppresses the returning WebSocket EntityChanged event using a session identifier. Gives instant UI feedback without double-rendering.
WebSocket for Core↔MLAudio flows upstream as binary frames and insights flow downstream as CloudEvents text frames on the same long-lived WebSocket. Supports bidirectional control events (DrainCommand, BackpressureEvent), replay cursors for reconnection, and speaker state push — capabilities that would require a separate control channel with HTTP streaming. Core acts as a protocol bridge: browser-facing WebSocket on the client side, service-to-service WebSocket on the ML side.
OPFS always-on shadow bufferEvery audio chunk is written to the browser's Origin Private File System via createSyncAccessHandle() in a dedicated Web Worker before (or instead of) being sent to Core. The buffer runs unconditionally — it captures audio regardless of Core or GCS connectivity. This separates audio capture (which must never fail) from transport (which can be retried). The buffer is cleared only after Core confirms all chunks are safely in GCS.
Segment-based GCS writes + hierarchical composeBrowser audio frames remain 100ms for low-latency ML, but Core aggregates frames into storage segments (default 2s, tunable 1-5s) keyed by sequence range (meetings/{id}/segments/{start_seq:08d}-{end_seq:08d}.webm). The manifest preserves frame-level gap recovery while reducing GCS object count and compose fan-in. At session end, Core composes segment objects into final audio.webm using GCS Compose — hierarchically in groups of ≤32 for recordings that exceed GCS's 32-object compose limit.
GCS as the indestructible recordingAudio always reaches GCS eventually, even across connectivity failures, because the OPFS shadow buffer guarantees local capture. Everything else (transcript, insights, tasks) can be rebuilt from the audio during post-meeting processing.
Task reconciliation for live recordingsPost-meeting re-processing replaces transcript segments, regenerates synthesis, and reconciles system-generated tasks from the final transcript. User-created tasks and user-edited system tasks are preserved; unedited system tasks from the live session can be replaced or removed. This keeps final tasks consistent with final transcript quality without clobbering user work.
Transcription status state machinepending → transcribing → synthesizing → completed gives the client granular progress without polling. Each transition fires an EntityChanged event. Fewer states reduce complexity while still distinguishing the two user-visible phases: transcript generation and insight synthesis.
Signed URL with client-side rotationClient streams audio directly from GCS (no Core proxy). Signed URLs expire after 1 hour. Client sets a timer to refresh before expiry for seamless playback.
Dual-layer silence detectionClient-side AnalyserNode catches mic issues instantly (no latency). Server-side chunk timeout catches broken streams the client can't detect. Neither alone covers both cases.
Concurrent session as pre-condition (not a separate flow)The check is a guard on Flow 1, not an independent workflow. Client checks on page load; Core enforces atomically before starting a session.
Batched LLM query (talking points + tasks)A single structured output call extracts both talking points and tasks from the same prompt. The rolling transcript buffer is loaded once into the prompt cache; adding a second extraction task to the same query costs almost nothing vs. a separate call.
Rolling transcript buffer with prompt cachingML appends each finalised segment to an in-memory buffer. Each call to the LLM sends [system instructions] [schema] [anchor] [recent window] [latest segment] [existing task list]. OpenAI caches from the prompt start, so the cached region ([system] + [schema] + [anchor]) grows as the anchor grows. When the context budget is exceeded, segments are dropped from the oldest non-anchored position (between anchor and recent window) — never from the front. Dropping from the front would change the content right after the static prefix, invalidating the entire transcript cache. Dropping from the middle preserves the cached prefix and keeps the most recent context intact. The existing task list in the dynamic suffix lets the LLM deduplicate naturally — it returns only tasks not already in the list, eliminating the need for a separate semantic similarity step.
Speaker state externalised to meeting_speaker_statesThe in-memory speaker_label → state map is mirrored to Postgres on meaningful transitions (matched, exhausted, manual). Core pushes current speaker states and voice profiles to ML on every session start and WebSocket reconnect (via StreamStartEvent), so ML reconstructs its map without a pull endpoint. Attempts are tracked in-memory only; an unmatched speaker restarting at 0 on recovery is acceptable since it retries a bounded number of times before exhausting again. Manual overrides written by Core (Flows 7/8) are immediately visible on any reconnect, so user resolutions are never lost or re-overridden by voice matching.
Progressive speaker matching with lock-onML tries to match each diarised speaker to an enrolled voice profile. Once a confident match is found, the speaker label is locked — no further voice comparison is done for that speaker. Unknown speakers fail fast after a bounded number of attempts. Manual state (set by user labelling) takes precedence and cannot be overridden by voice matching. Voice profile enrichment from session embeddings is deferred to post-meeting processing.
Sticky session affinity (not a backplane)Load balancer routes all WebSocket frames for a session to the same Core pod. No pod-to-pod event routing exists today. This is a known scaling constraint documented separately as a problem statement.
Session not resumable after tab close (v1)OPFS data persists beyond tab close, but the recording session does not. If the user closes the tab, the session ends and post-meeting processing runs on whatever audio reached GCS. Session resume is a future enhancement, captured as a separate problem statement.
WebSocket heartbeat (30s ping/pong)Detects zombie connections in seconds rather than waiting for TCP timeout (minutes). Two missed pongs trigger the client-side OPFS-bridges-the-gap path (Flow 16).
Tiered integrity checksHot-path audio frames use CRC32C for fast corruption detection at high frequency. Higher-integrity flows — OPFS manifest validation, post-stop gap upload batches, composition manifests, and final audio.webm versioning — include SHA-256. CRC32C protects the low-latency stream; SHA-256 protects durable artefacts and replayable recovery operations.
Pre-warm AssemblyAI on mic permissionThe upstream streaming session opens when the browser grants mic access, not when the first audio chunk arrives. Saves ~500–800ms on first-segment latency.
Sequential post-meeting pipelinePost-meeting processing is always sequential: batch transcription first (replaces live segments with higher-accuracy results), then synthesis (headline, summary, topics, talking points). Synthesis depends on the final transcript, so the stages cannot be parallelised. Task extraction is skipped for live recordings (tasks were captured during the session). Each stage updates the transcription status, giving the client granular progress.
Tunable insight cadenceLLM insight extraction is cadence-based by default: run after either a time threshold or final-segment threshold, whichever comes first (default 30s or 5 final segments). The thresholds are delivered in insight_policy, can be tuned per environment, and preserve a fast path for obvious action-item utterances. This protects cost and rate limits while keeping the contract stable.
GCS segment lifecycle: 24h TTL after composeOnce audio.webm is composed, segment objects (segments/{start_seq}-{end_seq}.webm) are no longer needed. A GCS lifecycle rule deletes them 24 hours after composition. The delay provides a safety window for debugging or re-composition.

Boundary Inventory

Every boundary shown in the diagrams above. Each becomes a contract on the Contracts page.

BoundaryFlowsFrom → ToProtocolData shape
Meeting CRUD1, 4App → CoreRESTCreate meeting (live recording); patch meeting notes (echo-suppressed)
Recording commands1, 9App → CoreWebSocketStartRecordingCommand, StopRecordingCommand
Audio streaming2App → Core → MLWS (binary) → ML WS (binary)Raw audio chunks (sequence-numbered, Core enriches with ml_session_id)
Live insights3a, 3bML → Core → AppML WS (CloudEvents) → Browser WSTalking points, tasks, embeddings, speaker matches, speaker exhausted
Task CRUD5, 6App → CoreRESTTask create/update/delete (idempotent create, echo-suppressed; cascading sub-task nesting)
Person creation8App → CoreRESTCreate person
Speaker labels7, 8App → CoreRESTSpeaker-to-person assignment, meeting-scoped (always references an existing person)
Notes auto-save4App → CoreRESTMeeting notes patch (debounced, echo-suppressed)
OPFS gap upload9, 16, 19App → CoreRESTRange-planned, sequence-numbered audio chunks from OPFS shadow buffer; Core deduplicates by sequence and checksum
Recording health14, 16Core → AppWebSocketRecordingHealthEvent for recoverable health transitions; RecordingErrorEvent for fatal/user-actionable failures
Duration warning10Core → AppWebSocketRecordingDurationWarningEvent
Concurrent session check1App → CoreRESTActive session read (read-only guard, no mutation)
Transcription status12ML → Core → AppREST → WebSocketTranscription status transitions + EntityChanged (transcription)
Signed URL13App → Core → GCSREST → GCS signed URLSigned URL fetch (404 while processing, 200 when ready; 1-hour expiry, client-side rotation)
Post-meeting trigger9, 10Core → MLPub/SubTranscriptionJob with audio_version and task reconciliation policy, published after drain completes and audio is composed
Synthesis write-back11ML → CoreRESTTranscript segments replace-all; meeting headline; synthesis artefacts (summary, topics, talking points); system-generated tasks

On this page