WordloopWordloop
WorkMeeting RecordingTechnical Design DocMilestones05 Live Audio Durability Recovery

Core — Audio Durability

Binary audio frame receipt, segment-based GCS storage, sequence tracking, gap detection, gap upload endpoint, resume handling, transcript degradation signalling, and AudioStoredProgressEvent.

Core — Audio Durability

Owner: Core Engineer Domain: core Complexity: L Prerequisite: Milestone 04 merged

When this slice is complete, Core durably stores every audio frame received during a live recording as aggregated GCS segment objects, tracks the highest contiguous sequence for gap detection, emits AudioStoredProgressEvent so the client can trim its OPFS buffer, accepts gap-recovery uploads via a new REST endpoint, handles WebSocket reconnect via ResumeRecordingCommand, emits GapUploadCompleteEvent on successful gap recovery, and flags the live transcription as degraded when ML reports audio gaps. The recording resource exposes full continuity state — including missing ranges — so clients and diagnostics can verify audio completeness at any point.

Terminology: In this slice, segment always refers to a GCS storage segment — an aggregated object containing multiple contiguous 100 ms audio frames. For transcript segments (speaker-attributed text fragments), see the Transcription contract. The segment manifest is the segment_manifest table that tracks metadata for each GCS storage segment.

Required Capabilities

Segment Storage

  • Core aggregates contiguous binary audio frames (100 ms each) into GCS segment objects with a default target duration of 2 seconds (tunable 1–5 s via config).
  • Segment objects are named meetings/{meeting_id}/segments/{start_seq:08d}-{end_seq:08d}.webm where start_seq and end_seq are the inclusive frame sequence range.
  • Core persists a composition manifest row per segment with: audio_version, start_sequence, end_sequence, started_at_ms, duration_ms, object_name, object_generation, byte_count, crc32c.
  • Duplicate sequences are acknowledged but not appended twice — CRC32C on the incoming frame detects exact duplicates.

Sequence Tracking

  • The recordings row tracks two sequence columns:
    • last_audio_sequence — highest sequence received from the browser (may have gaps below it).
    • last_stored_sequence — highest sequence written to any GCS segment (may have gaps below it). The difference between last_audio_sequence and last_stored_sequence indicates frames received but not yet flushed to GCS.
  • Core computes highest_contiguous_sequence dynamically from the segment manifest — the highest sequence N such that all sequences 1..N are covered by at least one stored segment. This value is not stored in the database; it is computed on read from the segment manifest.
  • Core detects missing ranges by comparing the received sequence space against stored segment coverage and writes a gap_plan JSONB array of [{start_sequence, end_sequence}] to the recordings row.

AudioStoredProgressEvent

  • Core emits AudioStoredProgressEvent on the browser WebSocket every 10 seconds or every 100 audio frames, whichever comes first.
  • The event payload includes: meeting_id, highest_contiguous_sequence, total_frames_stored, total_segments_stored — matching the contract in contracts/recording.mdx.
  • The client uses highest_contiguous_sequence to trim the OPFS shadow buffer during normal operation.

Resume Command Handling

  • Core accepts ResumeRecordingCommand on the browser WebSocket during an active recording. This command is sent by the app after a WebSocket reconnect.
  • The command carries last_client_sequence — the highest sequence the app has written to OPFS.
  • On receipt, Core computes the gap plan by comparing last_client_sequence against the segment manifest to determine which sequences are missing from GCS.
  • Core emits RecordingResumedEvent with highest_contiguous_sequence, missing_ranges[], gap_upload_deadline_at (30 minutes from now), and ml_session_id — matching the contract in contracts/recording.mdx.
  • Core resumes accepting binary audio frames on the same WebSocket after sending the resume event.

Gap Detection API

  • GET /meetings/{id}/recording/missing-chunks returns the range-based gap plan: meeting_id, highest_contiguous_sequence, missing_ranges[], accepted_mime_types, max_chunk_bytes.
  • Missing ranges are computed dynamically from the segment manifest and last_audio_sequence, not from a stale cached value.
  • The endpoint returns 404 if the meeting has never been recorded and 403 if it belongs to another user.

Gap Upload Endpoint

  • POST /meetings/{id}/recording/chunks accepts multipart/form-data with per-part fields: sequence (integer), started_at_ms (integer), duration_ms (integer), mime_type (string), sha256 (hex string), audio (binary).
  • Core verifies the SHA-256 checksum of every uploaded part against the declared sha256. On mismatch, the entire request is rejected with 422 Unprocessable Entity and error code audio_checksum_mismatch.
  • Core de-duplicates by (meeting_id, sequence, checksum) — chunks that match an existing stored segment are accepted but not stored again.
  • On success, Core merges gap chunks into segment objects, advances highest_contiguous_sequence, and returns 200 GapUploadResult: meeting_id, accepted_sequences[], remaining_missing_ranges[], last_contiguous_sequence, audio_version.
  • After processing gap chunks, if contiguous coverage advances, Core emits GapUploadCompleteEvent on the browser WebSocket with highest_contiguous_sequence and remaining_missing_ranges[] — matching the contract in contracts/recording.mdx. The client uses this to trim OPFS and update progress.
  • Returns 409 Conflict if audio_version has been sealed (audio already composed). Returns 413 Payload Too Large if a chunk exceeds max_chunk_bytes (1 MB default).
  • The endpoint requires an Idempotency-Key header per the infrastructure contract.

Transcript Degradation Signalling

  • When Core receives StreamDrainedEvent from ML with gap_count > 0, Core sets transcriptions.is_degraded = true for the live transcription associated with this recording session.
  • This is the only mechanism for setting transcript-level degradation — ML does not tag individual segments.
  • The post-meeting batch transcription (triggered by TranscriptionJob) creates a new transcription with is_degraded = false, assuming the composed audio.webm has full coverage after gap recovery.
  • The UI reads transcriptions.is_degraded to show: "This transcript may have gaps from a network interruption. The full transcript will be available after processing completes."

Recording Resource Extensions

  • GET /meetings/{id}/recording returns the full recording state including: last_received_sequence, highest_contiguous_sequence, missing_ranges[], audio_object_prefix, audio_version, gap_upload_deadline_at, degraded_reasons[].
  • Recording status transitions to awaiting_gap_upload after stop when gaps exist, with a gap_upload_deadline_at 30 minutes after stop.
  • After the deadline expires, Core seals the current audio_version and transitions to composing_audio using whatever segments are available.

Diagnostic Endpoint

  • GET /meetings/{id}/recording/chunk-inventory returns full chunk inventory: total_frames_received, total_segments_stored, highest_contiguous_sequence, gaps[], audio_version, total_bytes, composition_status, first_chunk_at, last_chunk_at.
  • This endpoint uses service auth (not bearer auth) — it is for operator diagnostics, not client use.

Dependencies

  • Milestone 04 (live-transcript-to-final-rebuild) must be merged — specifically the recordings table, RecordingService, and WebSocket binary frame handling.
  • The GCS client must support PUT object writes and GET object reads with CRC32C validation.
  • The recordings schema must be extended with gap_plan JSONB and the sequence tracking columns from the Postgres Schema TDD.

Domain Notes

Schema

The segment_manifest table is new for this milestone. The canonical schema (postgres.mdx) must be updated as part of this slice.

TableKey ColumnsChanges from Milestone 04
recordingsmeeting_id (PK), last_audio_sequence, last_stored_sequence, gap_planAdd last_stored_sequence BIGINT NOT NULL DEFAULT 0, add gap_plan JSONB
segment_manifest (NEW)id (PK), meeting_id (FK), audio_version, start_sequence, end_sequence, started_at_ms, duration_ms, object_name, object_generation, byte_count, crc32c, sha256New table — one row per GCS segment object
CREATE TABLE segment_manifest (
    id               UUID        PRIMARY KEY DEFAULT gen_random_uuid(),
    meeting_id       UUID        NOT NULL REFERENCES meetings (id) ON DELETE CASCADE,
    audio_version    INTEGER     NOT NULL DEFAULT 1,
    start_sequence   BIGINT      NOT NULL,
    end_sequence     BIGINT      NOT NULL,
    started_at_ms    BIGINT      NOT NULL,
    duration_ms      BIGINT      NOT NULL,
    object_name      TEXT        NOT NULL,
    object_generation BIGINT,
    byte_count       BIGINT      NOT NULL,
    crc32c           TEXT        NOT NULL,
    sha256           TEXT,       -- NULL for live segments; populated for gap-recovery segments
    created_at       TIMESTAMPTZ NOT NULL DEFAULT now(),

    UNIQUE (meeting_id, start_sequence, end_sequence)
);

-- Range-scan for gap detection: find all segments for a recording ordered by sequence
CREATE INDEX idx_segment_manifest_meeting_seq
    ON segment_manifest (meeting_id, start_sequence);

Data Flow Summary

Browser ──[binary WS frame]──► Core
                                 │
                                 ├──► Append to in-memory segment buffer
                                 │      └──► On target duration reached:
                                 │             PUT segment object to GCS
                                 │             INSERT segment_manifest row
                                 │             UPDATE recordings.last_stored_sequence
                                 │             Compute highest_contiguous_sequence
                                 │
                                 ├──► Forward to ML (unchanged from M04)
                                 │
                                 └──► Every 10s / 100 frames:
                                        Emit AudioStoredProgressEvent to browser WS

Gap Recovery Flow

Client ──[GET /missing-chunks]──► Core (reads segment_manifest, computes ranges)
       ◄── {highest_contiguous_sequence, missing_ranges[]}

Client ──[POST /chunks multipart]──► Core
                                      │
                                      ├──► Verify SHA-256 per part
                                      ├──► De-duplicate by (meeting_id, seq, checksum)
                                      ├──► Store to GCS segment object
                                      ├──► INSERT segment_manifest row
                                      ├──► Emit GapUploadCompleteEvent (if contiguous coverage advanced)
                                      └──► Return remaining_missing_ranges[]

Resume Flow

App ──[WS reconnect]──► Core
App ──[ResumeRecordingCommand {last_client_sequence}]──► Core
                                                          │
                                                          ├──► Compare last_client_sequence vs segment_manifest
                                                          ├──► Compute missing_ranges[]
                                                          └──► Emit RecordingResumedEvent
                                                                 {highest_contiguous_sequence, missing_ranges[],
                                                                  gap_upload_deadline_at, ml_session_id}

Degradation Model (Core's Role)

Core owns or relays all three degradation layers:

LayerCore's Role
Pipeline Health (Layer 1)Emits RecordingHealthEvent based on ML/GCS/provider state. Relays health transitions to the app via browser WebSocket. Maintains recording.degraded_reasons[].
Audio Completeness (Layer 2)Tracks sequences, stores segments, computes highest_contiguous_sequence, manages gap plan, orchestrates gap upload and composition.
Transcript Coverage (Layer 3)Reads gap_count from StreamDrainedEvent and sets transcriptions.is_degraded = true when gaps occurred.

Test Cases

Happy path

TestLocationAssertion
test_core_stores_audio_segments_during_active_recordingtest_coreStart a recording, send 40 binary audio frames (4 seconds), verify at least 2 GCS segment objects exist in segment_manifest and the objects are accessible in GCS.
test_core_emits_audio_stored_progress_eventtest_coreStart a recording, send 100 binary audio frames, verify the browser WebSocket receives at least one AudioStoredProgressEvent with highest_contiguous_sequence > 0.
test_core_emits_progress_event_after_100_framestest_coreStart a recording, send exactly 100 audio frames in under 10 seconds. Verify exactly one AudioStoredProgressEvent is received — validating the "every 100 frames" cadence trigger fires before the 10-second timer.
test_core_recording_state_includes_sequence_trackingtest_coreStart a recording, send 20 frames, call GET /meetings/{id}/recording, verify response includes last_received_sequence >= 20, highest_contiguous_sequence > 0, and missing_ranges is an empty array.
test_core_missing_chunks_returns_empty_when_contiguoustest_coreStart a recording, send 20 contiguous frames, call GET /meetings/{id}/recording/missing-chunks, verify missing_ranges is [] and highest_contiguous_sequence >= 20.

Gap detection and recovery

TestLocationAssertion
test_core_detects_gap_when_sequences_skiptest_coreStart a recording, send frames 1–10, then send frames 15–20 (skipping 11–14). Call GET /meetings/{id}/recording/missing-chunks. Verify missing_ranges contains [{start: 11, end: 14}].
test_core_gap_upload_fills_missing_rangetest_coreAfter creating a gap (sequences 11–14 missing), upload chunks for sequences 11–14 via POST /meetings/{id}/recording/chunks with valid SHA-256 checksums. Verify response remaining_missing_ranges is [] and last_contiguous_sequence >= 20.
test_core_gap_upload_deduplicates_existing_chunkstest_coreSend frame sequence 5 during recording. Upload the same sequence 5 again via gap upload. Verify response accepted_sequences includes 5 and no duplicate segment exists in segment_manifest.
test_core_gap_upload_rejects_checksum_mismatchtest_coreUpload a gap chunk with an incorrect sha256 value. Verify 422 Unprocessable Entity with error code audio_checksum_mismatch.
test_core_gap_upload_rejects_after_compositiontest_coreComplete a recording through composing_audio, then attempt a gap upload. Verify 409 Conflict indicating audio is already composed.
test_core_emits_gap_upload_complete_eventtest_coreAfter creating a gap (sequences 11–14 missing), upload those chunks via POST /chunks. Verify the browser WebSocket receives GapUploadCompleteEvent with updated highest_contiguous_sequence and remaining_missing_ranges: [].

Resume handling

TestLocationAssertion
test_core_handles_resume_commandtest_coreStart a recording, send frames 1–20. Simulate a WebSocket disconnect and reconnect. Send ResumeRecordingCommand with last_client_sequence: 30 (client has more frames than GCS). Verify Core responds with RecordingResumedEvent containing highest_contiguous_sequence, missing_ranges[] for the gap, gap_upload_deadline_at, and ml_session_id.

Transcript degradation

TestLocationAssertion
test_core_sets_transcription_degraded_on_gap_draintest_coreStart a recording with an associated transcription. Simulate StreamDrainedEvent with gap_count: 3 and total_gap_frames: 45. Verify transcriptions.is_degraded is set to true for the live transcription.

Deadline enforcement

TestLocationAssertion
test_core_seals_audio_version_after_gap_deadlinetest_coreStop a recording with gaps remaining. Advance time past gap_upload_deadline_at. Verify recording status transitions to composing_audio and subsequent gap upload requests are rejected with 409 Conflict.

Error handling

TestLocationAssertion
test_core_gap_upload_requires_idempotency_keytest_corePOST to /meetings/{id}/recording/chunks without an Idempotency-Key header. Verify 400 Bad Request.
test_core_missing_chunks_returns_404_for_unrecorded_meetingtest_coreCall GET /meetings/{id}/recording/missing-chunks for a meeting that has never been recorded. Verify 404.
test_core_chunk_inventory_requires_service_authtest_coreCall GET /meetings/{id}/recording/chunk-inventory with bearer (user) auth. Verify 403 Forbidden.

Completion Checklist

  • Code merged and deployed
  • Bet progress tests pass (./dev test bet meeting-recording)
  • Permanent service tests implemented per testing strategy
  • Code review completed
  • API review completed (new REST endpoints: GET /missing-chunks, POST /chunks, GET /chunk-inventory)
  • Testing review completed
  • Schema review completed (segment_manifest table, recordings column additions — postgres.mdx updated)
  • System documentation updated (architecture, data flows, API reference, database reference, runbooks — as applicable)

On this page