Core — Audio Durability
Binary audio frame receipt, segment-based GCS storage, sequence tracking, gap detection, gap upload endpoint, resume handling, transcript degradation signalling, and AudioStoredProgressEvent.
Core — Audio Durability
Owner: Core Engineer Domain: core Complexity: L Prerequisite: Milestone 04 merged
When this slice is complete, Core durably stores every audio frame received during a live recording as aggregated GCS segment objects, tracks the highest contiguous sequence for gap detection, emits AudioStoredProgressEvent so the client can trim its OPFS buffer, accepts gap-recovery uploads via a new REST endpoint, handles WebSocket reconnect via ResumeRecordingCommand, emits GapUploadCompleteEvent on successful gap recovery, and flags the live transcription as degraded when ML reports audio gaps. The recording resource exposes full continuity state — including missing ranges — so clients and diagnostics can verify audio completeness at any point.
Terminology: In this slice, segment always refers to a GCS storage segment — an aggregated object containing multiple contiguous 100 ms audio frames. For transcript segments (speaker-attributed text fragments), see the Transcription contract. The segment manifest is the
segment_manifesttable that tracks metadata for each GCS storage segment.
Required Capabilities
Segment Storage
- Core aggregates contiguous binary audio frames (100 ms each) into GCS segment objects with a default target duration of 2 seconds (tunable 1–5 s via config).
- Segment objects are named
meetings/{meeting_id}/segments/{start_seq:08d}-{end_seq:08d}.webmwherestart_seqandend_seqare the inclusive frame sequence range. - Core persists a composition manifest row per segment with:
audio_version,start_sequence,end_sequence,started_at_ms,duration_ms,object_name,object_generation,byte_count,crc32c. - Duplicate sequences are acknowledged but not appended twice — CRC32C on the incoming frame detects exact duplicates.
Sequence Tracking
- The
recordingsrow tracks two sequence columns:last_audio_sequence— highest sequence received from the browser (may have gaps below it).last_stored_sequence— highest sequence written to any GCS segment (may have gaps below it). The difference betweenlast_audio_sequenceandlast_stored_sequenceindicates frames received but not yet flushed to GCS.
- Core computes
highest_contiguous_sequencedynamically from the segment manifest — the highest sequenceNsuch that all sequences1..Nare covered by at least one stored segment. This value is not stored in the database; it is computed on read from the segment manifest. - Core detects missing ranges by comparing the received sequence space against stored segment coverage and writes a
gap_planJSONB array of[{start_sequence, end_sequence}]to therecordingsrow.
AudioStoredProgressEvent
- Core emits
AudioStoredProgressEventon the browser WebSocket every 10 seconds or every 100 audio frames, whichever comes first. - The event payload includes:
meeting_id,highest_contiguous_sequence,total_frames_stored,total_segments_stored— matching the contract incontracts/recording.mdx. - The client uses
highest_contiguous_sequenceto trim the OPFS shadow buffer during normal operation.
Resume Command Handling
- Core accepts
ResumeRecordingCommandon the browser WebSocket during an active recording. This command is sent by the app after a WebSocket reconnect. - The command carries
last_client_sequence— the highest sequence the app has written to OPFS. - On receipt, Core computes the gap plan by comparing
last_client_sequenceagainst the segment manifest to determine which sequences are missing from GCS. - Core emits
RecordingResumedEventwithhighest_contiguous_sequence,missing_ranges[],gap_upload_deadline_at(30 minutes from now), andml_session_id— matching the contract incontracts/recording.mdx. - Core resumes accepting binary audio frames on the same WebSocket after sending the resume event.
Gap Detection API
-
GET /meetings/{id}/recording/missing-chunksreturns the range-based gap plan:meeting_id,highest_contiguous_sequence,missing_ranges[],accepted_mime_types,max_chunk_bytes. - Missing ranges are computed dynamically from the segment manifest and
last_audio_sequence, not from a stale cached value. - The endpoint returns
404if the meeting has never been recorded and403if it belongs to another user.
Gap Upload Endpoint
-
POST /meetings/{id}/recording/chunksacceptsmultipart/form-datawith per-part fields:sequence(integer),started_at_ms(integer),duration_ms(integer),mime_type(string),sha256(hex string),audio(binary). - Core verifies the SHA-256 checksum of every uploaded part against the declared
sha256. On mismatch, the entire request is rejected with422 Unprocessable Entityand error codeaudio_checksum_mismatch. - Core de-duplicates by
(meeting_id, sequence, checksum)— chunks that match an existing stored segment are accepted but not stored again. - On success, Core merges gap chunks into segment objects, advances
highest_contiguous_sequence, and returns200 GapUploadResult:meeting_id,accepted_sequences[],remaining_missing_ranges[],last_contiguous_sequence,audio_version. - After processing gap chunks, if contiguous coverage advances, Core emits
GapUploadCompleteEventon the browser WebSocket withhighest_contiguous_sequenceandremaining_missing_ranges[]— matching the contract incontracts/recording.mdx. The client uses this to trim OPFS and update progress. - Returns
409 Conflictifaudio_versionhas been sealed (audio already composed). Returns413 Payload Too Largeif a chunk exceedsmax_chunk_bytes(1 MB default). - The endpoint requires an
Idempotency-Keyheader per the infrastructure contract.
Transcript Degradation Signalling
- When Core receives
StreamDrainedEventfrom ML withgap_count > 0, Core setstranscriptions.is_degraded = truefor the live transcription associated with this recording session. - This is the only mechanism for setting transcript-level degradation — ML does not tag individual segments.
- The post-meeting batch transcription (triggered by
TranscriptionJob) creates a new transcription withis_degraded = false, assuming the composedaudio.webmhas full coverage after gap recovery. - The UI reads
transcriptions.is_degradedto show: "This transcript may have gaps from a network interruption. The full transcript will be available after processing completes."
Recording Resource Extensions
-
GET /meetings/{id}/recordingreturns the full recording state including:last_received_sequence,highest_contiguous_sequence,missing_ranges[],audio_object_prefix,audio_version,gap_upload_deadline_at,degraded_reasons[]. - Recording status transitions to
awaiting_gap_uploadafter stop when gaps exist, with agap_upload_deadline_at30 minutes after stop. - After the deadline expires, Core seals the current
audio_versionand transitions tocomposing_audiousing whatever segments are available.
Diagnostic Endpoint
-
GET /meetings/{id}/recording/chunk-inventoryreturns full chunk inventory:total_frames_received,total_segments_stored,highest_contiguous_sequence,gaps[],audio_version,total_bytes,composition_status,first_chunk_at,last_chunk_at. - This endpoint uses service auth (not bearer auth) — it is for operator diagnostics, not client use.
Dependencies
- Milestone 04 (
live-transcript-to-final-rebuild) must be merged — specifically therecordingstable,RecordingService, and WebSocket binary frame handling. - The GCS client must support
PUTobject writes andGETobject reads with CRC32C validation. - The
recordingsschema must be extended withgap_plan JSONBand the sequence tracking columns from the Postgres Schema TDD.
Domain Notes
Schema
The segment_manifest table is new for this milestone. The canonical schema (postgres.mdx) must be updated as part of this slice.
| Table | Key Columns | Changes from Milestone 04 |
|---|---|---|
recordings | meeting_id (PK), last_audio_sequence, last_stored_sequence, gap_plan | Add last_stored_sequence BIGINT NOT NULL DEFAULT 0, add gap_plan JSONB |
segment_manifest (NEW) | id (PK), meeting_id (FK), audio_version, start_sequence, end_sequence, started_at_ms, duration_ms, object_name, object_generation, byte_count, crc32c, sha256 | New table — one row per GCS segment object |
CREATE TABLE segment_manifest (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
meeting_id UUID NOT NULL REFERENCES meetings (id) ON DELETE CASCADE,
audio_version INTEGER NOT NULL DEFAULT 1,
start_sequence BIGINT NOT NULL,
end_sequence BIGINT NOT NULL,
started_at_ms BIGINT NOT NULL,
duration_ms BIGINT NOT NULL,
object_name TEXT NOT NULL,
object_generation BIGINT,
byte_count BIGINT NOT NULL,
crc32c TEXT NOT NULL,
sha256 TEXT, -- NULL for live segments; populated for gap-recovery segments
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
UNIQUE (meeting_id, start_sequence, end_sequence)
);
-- Range-scan for gap detection: find all segments for a recording ordered by sequence
CREATE INDEX idx_segment_manifest_meeting_seq
ON segment_manifest (meeting_id, start_sequence);Data Flow Summary
Browser ──[binary WS frame]──► Core
│
├──► Append to in-memory segment buffer
│ └──► On target duration reached:
│ PUT segment object to GCS
│ INSERT segment_manifest row
│ UPDATE recordings.last_stored_sequence
│ Compute highest_contiguous_sequence
│
├──► Forward to ML (unchanged from M04)
│
└──► Every 10s / 100 frames:
Emit AudioStoredProgressEvent to browser WSGap Recovery Flow
Client ──[GET /missing-chunks]──► Core (reads segment_manifest, computes ranges)
◄── {highest_contiguous_sequence, missing_ranges[]}
Client ──[POST /chunks multipart]──► Core
│
├──► Verify SHA-256 per part
├──► De-duplicate by (meeting_id, seq, checksum)
├──► Store to GCS segment object
├──► INSERT segment_manifest row
├──► Emit GapUploadCompleteEvent (if contiguous coverage advanced)
└──► Return remaining_missing_ranges[]Resume Flow
App ──[WS reconnect]──► Core
App ──[ResumeRecordingCommand {last_client_sequence}]──► Core
│
├──► Compare last_client_sequence vs segment_manifest
├──► Compute missing_ranges[]
└──► Emit RecordingResumedEvent
{highest_contiguous_sequence, missing_ranges[],
gap_upload_deadline_at, ml_session_id}Degradation Model (Core's Role)
Core owns or relays all three degradation layers:
| Layer | Core's Role |
|---|---|
| Pipeline Health (Layer 1) | Emits RecordingHealthEvent based on ML/GCS/provider state. Relays health transitions to the app via browser WebSocket. Maintains recording.degraded_reasons[]. |
| Audio Completeness (Layer 2) | Tracks sequences, stores segments, computes highest_contiguous_sequence, manages gap plan, orchestrates gap upload and composition. |
| Transcript Coverage (Layer 3) | Reads gap_count from StreamDrainedEvent and sets transcriptions.is_degraded = true when gaps occurred. |
Test Cases
Happy path
| Test | Location | Assertion |
|---|---|---|
test_core_stores_audio_segments_during_active_recording | test_core | Start a recording, send 40 binary audio frames (4 seconds), verify at least 2 GCS segment objects exist in segment_manifest and the objects are accessible in GCS. |
test_core_emits_audio_stored_progress_event | test_core | Start a recording, send 100 binary audio frames, verify the browser WebSocket receives at least one AudioStoredProgressEvent with highest_contiguous_sequence > 0. |
test_core_emits_progress_event_after_100_frames | test_core | Start a recording, send exactly 100 audio frames in under 10 seconds. Verify exactly one AudioStoredProgressEvent is received — validating the "every 100 frames" cadence trigger fires before the 10-second timer. |
test_core_recording_state_includes_sequence_tracking | test_core | Start a recording, send 20 frames, call GET /meetings/{id}/recording, verify response includes last_received_sequence >= 20, highest_contiguous_sequence > 0, and missing_ranges is an empty array. |
test_core_missing_chunks_returns_empty_when_contiguous | test_core | Start a recording, send 20 contiguous frames, call GET /meetings/{id}/recording/missing-chunks, verify missing_ranges is [] and highest_contiguous_sequence >= 20. |
Gap detection and recovery
| Test | Location | Assertion |
|---|---|---|
test_core_detects_gap_when_sequences_skip | test_core | Start a recording, send frames 1–10, then send frames 15–20 (skipping 11–14). Call GET /meetings/{id}/recording/missing-chunks. Verify missing_ranges contains [{start: 11, end: 14}]. |
test_core_gap_upload_fills_missing_range | test_core | After creating a gap (sequences 11–14 missing), upload chunks for sequences 11–14 via POST /meetings/{id}/recording/chunks with valid SHA-256 checksums. Verify response remaining_missing_ranges is [] and last_contiguous_sequence >= 20. |
test_core_gap_upload_deduplicates_existing_chunks | test_core | Send frame sequence 5 during recording. Upload the same sequence 5 again via gap upload. Verify response accepted_sequences includes 5 and no duplicate segment exists in segment_manifest. |
test_core_gap_upload_rejects_checksum_mismatch | test_core | Upload a gap chunk with an incorrect sha256 value. Verify 422 Unprocessable Entity with error code audio_checksum_mismatch. |
test_core_gap_upload_rejects_after_composition | test_core | Complete a recording through composing_audio, then attempt a gap upload. Verify 409 Conflict indicating audio is already composed. |
test_core_emits_gap_upload_complete_event | test_core | After creating a gap (sequences 11–14 missing), upload those chunks via POST /chunks. Verify the browser WebSocket receives GapUploadCompleteEvent with updated highest_contiguous_sequence and remaining_missing_ranges: []. |
Resume handling
| Test | Location | Assertion |
|---|---|---|
test_core_handles_resume_command | test_core | Start a recording, send frames 1–20. Simulate a WebSocket disconnect and reconnect. Send ResumeRecordingCommand with last_client_sequence: 30 (client has more frames than GCS). Verify Core responds with RecordingResumedEvent containing highest_contiguous_sequence, missing_ranges[] for the gap, gap_upload_deadline_at, and ml_session_id. |
Transcript degradation
| Test | Location | Assertion |
|---|---|---|
test_core_sets_transcription_degraded_on_gap_drain | test_core | Start a recording with an associated transcription. Simulate StreamDrainedEvent with gap_count: 3 and total_gap_frames: 45. Verify transcriptions.is_degraded is set to true for the live transcription. |
Deadline enforcement
| Test | Location | Assertion |
|---|---|---|
test_core_seals_audio_version_after_gap_deadline | test_core | Stop a recording with gaps remaining. Advance time past gap_upload_deadline_at. Verify recording status transitions to composing_audio and subsequent gap upload requests are rejected with 409 Conflict. |
Error handling
| Test | Location | Assertion |
|---|---|---|
test_core_gap_upload_requires_idempotency_key | test_core | POST to /meetings/{id}/recording/chunks without an Idempotency-Key header. Verify 400 Bad Request. |
test_core_missing_chunks_returns_404_for_unrecorded_meeting | test_core | Call GET /meetings/{id}/recording/missing-chunks for a meeting that has never been recorded. Verify 404. |
test_core_chunk_inventory_requires_service_auth | test_core | Call GET /meetings/{id}/recording/chunk-inventory with bearer (user) auth. Verify 403 Forbidden. |
Completion Checklist
- Code merged and deployed
- Bet progress tests pass (
./dev test bet meeting-recording) - Permanent service tests implemented per testing strategy
- Code review completed
- API review completed (new REST endpoints:
GET /missing-chunks,POST /chunks,GET /chunk-inventory) - Testing review completed
- Schema review completed (
segment_manifesttable,recordingscolumn additions —postgres.mdxupdated) - System documentation updated (architecture, data flows, API reference, database reference, runbooks — as applicable)
App — Recording Durability UI
OPFS shadow buffer worker, MediaRecorder → OPFS → WebSocket pipeline, worker-centric architecture with MessagePort proxy, bufferedAmount backpressure, gap upload on reconnect/stop, and RecordingHealthEvent banner states.
ML — Transcription Recovery
Gap-tolerant audio input handling, graceful degradation for non-contiguous sequences, backpressure recovery signalling, and session-level gap diagnostics.