The format / strict-v0.1

Small files.
Explicit meaning.

A recording’s identity, its timed filtering information and your edits are related pieces. CBFP keeps their roles distinct.

Read the draft

These JSON Schema 2020-12 files describe the structural portion of strict-v0.1, the private SDK’s inspection profile. The underlying MAIN format is choirboy.sidecar/2; a bundle uses manifest version 1.

This is a review draft. The profile does not settle the seven compatibility differences between existing readers. A structurally valid file is not automatically safe to reuse for playback.

The file family

A .cbfp bundle is a ZIP archive containing manifest.json and flat, identity-keyed sidecar files. The MAIN sidecar is named <sha256>.cbfp.json. The manifest identifies the included recordings; each MAIN describes one recording and its filtering spans.

LayerFormatPurpose
MAINchoirboy.sidecar/2Recording metadata and base filtering spans
Editschoirboy.edits/1Personal additions, changes and deletions by persistent span ID
Imported layerchoirboy.import/1A separate imported filtering layer
Captions / transcriptchoirboy.captions/1
choirboy.transcript/1
Caption or timed-word information
Aliaseschoirboy.alias/1Related identity references

The prototype inspects MAIN and bounded bundle contents. It does not merge edit layers or validate all these sidecar formats. Its writer produces one MAIN and a manifest, with no transcript payload or verified signature.

MAIN & timed spans

FieldMeaning in this profile
sha25664 lowercase hexadecimal characters for an exact file identity.
songTitle, artist and integer durationMs. The historical field name also appears for nonmusic recordings.
source / revisionOrigin description and unsigned 32-bit revision counter. Neither establishes trust.
spansUp to 20,000 timed filtering decisions. Each needs a nonempty, unique ID.
analysisRan / completeScanAnalysis and completion claims. Missing values are treated as false; an empty span list is not proof of a complete scan.
editsRefOptional reference to this identity’s .edits.json layer.
landmarkOptional acoustic landmark records, distinct from filtering spans.

A span contains id, word, category, startMs and endMs. Optional defaults are action: "mute", origin: "auto" and confidence: 0.9. A label may contain up to 140 printable ASCII characters.

Times are nonnegative integer milliseconds. A span must start before it ends and, when duration is known and nonzero, end within that duration. The profile bounds time values at 263 − 1; consumers must preserve integers without rounding. Unknown metadata may be retained, but it carries no automatic authority.

JSON Schema’s default is documentation, not an instruction that every validator inserts a value. JSON numbers such as 1.0 also require care: a schema can treat them as integers while the Python inspection profile rejects a parsed float for an integer field.

Identity & timing

An exact file hash covers the complete file’s bytes, including its container. Re-encoding changes that hash. An acoustic fingerprint can help identify a recording across some file differences, but a matching title or a short acoustic match does not establish timing across an entire alternate cut.

Filtering spans use milliseconds. Landmark anchors use frame indices, where one frame is exactly 256000 / 11025 milliseconds. For landmark matching, offset is reference minus query: query_ms = reference_ms − offset_ms.

Chromaprint’s convenience matcher uses a different offset convention. Its raw-frame representation is also separate from landmark records. The prototype does not generate or match either acoustic fingerprint type.

Before reuse, a player needs evidence for the recording, the relevant timing and the scan’s coverage, plus permission to use the filtering record. A schema check alone supplies none of those decisions.

Landmark records

A landmark object has algo: "landmark-1", a count, and records encoded as standard padded base64. Each decoded record is eight bytes: a 32-bit hash followed by a 32-bit anchor frame, both little-endian.

The inspection profile allows 1–8,000 records, requires nondecreasing frames and rejects frames greater than 231 − 1. It rejects partial records, noncanonical base64 and a count that disagrees with the decoded bytes.

The JSON Schema checks the object’s shape and string bounds. Decoding, count agreement and frame checks remain separate. These bytes describe existing landmarks; encoding them does not generate acoustic landmarks from audio.

Walk through an encoded record →

Beyond JSON Schema

Use bounded parsing and semantic validation after structural validation. The prototype’s further checks include:

  • Duplicate JSON keys; unique span IDs; valid intervals and duration agreement.
  • An edit reference, manifest identity and MAIN identity that agree with their filenames.
  • Decoded landmark bytes, canonical padding, actual count and ordered, signed-safe frames.
  • Flat ZIP names, duplicate entries, encrypted or symbolic-link entries, and actual inflated bytes.

The inspection limits are 48 MiB compressed, 32 MiB inflated in total, 4,096 ZIP entries, 2,000 tracks and 20,000 spans per MAIN. Other readers have different limits; a writer’s success does not guarantee their acceptance.

Parsing, server admission, authenticity, transcript authorization and playback eligibility are separate checks. Version 1’s signature metadata is reserved; the prototype does not verify signatures. The schemas do not implement an admission policy or approve content as fully analyzed.