Cut — Post-Production Platform Architecture

Multi-protocol live ingest (SDI/NDI/RTMP/RTSP/SRT/WebRTC) with room-composite egress, Twelve Labs video understanding with transcript + visual search, Whisper speech-to-text + Demucs stem separation, and AI auto-tagging with face detection + visual similarity clustering.

System & Performance Metrics

Live Ingest Latency

SDI < 100ms, NDI < 150ms, RTMP < 2s (buffer)

Twelve Labs Indexing

1x realtime for 30fps content

Whisper Transcription

<0.3x realtime on GPU (batch pipeline)

Demucs Separation

GPU-accelerated 4-stem output (vocals/drums/bass/other)

AIVideo Search

<1s per semantic query across indexed corpus

Auto-Organization

500 clips/min via concurrent processing

Dailies Generation

<5min for 1-hour raw footage (ProRes proxy)

Review Portal

<3s shared link generation with watermark + expiry

1. Multi-Protocol Live Ingest & Egress

ProtocolDirectionUse Case
SDI (3G/6G/12G)IngestBroadcast truck, studio M/E switcher
NDI (HX/NX)Ingest + EgressLocal network cameras, low-latency
RTMPIngestWebcam, OBS, social media streams
RTSPIngestIP cameras, encoders
SRTIngestLong-distance transport, error correction
WebRTCEgressBrowser playback, low-latency
WHIPIngest (passthrough)Direct real-time room ingress

2. AI Video Understanding (Twelve Labs)

Cut integrates Twelve Labs for video indexing, semantic search, transcript generation, and visual detection.

Indexing: point Cut at a video and it builds a searchable index — transcription, visual (scene/shot) embeddings, and on-screen text (OCR) embeddings, with automatic language detection or an explicit language hint.

Video search: natural-language queries ("chef plating dishes") search across visual content, spoken conversation, on-screen text, and logos, with a confidence threshold and result ranking by score or timestamp.

Transcript segments: word-level timestamps with confidence scores and optional speaker diarization, exportable as SRT/VTT captions.

Visual search capabilities:

  • Find frames containing a specific person
  • Find frames matching a scene description ("crowded restaurant kitchen")
  • Find frames with on-screen text ("menu board with prices")
  • Find frames with brand logos
  • Search by visual similarity across the entire video corpus

Example queries: "Find all scenes with chef Gordon plating dishes," "Find any frame where the menu board is visible," "Find all takes where talent says the word 'special'," "Find all instances of the company logo in the lower third."

3. Speech-to-Text & Stem Separation

Cut's audio processing handles transcription via Whisper AI and stem separation via Demucs.

Whisper transcription: choose a model size (tiny through large), transcribe or translate, with optional voice activity detection and per-word timestamps. Output as JSON, SRT, VTT, or plain text. GPU acceleration is used for medium/large models, and transcription runs as a background batch job so editing isn't blocked.

Demucs stem separation: splits a track into 4 stems (vocals/drums/bass/other) using the htdemucs model, exported as WAV. Runs on GPU-accelerated batch workers for fast turnaround, and stem downloads are available via time-limited secure links.

Progress broadcasting: long-running audio jobs stream live progress (downloading → processing → uploading) back to the editor.

4. AI Tagging + Face Detection

Cut's smart organization auto-tags media using AI analysis and provides face detection, visual similarity clustering, and rule-based organization.

Tag sources: manual tagging plus 7 AI-driven sources — detected objects, scene type, spoken transcript, on-screen text (OCR), face detection, dominant color palette, and audio classification (music/speech/noise) — alongside a rule engine for custom logic. Content-type labels (interview/B-roll/promo) are a tag category populated by those sources, not a separate detector.

Tag categories: object, person, location, action, emotion, color, style, content type, technical (resolution/frame rate/codec), and fully custom tags.

Each tag carries a confidence score and, where relevant, timestamps or bounding boxes showing exactly where in the frame and timeline it applies.

Face detection & grouping: faces are clustered across your media, you assign a name to a cluster once, and new footage featuring that person auto-tags going forward.

Rule-based organization: define conditions (filename, duration, date created, and more) and actions (add/remove tag, move to collection, set a custom field) — with a dry-run mode so you can preview a rule before it runs automatically.

Analysis runs as a background pipeline with automatic retry on transient failures, and previously-analyzed media is skipped so re-running organization doesn't waste time.

5. Secure Review Links & Watermarking

Cut's review portal generates secure review links for stakeholders with configurable watermarks, expiry, and download restrictions.

Review sessions carry an expiration, optional view-count limits, and per-session watermark configuration (text, position — center/corner/tiled — opacity, font size). Download settings can restrict quality (proxy/high/original), set an expiration, or cap the number of downloads. Optional view notifications alert you the moment a stakeholder opens the link.

Token security: each review link uses a cryptographically random token, stored in a way that makes the token space computationally infeasible to guess or enumerate.

Watermarking: applied at serve time — not baked into your source media — so the same original file can be shared with different watermarks per recipient. Supports both video and image formats, including watermarking of live HLS streams (in the manifest and segments).

Stakeholder notifications: get notified when someone views a share (who, when, approximate location). The review workflow supports Approve / Request Changes / Reject, with comments timestamped to specific frames.

6. Timeline Model

Cut's timeline manages the nonlinear editing model, track management, and clip manipulation.

Structure: a timeline holds a frame rate (23.976 through 60fps), a resolution (1080p, 4K, and beyond), one or more video/audio/subtitle tracks, and clips referencing your source media with independent in/out points, speed ramping (0.25×–4×), transitions, filters, color grades, opacity, and volume.

Editing operations: trim, split, delete (with ripple or gap-close), extract (lift to clipboard), overwrite, insert (ripple forward), slip, slide, frame-accurate snapping, and transitions (cross-dissolve, wipe, dip-to-black) — all non-destructive.

7. Automated Dailies QC

Cut's dailies QC processes incoming footage for quality control, generating dailies packages for review.

QC checks per clip: exposure, focus, audio levels (peak/average dB), corrupted-frame detection, aspect ratio, color cast, and audio sync — each scored with pass/fail thresholds configurable per project. A ProRes proxy and thumbnail strip are generated automatically for review.

Dailies package: groups all clips from a shoot day/camera card together with their QC results, a review link, and size/transcoding metadata — ready to hand to a DIT or editor for sign-off.

8. Multi-Camera Conference Recording

Cut's conference orchestration manages multi-camera recording for corporate events, conferences, and live broadcasts.

Recording sessions support multiple camera sources, an optional live composite layout, a choice of recording format (individual feeds, composite, or both), an automatic maximum-duration stop, auto-publish to the review portal after the session ends, and per-participant metadata (name, email, role — speaker/panelist/attendee).

Sessions move through setup → live recording → post-processing (proxy generation + QC) → review → long-term archive, supporting everything from a single-camera talk to a full multi-camera conference.

9. Supporting Capabilities

Media versioning: full version history for every media asset — major versions (editorial cuts) and minor versions (color tweaks) — with restore, compare, and merge across versions, integrated with project versioning on the timeline.

Approval workflow: sequential sign-off across stakeholders (e.g., Producer → Director → Network → Legal → Final), with pending/changes-requested/approved/rejected states, per-stakeholder timestamps and comments, deadline tracking with escalation, and review-portal comments surfacing directly into the approval flow.

Built-in effects library: beyond AI-assisted masking, Cut ships a parameterized effects engine covering the categories most editors buy as separate After Effects plugins — included, with no plugin marketplace. 40+ effects across 100+ presets:

Particle systems snow · rain · fire · smoke · dust · sparks Optical lens flare · halation · bloom · chromatic aberration · vignette · film grain · light leak Trapcode-class Shine (volumetric light rays) · Starglow · 3D Stroke · Sound Keys (audio-driven keyframes) Stylize cartoon · pencil sketch · oil paint · watercolor · neon Distortion / warp wave warp · mesh warp · ripple · magnify Blur radial · motion · camera-lens · bokeh Transitions morph · glitch · page turn · advanced dissolve · motion-blur Professional key chroma key · luma key · difference matte · garbage matte Blend modes 18 (multiply, screen, overlay, soft/hard light, color dodge/burn, darken, lighten, difference, exclusion, hue, …)

Effects are stored as parameterized configs and applied as non-destructive adjustment layers on the timeline, with GPU/canvas rendering happening client-side.

Transcoding pipeline: FFmpeg-based, with output formats spanning ProRes 422/4444, DNxHD, H.264, H.265, and VP9. HDR metadata (HDR10, HLG, Dolby Vision) is preserved through the pipeline, along with color space (Rec.709, Rec.2020, DCI-P3) and audio format (AAC, PCM, Dolby Digital, Dolby Atmos). GPU-accelerated workers handle the heavy lifting.

Color grading: LUT-based grading (.cube files), CDL grades (lift/gamma/gain per channel), HDR tone mapping for broadcast delivery, vendor-specific color space conversion, and DaVinci Resolve XML/EDL import for roundtrip workflows.

Magic Mask: AI-assisted rotoscoping with frame-by-frame object tracking, exportable as a matte, alpha channel, or masked video — or applied directly as a layer on the Cut timeline. Presets tune detection and edge handling for people, products, animals, and text.

Composes with the rest of the OS