From f01ddb8bc517ccffaff5374087c943c50e1c0f72 Mon Sep 17 00:00:00 2001 From: Shoaib Date: Mon, 7 Sep 2026 22:27:22 +0530 Subject: [PATCH 1/2] render: add --height and --crf; per-segment zoom MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The renderer encodes video twice — once per segment on extract, then again to composite overlays and burn in subtitles. Both CRFs were hardcoded and the scale was an unconditional `scale=1920:-2`, so there was no way to deliver a 4K source at anything but 1080p, and no way to raise the extract CRF that sets the quality ceiling. Symptom: a 4K talking-head render looked visibly soft. Bits-per-pixel turned out to be identical between source and output (0.253 vs 0.255) — the loss was discarded pixels plus a second generation, not bitrate. --height output height (default 1080, 720 with --draft); width follows the source aspect --crf CRF for the extract, i.e. the quality ceiling (default 16 final / 22 preview / 28 draft) The composite encode is now derived as (crf - 2) instead of a hardcoded 18, so the second generation adds minimal further loss, and final defaults to slow/16 rather than fast/20. Re-rendering one edit at --height 2160 --crf 16 gave 0.451 bits/px with no downscale generation. Also adds two EDL fields: ranges[].zoom (+ zoom_x) per-segment push-in, to disguise jump cuts on a static single-camera shot. Deliberately a number, not a filter string: a hardcoded crop is only correct at one output height, and any per-segment dimension mismatch breaks the -c copy concat. Relative ffmpeg expressions do not fix this either — rounding to even dimensions stops crop round-tripping to the exact original size (1920 comes back as 1918). probe_scaled_dims() mirrors the scale expression instead, so the crop is computed from the real post-scale dimensions. Verified with 13 mixed 1.00x/1.06x/1.12x segments all landing at exactly 3840x2160. audio_filter a global audio chain applied per segment BEFORE the 30ms fades, so the fades stay on the true segment edges (Hard Rule 3). Intended for denoise/EQ on noisy location audio. display_dims() factors out the rotation-aware probe that is_portrait_source() was already doing inline. --- SKILL.md | 15 +++++- helpers/render.py | 133 ++++++++++++++++++++++++++++++++++++++++++---- 2 files changed, 136 insertions(+), 12 deletions(-) diff --git a/SKILL.md b/SKILL.md index fa4b776d..42767230 100644 --- a/SKILL.md +++ b/SKILL.md @@ -75,7 +75,7 @@ Helpers (`helpers/transcribe.py`, `helpers/render.py`, etc.) live alongside this - **`transcribe_batch.py `** — 4-worker parallel transcription. Use for multi-take. - **`pack_transcripts.py --edit-dir `** — `transcripts/*.json` → `takes_packed.md` (phrase-level, break on silence ≥ 0.5s). - **`timeline_view.py