Skip to content

Commit bc4a21e

Browse files
jmjavacursoragent
andauthored
Fail closed when timestamps runs with empty segments.all. (#167)
Co-authored-by: Cursor <cursoragent@cursor.com>
1 parent c1dcec0 commit bc4a21e

4 files changed

Lines changed: 20 additions & 11 deletions

File tree

‎AGENTS.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -69,7 +69,7 @@ Commands registered on the **`docgen`** CLI include:
6969
- **`gui`** — desktop window over the same Vue/Flask UI (`pip install 'docgen[gui]'` for pywebview). ``--smoke`` is a headless HTTP check. PyInstaller spec: ``packaging/docgen-gui.spec``. Frozen apps resolve templates/static/benchmark JSON via ``docgen.resources``.
7070
- **`freeze`** — ``docgen freeze`` builds the **`docgen-gui`** onedir (`pip install 'docgen[packaging]'`). Optional ``--smoke`` runs the binary headless. Do not run a full freeze in routine pytest; set ``DOCGEN_FREEZE_SMOKE=1`` for the optional test.
7171
- **`tts`** — text-to-speech for segment files (OpenAI or xAI `/v1/tts`).
72-
- **`timestamps`** — word/segment timing (`timing.json`). Default engine **`local`** aligns the known narration text against the mp3 offline (ffmpeg silencedetect, no API); **`--engine whisper`** uses OpenAI whisper-1 or xAI `/v1/stt` when `ai.provider` is grok. Both emit the same Whisper-shaped blocks. Failed ffmpeg silencedetect raises `AlignmentError` (empty stderr is not treated as full-span speech). OpenAI whisper-1 word/segment `start`/`end` must be finite JSON numbers (bool/NaN raise `AIError`). Grok `/v1/stt` rejects empty word tokens and inverted `end < start` intervals (`AIError`).
72+
- **`timestamps`** — word/segment timing (`timing.json`). Default engine **`local`** aligns the known narration text against the mp3 offline (ffmpeg silencedetect, no API); **`--engine whisper`** uses OpenAI whisper-1 or xAI `/v1/stt` when `ai.provider` is grok. Both emit the same Whisper-shaped blocks. Failed ffmpeg silencedetect raises `AlignmentError` (empty stderr is not treated as full-span speech). OpenAI whisper-1 word/segment `start`/`end` must be finite JSON numbers (bool/NaN raise `AIError`). Grok `/v1/stt` rejects empty word tokens and inverted `end < start` intervals (`AIError`). Empty ``segments.all`` raises ``TimestampError`` (same as TTS) and does not leave a stale ``timing.json`` as success.
7373
- **`image-generate`** — render scene-spec **image elements** (`image:` + `prompt:` boxes) via OpenAI Images or xAI Imagine into the bundle (also runs for missing assets inside `generate-all`).
7474
- **`manim`** — render Manim scenes declared in config.
7575
- **`compose`** — mux narration audio with visual sources via ffmpeg. A ``type: mixed`` row raises ``ComposeError`` if any listed source is missing (no silent subset mux).

‎README.md‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -40,7 +40,9 @@ If you still need the legacy behaviour, pin a pre-removal commit
4040
OpenAI `whisper-1` or xAI `/v1/stt` when the provider is Grok. Both engines
4141
write the same `timing.json` shape. OpenAI whisper-1 word/segment `start`/`end`
4242
must be finite JSON numbers (bool/NaN raise `AIError`). Grok `/v1/stt` rejects
43-
empty word tokens and inverted `end < start` intervals (`AIError`).
43+
empty word tokens and inverted `end < start` intervals (`AIError`). Empty
44+
`segments.all` raises `TimestampError` (same as TTS) and does not leave a
45+
stale `timing.json` as success.
4446
- **Manim animations (default: declarative scene specs)** — primary visual surface.
4547
Prefer **`animations/specs/*.scene.yaml`** via **`docgen scene-spec-generate`**
4648
+ **`scene-compile`**. On **`generate-all`**, if no specs exist yet, the pipeline
@@ -238,7 +240,7 @@ docgen --repo /path/to/your-project generate-all
238240
| `docgen gui [--view benchmark] [--browser] [--smoke]` | Desktop GUI (Vue + Flask). Install `docgen[gui]` for a pywebview window; `--browser` uses the system browser; `--smoke` is a headless HTTP check |
239241
| `docgen freeze [--dist DIR] [--smoke]` | PyInstaller onedir for **`docgen-gui` only** (`pip install 'docgen[packaging]'`). Not the full Manim CLI |
240242
| `docgen tts [--segment 01] [--dry-run]` | Generate TTS audio |
241-
| `docgen timestamps [--engine local\|whisper]` | Extract word/segment timestamps from TTS audio → `timing.json` (default `local`: offline narration-text alignment; `whisper`: OpenAI transcription) |
243+
| `docgen timestamps [--engine local\|whisper]` | Extract word/segment timestamps from TTS audio → `timing.json` (default `local`: offline narration-text alignment; `whisper`: OpenAI transcription). Empty `segments.all` is `TimestampError` (stale `timing.json` is not success) |
242244
| `docgen image-generate [--segment 01 \| --all \| --spec PATH] [--force] [--dry-run] [--model …] [--size …]` | Generate scene-spec image assets (`image:` + `prompt:` boxes) via the OpenAI Images API into the bundle |
243245
| `docgen manim [--scene StackDAGScene]` | Render Manim animations |
244246
| `docgen compose [01 02 03] [--ffmpeg-timeout 900]` | Compose segments (audio + video) |

‎src/docgen/timestamps.py‎

Lines changed: 7 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -229,8 +229,9 @@ def extract_all(self, engine: str | None = None) -> None:
229229
"""Extract timestamps for ``segments.all`` and write timing.json.
230230
231231
Walks configured segment ids via :meth:`Config.find_segment_asset` (no
232-
``*.mp3`` glob). Missing audio for a listed segment is an error. With
233-
no ``segments.all`` entries, existing ``timing.json`` is left unchanged.
232+
``*.mp3`` glob). Missing audio for a listed segment is an error. An
233+
empty ``segments.all`` raises :class:`TimestampError` (same as TTS)
234+
and does not treat an existing ``timing.json`` as success.
234235
235236
Successful runs **merge** stems into the existing file (same as the
236237
wizard per-segment timestamps step) so extra keys not in
@@ -254,8 +255,10 @@ def extract_all(self, engine: str | None = None) -> None:
254255

255256
seg_ids = [str(s) for s in self.config.segments_all]
256257
if not seg_ids:
257-
print("[timestamps] segments.all is empty; leaving timing.json unchanged")
258-
return
258+
raise TimestampError(
259+
"segments.all is empty — add segment ids in docgen.yaml "
260+
"(or hints + yaml-generate) before timestamps"
261+
)
259262

260263
missing: list[str] = []
261264
jobs: list[tuple[str, Path]] = []

‎tests/test_timestamps_local.py‎

Lines changed: 8 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -107,16 +107,20 @@ def test_no_mp3s_fails_when_segments_listed(self, cfg) -> None:
107107
TimestampExtractor(cfg).extract_all()
108108
assert json.loads(out.read_text(encoding="utf-8")) == {"keep": True}
109109

110-
def test_empty_segments_all_leaves_existing_timing_json(self, tmp_path) -> None:
110+
def test_empty_segments_all_stale_timing_json_raises(self, tmp_path) -> None:
111111
(tmp_path / "docgen.yaml").write_text(
112112
yaml.dump({"segments": {"all": []}}), encoding="utf-8"
113113
)
114114
cfg = Config.from_yaml(tmp_path / "docgen.yaml")
115115
out = cfg.animations_dir / "timing.json"
116116
out.parent.mkdir(parents=True, exist_ok=True)
117-
out.write_text('{"keep": true}\n', encoding="utf-8")
118-
TimestampExtractor(cfg).extract_all()
119-
assert json.loads(out.read_text(encoding="utf-8")) == {"keep": True}
117+
stale = {"keep": True, "legacy-stem": {"text": "stale"}}
118+
out.write_text(json.dumps(stale) + "\n", encoding="utf-8")
119+
from docgen.timestamps import TimestampError
120+
121+
with pytest.raises(TimestampError, match="segments.all is empty"):
122+
TimestampExtractor(cfg).extract_all()
123+
assert json.loads(out.read_text(encoding="utf-8")) == stale
120124

121125
def test_heading_only_narration_fails_loud(self, cfg, monkeypatch) -> None:
122126
_fake_audio_env(monkeypatch)

0 commit comments

Comments
 (0)