Skip to content

Add directors-cut: an LLM directs, PixVerse shoots with native audio - #10

Merged
SamyMe merged 2 commits into
masterfrom
edenai-assistant/directors-cut
Sep 29, 2026
Merged

SamyMe merged 2 commits into
masterfrom
edenai-assistant/directors-cut

Conversation

@SamyMe

@SamyMe SamyMe commented Sep 29, 2026

Copy link
Copy Markdown
Member

Summary

This adds a fourth demo, directors-cut/, which shows off PixVerse's video models.

You describe a film in a sentence. An LLM director writes it as a strict JSON shot list: a "style bible" that describes the characters identically in every shot, plus N shots, each with its action, camera move, sound and length. PixVerse shoots every shot with native audio, and ffmpeg cuts them behind a title card. Any shot can be reshot with a new note, which re-renders only that shot and re-cuts the film.

Shooting modes:

  • Every shot at once: parallel, about 60–65 s.
  • Continuity: each shot starts from the last frame of the one before (image-to-video).
  • One multi-shot take: V6 or V5.5 cuts the whole sequence itself in one generation.

Directors: eight LLMs, grouped on the page. Each is pinned to an endpoint that supports JSON-schema output, with a fallback serving the same model elsewhere.

  • Anthropic: Claude Sonnet (the default) and Claude Opus.
  • OpenAI: GPT-6 Astra, Sol and Luna, served by Azure.
  • Open weights: Kimi K3, GLM 5.3, and DeepSeek V4 Pro (served by Nebius).

Page: a light theme (off-white background, white panels, dark video frames), with a live Direct → Shoot → Edit pipeline, storyboard cards, the final cut and a cost table.

Testing (real key, 2026-09-29)

  • 4-shot film, Claude Sonnet directing: 66 s, $0.78, a 16 s film. The same rider, kit and bike appear across all four shots, every shot has native audio, and each follows its camera move.
  • Reshoot of shot 3: 46 s, $0.19, matching the new note.
  • 3-shot film through the page, GPT-6 Sol directing: 60 s, $0.59. The film loads and plays in the page.
  • All eight directors, same premise: each returned a valid 4-shot plan, in 4.5–23 s. GLM 5.3 took 44–90 s.
  • Offline with fakes: all three modes, reshoots, silent and mismatched clips, 16:9 / 21:9 / 9:16, and the page flows.

Notes

These are problems on the providers' side or in the Eden AI platform, not in the demo:

  • OpenAI's own endpoint returns 401 Incorrect API key for GPT-6 (and for OpenAI TTS), with is_byok: false. That's Eden AI's platform OpenAI key. The demo uses Azure, with OpenAI as the fallback.
  • DeepSeek's own endpoint rejects the JSON-schema response format, so the demo uses Nebius.
  • GLM 5.3 can't have its thinking turned off (Z.ai rejects the request), so it's labelled as slower.
  • Eden AI's MCP video_generation tool takes no provider_params, so it can't switch on PixVerse audio or multi-shot. The demo calls the REST API instead.
  • PixVerse options on Eden AI: camera_movement, style and last_frame_image aren't on the allow-list, so camera moves are written into the prompts.

🤖 Generated with Claude Code

SamyMe and others added 2 commits September 29, 2026 14:36
A premise goes to Claude, which writes the film as a strict JSON shot
list: a style bible that describes the characters identically in every
shot, plus N shots with action, camera move, sound and length. PixVerse
renders each shot with native audio, in one of three modes: all at
once, chained (each shot starts from the previous shot's last frame),
or one multi-shot take that PixVerse cuts itself. ffmpeg normalizes the
shots, fills in silence where a shot has no sound, and cuts them behind
a title card. Any shot can be reshot from the page with a new note.

Tested with a real key: a 4-shot, 16 s film in 66 s for $0.78, with a
consistent rider across shots, and a reshoot in 46 s for $0.19. All
three modes, reshoots and the page were tested offline with fakes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ght theme

- The director picker now offers eight LLMs, grouped on the page:
  - Anthropic: Claude Sonnet (default) and Claude Opus.
  - OpenAI: GPT-6 Astra, Sol and Luna.
  - Open weights: Kimi K3, GLM 5.3 and DeepSeek V4 Pro.
  Each is pinned to an endpoint that supports JSON-schema output, with a
  fallback serving the same model elsewhere. GPT-6 is served by Azure,
  because OpenAI's own endpoint returned 401s on the account. DeepSeek
  V4 Pro is served by Nebius, because DeepSeek's own endpoint rejects
  JSON schema. Kimi K3 uses reasoning_effort low (23 s instead of 58 s).
  GLM 5.3 always thinks, so it's labelled as slower.
- Room for reasoning models: max_tokens 8000, a 240 s timeout on
  director calls, and no temperature when reasoning_effort is set.
- The page gets a light theme: off-white background, white panels, dark
  text, and dark frames kept for the video.
- READMEs: a measured "Pick your director" table (every director
  returned a valid 4-shot plan), and the cookbook index lists the new
  directors.

Tested with a real key: all eight directors on the same premise, plus a
full 3-shot film through the page with GPT-6 Sol (60 s, $0.59).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@SamyMe
SamyMe merged commit 444c716 into master Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant