diff --git a/AGENTS.md b/AGENTS.md index 6c18823..89050e0 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -84,7 +84,7 @@ $$\Delta_{n+1} = 2 X_m \Delta_n + \Delta_n^2 + \Delta c$$ `Renderer.js` automatically selects the cheapest sufficient WGSL pipeline based on `scale`: -- **Tier 1 (`fs_main_f32`)**: `scale > 1.0e-6` — Native hardware `f32` (~7 decimal digits). +- **Tier 1 (`fs_main_f32`)**: `scale > 1.0e-6` — Native hardware `f32` (~7 decimal digits). At shallow zooms (`scale > 1.0e-4`, where `f32` resolves $c = X_1 + \Delta c$ and $Z_n$ below a pixel), `in_main_bulbs` skips pixels inside the main cardioid and period-2 bulb before the loop, and a Brent periodicity check (with a 128-step capped checkpoint window) terminates periodic interior orbits early. - **Tier 2 (`fs_main_ds`)**: `1.0e-6 >= scale > 1.0e-13` — Knuth/Dekker Double-Single (`vec2`, ~14 decimal digits). - **Tier 3 (`fs_main_qs`)**: `scale <= 1.0e-13` — Quad-Single (`vec4`, ~28 decimal digits). - All three fragment entry points share a single colouring helper `compute_color(i, zn_sq, zn_sp, uv)` in `src/renderer/shader.wgsl`. @@ -94,6 +94,7 @@ $$\Delta_{n+1} = 2 X_m \Delta_n + \Delta_n^2 + \Delta c$$ In `shader.wgsl`, the pixel's total iteration count `i` (`0..uniforms.iter`) is decoupled from the reference orbit lookup index `m` (`0..uniforms.ref_iter`): +- Each iteration loads only $X_{m+1}$ (`reference_orbit[m + 1u]`) and carries it forward as the next step's $X_m$ (resetting to $X_0 = 0$ on rebase). - At each step, the shader computes the true orbit position $Z = X_{m+1} + \Delta$. - Whenever $|Z|^2 < |\Delta|^2$ (the pixel's orbit passes closer to the critical point $X_0 = (0, 0)$ than to the reference orbit $X_{m+1}$) **or** `next_m >= uniforms.ref_iter` (the reference orbit itself escaped at step `ref_iter`), the shader **rebases** the perturbation: $$\Delta \leftarrow Z, \quad m \leftarrow 0$$ @@ -123,7 +124,7 @@ While dragging/zooming or waiting for the WASM worker to compute a new reference 1. `#fractal-bg` (a 2D canvas behind `#fractal`) is updated via `_updateBackgroundCanvas()` (`bgCtx.drawImage(this.canvas, 0, 0)`) whenever a render completes (`state.nextRow >= canvas.height`), including single-slice low-resolution interactive previews. 2. When the viewport transitions from a completed low-res preview to a multi-slice progressive high-res render (`this.canvas.width` / `height` resize), `#fractal-bg` retains the low-res preview underneath `#fractal` so there is never a black flash or stale-frame jump while progressive slices fill in. 3. Progressive high-res slices in `Fractious.frame()` are gated on `device.queue.onSubmittedWorkDone()` and a monotonically increasing `_renderGeneration` token so GPU submissions never saturate the browser compositor queue and user input can preempt in-flight progressive passes within a single frame. -4. Progressive slice height is adaptive: `Renderer.recordSliceTime()` learns per-tier GPU throughput (in worst-case ops per ms) from each slice's submit-to-done time, and `_sliceRows()` sizes the next slice to about `TARGET_SLICE_MS` (10ms). Slower measurements apply immediately, faster ones at most double the estimate, and slices are capped at `MAX_SLICE_BUDGETS` × the fixed worst-case budget to bound stalls when a slice runs into interior. Each full-res render records a `fractious:full-res` User Timing measure for traces. +4. Progressive slice height is adaptive: `Renderer.recordSliceTime()` learns per-tier GPU throughput (in worst-case ops per ms) from each slice's submit-to-done time, and `_sliceRows()` sizes the next slice (`Math.ceil`) to about `TARGET_SLICE_MS` (32ms, above a 16.7ms 60Hz VSync frame on mobile). Slower measurements apply immediately, faster ones at most double the estimate, and slices are capped at `MAX_SLICE_BUDGETS` × the fixed worst-case budget to bound stalls when a slice runs into interior. Each full-res render records a `fractious:full-res` User Timing measure for traces. --- diff --git a/src/Renderer.js b/src/Renderer.js index a3d503e..430d4b4 100644 --- a/src/Renderer.js +++ b/src/Renderer.js @@ -5,9 +5,10 @@ import postShaderCode from './renderer/post.wgsl?raw'; const INTERACTION_MAX_OPS = 20000000; const PROGRESSIVE_MAX_OPS = 25000000; const INTERACTION_SCALE_LIMIT = 0.5; -// Progressive slices are sized from measured GPU throughput to take about this long, -// so each frame does a useful amount of work while input stays responsive. -const TARGET_SLICE_MS = 10; +// Progressive slices are sized from measured GPU throughput to take about this long +// (above a 16.7ms 60Hz VSync frame, since swapchain presentation + IPC has a ~15ms floor +// on mobile GPUs), so each frame does useful work while input stays responsive. +const TARGET_SLICE_MS = 32; // Throughput is measured in worst-case ops (every pixel reaching max iterations), so a // slice can cost more than predicted if it runs into interior. Cap slices at this many // of the old fixed worst-case budgets to bound any single stall. @@ -285,7 +286,7 @@ export class Renderer { opsPerMs * TARGET_SLICE_MS, budget * MAX_SLICE_BUDGETS, ); - return Math.min(remaining, Math.max(1, Math.floor(ops / rowOps))); + return Math.min(remaining, Math.ceil(ops / rowOps)); } _getSliceGeometry(config, state, tier) { diff --git a/src/Renderer.test.js b/src/Renderer.test.js index 92f93d4..cc193f3 100644 --- a/src/Renderer.test.js +++ b/src/Renderer.test.js @@ -27,8 +27,8 @@ describe('Renderer adaptive progressive slices', () => { it('starts from the fixed worst-case budget before anything is measured', () => { const renderer = createRenderer(); - // 25M ops / (1000 px * 10000 iter) = 2.5 rows - expect(renderer._sliceRows(config, progressiveState(), tier)).toBe(2); + // ceil(25M ops / (1000 px * 10000 iter)) = ceil(2.5) = 3 rows + expect(renderer._sliceRows(config, progressiveState(), tier)).toBe(3); }); it('grows slices at most twofold per measurement and shrinks immediately', () => { @@ -45,9 +45,13 @@ describe('Renderer adaptive progressive slices', () => { it('sizes slices to the target time, capped by the stall limit and the rows left', () => { const renderer = createRenderer(); - renderer.throughput.set('QS', 4000000); // 40M ops per 10ms = 4 rows + renderer.throughput.set('QS', 1250000); // 40M ops per 32ms = 4 rows expect(renderer._sliceRows(config, progressiveState(), tier)).toBe(4); + // A 1-row slice finishing under TARGET_SLICE_MS (e.g. 16ms VSync floor) always grows + renderer.recordSliceTime({ tier: 'QS', ops: 10000000 }, 16); + expect(renderer._sliceRows(config, progressiveState(), tier)).toBe(2); + renderer.throughput.set('QS', 1e12); // capped at 16 * 25M ops = 40 rows expect(renderer._sliceRows(config, progressiveState(), tier)).toBe(40); expect(renderer._sliceRows(config, progressiveState(990), tier)).toBe(10); diff --git a/src/renderer/shader.wgsl b/src/renderer/shader.wgsl index 197a75a..41d83eb 100644 --- a/src/renderer/shader.wgsl +++ b/src/renderer/shader.wgsl @@ -372,6 +372,19 @@ fn compute_rotated_uv(uv: vec2) -> vec2 { ); } +// Interior shortcuts are only safe where f32 resolves c and z below a pixel. +const INTERIOR_CHECK_MIN_ZOOM: f32 = 1.0e-4; + +// Closed-form membership of the main cardioid and the period-2 bulb, which hold most +// of the interior (and therefore most of the iterations) in low-zoom views. +fn in_main_bulbs(c: vec2) -> bool { + let x = c.x - 0.25; + let y2 = c.y * c.y; + let q = x * x + y2; + let xp = c.x + 1.0; + return q * (q + x) <= 0.25 * y2 || xp * xp + y2 <= 0.0625; +} + fn compute_color(i: u32, zn_sq: f32, zn_sp: vec2, uv: vec2) -> vec4 { if (i >= uniforms.iter) { return vec4(0.0, 0.0, 0.0, 1.0); @@ -405,6 +418,23 @@ fn fs_main_f32(@location(0) uv: vec2) -> @location(0) vec4 { let c_delta_re = center_x + dx; let c_delta_im = center_y + dy; + + // X_1 = c_ref, so this is the pixel's absolute c. + let check_interior = uniforms.zoom.x > INTERIOR_CHECK_MIN_ZOOM; + let c_ref = reference_orbit[1]; + if (check_interior && + in_main_bulbs(vec2(c_ref.re.x + c_delta_re, c_ref.im.x + c_delta_im))) { + return compute_color(uniforms.iter, 0.0, vec2(0.0), uv); + } + // Brent-style periodicity check: an orbit that returns (within a small fraction of + // a pixel, floored above f32 rounding noise) to a saved point is periodic, so the + // pixel is interior. Cap the checkpoint window so slowly converging boundary orbits + // refresh their saved point regularly. + let period_tol = clamp(uniforms.zoom.x * 2.0e-4, 2.5e-7, 5.0e-5); + let period_tol_sq = period_tol * period_tol; + var period_z = vec2(0.0, 0.0); + var period_len: u32 = 8u; + var period_step: u32 = 0u; var delta_re = 0.0; var delta_im = 0.0; @@ -413,12 +443,12 @@ fn fs_main_f32(@location(0) uv: vec2) -> @location(0) vec4 { var m: u32 = 0u; var zn_sq: f32 = 0.0; var zn_sp = vec2(0.0, 0.0); + // X_m, carried over from the previous iteration's X_{m+1} (X_0 = 0). + var raw_xm = OrbitPoint(vec4(0.0), vec4(0.0)); loop { if (i >= uniforms.iter) { break; } - // Load X_m (Reference in F32) - let raw_xm = reference_orbit[m]; let x_re = raw_xm.re.x; let x_im = raw_xm.im.x; @@ -446,14 +476,30 @@ fn fs_main_f32(@location(0) uv: vec2) -> @location(0) vec4 { zn_sp = vec2(zn_re, zn_im); break; } + + if (check_interior) { + let dz = vec2(zn_re, zn_im) - period_z; + if (dot(dz, dz) < period_tol_sq) { + i = uniforms.iter; + break; + } + period_step = period_step + 1u; + if (period_step == period_len) { + period_step = 0u; + period_len = min(period_len * 2u, 128u); + period_z = vec2(zn_re, zn_im); + } + } let delta_norm_sq = delta_re * delta_re + delta_im * delta_im; if (zn_sq < delta_norm_sq || next_m >= uniforms.ref_iter) { delta_re = zn_re; delta_im = zn_im; m = 0u; + raw_xm = OrbitPoint(vec4(0.0), vec4(0.0)); } else { m = next_m; + raw_xm = raw_xm_next; } } @@ -483,12 +529,12 @@ fn fs_main_ds(@location(0) uv: vec2) -> @location(0) vec4 { var m: u32 = 0u; var zn_sq: f32 = 0.0; var zn_sp = vec2(0.0, 0.0); + // X_m, carried over from the previous iteration's X_{m+1} (X_0 = 0). + var raw_xm = OrbitPoint(vec4(0.0), vec4(0.0)); loop { if (i >= uniforms.iter) { break; } - // Load X_m (Reference in Double-Single) - let raw_xm = reference_orbit[m]; let x_re = vec2(raw_xm.re.x, raw_xm.re.y); let x_im = vec2(raw_xm.im.x, raw_xm.im.y); let xm = ds_complex(x_re, x_im); @@ -524,8 +570,10 @@ fn fs_main_ds(@location(0) uv: vec2) -> @location(0) vec4 { ); delta = dc_add(xm_next, delta); m = 0u; + raw_xm = OrbitPoint(vec4(0.0), vec4(0.0)); } else { m = next_m; + raw_xm = raw_xm_next; } } @@ -554,12 +602,12 @@ fn fs_main_qs(@location(0) uv: vec2) -> @location(0) vec4 { var m: u32 = 0u; var zn_sq: f32 = 0.0; var zn_sp = vec2(0.0, 0.0); + // X_m, carried over from the previous iteration's X_{m+1} (X_0 = 0). + var raw_xm = OrbitPoint(vec4(0.0), vec4(0.0)); loop { if (i >= uniforms.iter) { break; } - // Load X_m (Reference) - let raw_xm = reference_orbit[m]; let xm = qs_complex(raw_xm.re, raw_xm.im); // delta_{n+1} = 2 * X_m * delta_n + delta_n^2 + delta_0 @@ -590,8 +638,10 @@ fn fs_main_qs(@location(0) uv: vec2) -> @location(0) vec4 { let xm_next = qs_complex(raw_xm_next.re, raw_xm_next.im); delta = qc_add(xm_next, delta); m = 0u; + raw_xm = OrbitPoint(vec4(0.0), vec4(0.0)); } else { m = next_m; + raw_xm = raw_xm_next; } }