Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,9 +115,9 @@ In `shader.wgsl`, the pixel's total iteration count `i` (`0..uniforms.iter`) is
- Byte `76`: `ref_iter` (`u32`, valid pre-escape length of `reference_orbit`)
- Bytes `80..95`: `sa_params` (`vec4<f32>`: `[bitcast<f32>(skip_iter), inv_rmax_hi, inv_rmax_lo, 0.0]`)
- Bytes `96..223`: `sa_coeffs` (`array<vec4<f32>, 8>`: Double-Single normalized Series Approximation coefficients $a_1, \dots, a_8$ on the disk $|\Delta c| \le R_{\text{max}}$)
- **`reference_orbit` Storage Buffer (`(calcIter + 1) * 32 bytes`)**:
- Each orbit step `m` occupies `2 × vec4<f32>` (`32 bytes`): `[zx0, zx1, zx2, zx3]` at byte offset `m * 32` (`m * 2u`) and `[zy0, zy1, zy2, zy3]` at byte offset `m * 32 + 16` (`m * 2u + 1u`).
- The orbit ends at the anchor's escape point (or `calcIter`), so `ref_iter` is simply `orbit.byteLength / 32 - 1`.
- **`reference_orbit` Storage Buffer (`(calcIter + 1) * 16 bytes`)**:
- Each orbit step `m` occupies a single `vec4<f32>` (`16 bytes`): `[zx_hi, zx_lo, zy_hi, zy_lo]` at byte offset `m * 16`.
- The orbit ends at the anchor's escape point (or `calcIter`), so `ref_iter` is simply `orbit.byteLength / 16 - 1`.

### 3.4 Dual-Canvas Zero-Latency Interaction (`#fractal-bg` + `#fractal`)

Expand All @@ -126,7 +126,7 @@ While dragging/zooming or waiting for the WASM worker to compute a new reference
1. `#fractal-bg` (a 2D canvas behind `#fractal`) is updated via `_updateBackgroundCanvas()` (`bgCtx.drawImage(this.canvas, 0, 0)`) whenever a render completes (`state.nextRow >= canvas.height`), including single-slice low-resolution interactive previews.
2. When the viewport transitions from a completed low-res preview to a multi-slice progressive high-res render (`this.canvas.width` / `height` resize), `#fractal-bg` retains the low-res preview underneath `#fractal` so there is never a black flash or stale-frame jump while progressive slices fill in.
3. Progressive high-res slices in `Fractious.frame()` are gated on `device.queue.onSubmittedWorkDone()` and a monotonically increasing `_renderGeneration` token so GPU submissions never saturate the browser compositor queue and user input can preempt in-flight progressive passes within a single frame.
4. Progressive slice height is adaptive: `Renderer.recordSliceTime()` learns per-tier GPU throughput (in worst-case ops per ms) from each slice's submit-to-done time, and `_sliceRows()` sizes the next slice (`Math.ceil`) to about `TARGET_SLICE_MS` (32ms, above a 16.7ms 60Hz VSync frame on mobile). Slower measurements apply immediately, faster ones at most double the estimate, and slices are floored at `MIN_SLICE_ROWS` (32 rows, one mobile TBDR hardware tile strip) and capped at `MAX_SLICE_BUDGETS` × the fixed worst-case budget so high-iteration serial latency never collapses slices to 1 row and exterior-to-interior transitions stay bounded. Each full-res render records a `fractious:full-res` User Timing measure for traces.
4. Progressive slice height is adaptive: `Renderer.recordSliceTime()` learns per-tier GPU throughput (in worst-case ops per ms) from each slice's submit-to-done time, and `_sliceRows()` sizes the next slice (`Math.ceil`) to about `TARGET_SLICE_MS` (32ms, above a 16.7ms 60Hz VSync frame on mobile). Slower measurements apply immediately, faster ones at most double the estimate, and slices are floored at `MIN_SLICE_ROWS` (32 rows, one mobile TBDR hardware tile strip) and capped at `Math.max(budget * MAX_SLICE_BUDGETS, MIN_SLICE_ROWS * 2 * rowOps)` so high-iteration serial latency never collapses slices to 1 row, ultra-deep exterior slices can still scale up to `2 × MIN_SLICE_ROWS`, and exterior-to-interior transitions stay bounded. Each full-res render records a `fractious:full-res` User Timing measure for traces.

---

Expand Down
11 changes: 6 additions & 5 deletions src/Renderer.js
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,7 @@ export class Renderer {
});

// initial minimal size, will be updated when orbit arrives
this.referenceOrbitSize = (200 + 1) * 8 * 4;
this.referenceOrbitSize = (200 + 1) * 4 * 4;
this.referenceOrbitBuffer = this.device.createBuffer({
size: Math.max(this.referenceOrbitSize, 16),
usage: GPUBufferUsage.STORAGE | GPUBufferUsage.COPY_DST,
Expand Down Expand Up @@ -188,7 +188,7 @@ export class Renderer {
const requiredSize = orbitArrayBuffer.byteLength;
// The worker truncates the orbit at its escape point, so the last stored
// point index is the valid reference length.
this.referenceOrbitMaxIter = Math.max(0, Math.floor(requiredSize / 32) - 1);
this.referenceOrbitMaxIter = Math.max(0, Math.floor(requiredSize / 16) - 1);
if (sa && sa.length === 36) {
this.saUniformView.set(sa);
this.skipIter = Math.floor(this.uniformDataView.getFloat32(80, true));
Expand Down Expand Up @@ -300,10 +300,11 @@ export class Renderer {
const rowOps = this.canvas.width * this._effectiveIter(config);
const budget = PROGRESSIVE_MAX_OPS * tier.opsMultiplier;
const opsPerMs = this.throughput.get(tier.name) || budget / TARGET_SLICE_MS;
const ops = Math.min(
opsPerMs * TARGET_SLICE_MS,
const maxOps = Math.max(
budget * MAX_SLICE_BUDGETS,
MIN_SLICE_ROWS * 2 * rowOps,
);
const ops = Math.min(opsPerMs * TARGET_SLICE_MS, maxOps);
return Math.min(
remaining,
Math.max(MIN_SLICE_ROWS, Math.ceil(ops / rowOps)),
Expand Down Expand Up @@ -384,7 +385,7 @@ export class Renderer {

const maxBufferIter = Math.max(
0,
Math.floor(this.referenceOrbitSize / 32) - 1,
Math.floor(this.referenceOrbitSize / 16) - 1,
);
const refIter = Math.max(
1,
Expand Down
4 changes: 4 additions & 0 deletions src/Renderer.test.js
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,10 @@ describe('Renderer adaptive progressive slices', () => {

renderer.throughput.set('QS', 1e12); // capped at 16 * 200M ops = 640 rows
expect(renderer._sliceRows(config, progressiveState(), tier)).toBe(640);
// At ultra-high iterations where 16 * budget < 32 rows, fast slices can still scale up to 2 * MIN_SLICE_ROWS (64)
expect(
renderer._sliceRows({ iter: 150000 }, progressiveState(), tier),
).toBe(64);
expect(renderer._sliceRows(config, progressiveState(990), tier)).toBe(10);
});

Expand Down
78 changes: 39 additions & 39 deletions src/renderer/shader.wgsl
Original file line number Diff line number Diff line change
Expand Up @@ -232,13 +232,8 @@ struct Uniforms {
sa_coeffs: array<vec4<f32>, 8>,
};

struct OrbitPoint {
re: vec4<f32>,
im: vec4<f32>,
};

@group(0) @binding(0) var<uniform> uniforms: Uniforms;
@group(0) @binding(1) var<storage, read> reference_orbit: array<OrbitPoint>;
@group(0) @binding(1) var<storage, read> reference_orbit: array<vec4<f32>>;

struct VertexOutput {
@builtin(position) position: vec4<f32>,
Expand Down Expand Up @@ -276,7 +271,7 @@ fn ds_sub(a: vec2<f32>, b: vec2<f32>) -> vec2<f32> {
fn ds_mul(a: vec2<f32>, b: vec2<f32>) -> vec2<f32> {
let p = a.x * b.x;
let e1 = fma(a.x, b.x, -p);
let e2 = a.y * b.x + a.x * b.y;
let e2 = fma(a.y, b.x, a.x * b.y);
let hi = p + e2;
let lo = e1 + e2 - (hi - p);
return vec2<f32>(hi, lo);
Expand Down Expand Up @@ -397,16 +392,19 @@ fn fs_main_f32(@location(0) uv: vec2<f32>) -> @location(0) vec4<f32> {
let check_interior = uniforms.zoom.x > INTERIOR_CHECK_MIN_ZOOM;
let c_ref = reference_orbit[1];
if (check_interior &&
in_main_bulbs(vec2<f32>(c_ref.re.x + c_delta_re, c_ref.im.x + c_delta_im))) {
in_main_bulbs(vec2<f32>(c_ref.x + c_delta_re, c_ref.z + c_delta_im))) {
return compute_color(uniforms.iter, 0.0, vec2<f32>(0.0), uv);
}
// Brent-style periodicity check: an orbit that returns (within a small fraction of
// a pixel, floored above f32 rounding noise) to a saved point is periodic, so the
// pixel is interior. Cap the checkpoint window so slowly converging boundary orbits
// refresh their saved point regularly.
let period_tol = clamp(uniforms.zoom.x * 2.0e-4, 2.5e-7, 5.0e-5);
// Brent-style periodicity check: at shallow zooms compare absolute Z_n; at deeper
// zooms compare perturbations delta_n when next_m == period_m (so X_m cancels out).
let period_tol = select(
uniforms.zoom.x * 2.0e-4,
clamp(uniforms.zoom.x * 2.0e-4, 2.5e-7, 5.0e-5),
check_interior
);
let period_tol_sq = period_tol * period_tol;
var period_z = vec2<f32>(0.0, 0.0);
var period_m: u32 = 0u;
var period_len: u32 = 8u;
var period_step: u32 = 0u;

Expand All @@ -430,8 +428,8 @@ fn fs_main_f32(@location(0) uv: vec2<f32>) -> @location(0) vec4<f32> {
i = skip_iter;
m = skip_iter;
let init_xm = reference_orbit[skip_iter];
x_re = init_xm.re.x;
x_im = init_xm.im.x;
x_re = init_xm.x;
x_im = init_xm.z;
}

loop {
Expand All @@ -447,8 +445,8 @@ fn fs_main_f32(@location(0) uv: vec2<f32>) -> @location(0) vec4<f32> {

let next_m = m + 1u;
let raw_xm_next = reference_orbit[next_m];
let nx_re = raw_xm_next.re.x;
let nx_im = raw_xm_next.im.x;
let nx_re = raw_xm_next.x;
let nx_im = raw_xm_next.z;

// Compute zn in single precision
let zn_re = nx_re + delta_re;
Expand All @@ -462,18 +460,20 @@ fn fs_main_f32(@location(0) uv: vec2<f32>) -> @location(0) vec4<f32> {
break;
}

if (check_interior) {
let dz = vec2<f32>(zn_re, zn_im) - period_z;
let cur_z = select(vec2<f32>(delta_re, delta_im), vec2<f32>(zn_re, zn_im), check_interior);
if (check_interior || next_m == period_m) {
let dz = cur_z - period_z;
if (dot(dz, dz) < period_tol_sq) {
i = uniforms.iter;
break;
}
period_step = period_step + 1u;
if (period_step == period_len) {
period_step = 0u;
period_len = min(period_len * 2u, 128u);
period_z = vec2<f32>(zn_re, zn_im);
}
}
period_step = period_step + 1u;
if (period_step == period_len) {
period_step = 0u;
period_len = min(period_len * 2u, 128u);
period_z = cur_z;
period_m = next_m;
}

let delta_norm_sq = fma(delta_re, delta_re, delta_im * delta_im);
Expand Down Expand Up @@ -526,10 +526,7 @@ fn fs_main_ds(@location(0) uv: vec2<f32>) -> @location(0) vec4<f32> {
i = skip_iter;
m = skip_iter;
let init_xm = reference_orbit[skip_iter];
xm = ds_complex(
vec2<f32>(init_xm.re.x, init_xm.re.y),
vec2<f32>(init_xm.im.x, init_xm.im.y)
);
xm = ds_complex(init_xm.xy, init_xm.zw);
}

loop {
Expand All @@ -541,10 +538,7 @@ fn fs_main_ds(@location(0) uv: vec2<f32>) -> @location(0) vec4<f32> {

let next_m = m + 1u;
let raw_xm_next = reference_orbit[next_m];
let xm_next = ds_complex(
vec2<f32>(raw_xm_next.re.x, raw_xm_next.re.y),
vec2<f32>(raw_xm_next.im.x, raw_xm_next.im.y)
);
let xm_next = ds_complex(raw_xm_next.xy, raw_xm_next.zw);

// Compute zn in single precision
let zn_re = xm_next.re.x + delta.re.x;
Expand Down Expand Up @@ -595,7 +589,7 @@ fn fs_main_qs(@location(0) uv: vec2<f32>) -> @location(0) vec4<f32> {
var zn_sq: f32 = 0.0;
var zn_sp = vec2<f32>(0.0, 0.0);
// X_m, carried over from the previous iteration's X_{m+1} (X_0 = 0).
var raw_xm = OrbitPoint(vec4<f32>(0.0), vec4<f32>(0.0));
var raw_xm = vec4<f32>(0.0);

let skip_iter = u32(uniforms.sa_params.x);
let u_norm = vec2<f32>(c_delta_re.x, c_delta_im.x) * uniforms.sa_params.y;
Expand All @@ -614,15 +608,18 @@ fn fs_main_qs(@location(0) uv: vec2<f32>) -> @location(0) vec4<f32> {
if (i >= uniforms.iter) { break; }

// delta_{n+1} = (2 * X_m + delta_n) * delta_n + delta_0
let w = qs_complex(qs_add(raw_xm.re * 2.0, delta.re), qs_add(raw_xm.im * 2.0, delta.im));
let w = qs_complex(
qs_add(vec4<f32>(raw_xm.xy * 2.0, 0.0, 0.0), delta.re),
qs_add(vec4<f32>(raw_xm.zw * 2.0, 0.0, 0.0), delta.im)
);
delta = qc_add(qc_mul(w, delta), c_delta_qs);

let next_m = m + 1u;
let raw_xm_next = reference_orbit[next_m];

// Compute zn in single precision for escape check & coloring
let zn_re = raw_xm_next.re.x + delta.re.x;
let zn_im = raw_xm_next.im.x + delta.im.x;
let zn_re = raw_xm_next.x + delta.re.x;
let zn_im = raw_xm_next.z + delta.im.x;
zn_sq = fma(zn_re, zn_re, zn_im * zn_im);

i = i + 1u;
Expand All @@ -634,10 +631,13 @@ fn fs_main_qs(@location(0) uv: vec2<f32>) -> @location(0) vec4<f32> {

let delta_norm_sq = fma(delta.re.x, delta.re.x, delta.im.x * delta.im.x);
if (zn_sq < delta_norm_sq || next_m >= uniforms.ref_iter) {
let xm_next = qs_complex(raw_xm_next.re, raw_xm_next.im);
let xm_next = qs_complex(
vec4<f32>(raw_xm_next.xy, 0.0, 0.0),
vec4<f32>(raw_xm_next.zw, 0.0, 0.0)
);
delta = qc_add(xm_next, delta);
m = 0u;
raw_xm = OrbitPoint(vec4<f32>(0.0), vec4<f32>(0.0));
raw_xm = vec4<f32>(0.0);
} else {
m = next_m;
raw_xm = raw_xm_next;
Expand Down
Loading
Loading