Skip to content

perf(render,wasm): add Order-8 Series Approximation and factor perturbation loops - #96

Merged
rowan-m merged 1 commit into
mainfrom
perf/series-approx-and-factored-perturbation
Sep 25, 2026
Merged

rowan-m merged 1 commit into
mainfrom
perf/series-approx-and-factored-perturbation

Conversation

@rowan-m

@rowan-m rowan-m commented Sep 25, 2026

Copy link
Copy Markdown
Owner

Summary

Maximises deep-zoom CPU/WASM and WebGPU rendering throughput while keeping progressive slices responsive on mobile GPUs:

  1. Order-8 Circle-Probed Series Approximation (wasm/src/lib.rs, src/renderer/shader.wgsl):

    • Evolves an 8-term Taylor polynomial $\Delta_m(u) \approx \sum_{k=1}^8 a_k^{(m)} u^k$ normalized to the viewport bounding disk $u = \Delta c / R_{\text{max}}$ ($|u| \le 1$, preventing f64 exponent overflow/underflow at arbitrary zoom depths).
    • Validates the polynomial up to the first close approach / rebase against 8 boundary probes on $|u| = 1$ (by the Maximum Modulus Principle, maximum truncation and non-linearity error over the disk occurs on its boundary).
    • Transfers the Double-Single normalized coefficients to the GPU so pixels inside the validated disk (dot(u_norm, u_norm) <= 1.0) jump directly to i = skip_iter, m = skip_iter (skipping 12,584 iterations per pixel on z = 17.345), while out-of-disk pixels during large interactive drags seamlessly start at 0.
  2. Factored Perturbation Recurrence Across All 3 Shader Tiers (src/renderer/shader.wgsl):

    • Replaces $2 X_m \Delta_n + \Delta_n^2 + \Delta c$ with $(2 X_m + \Delta_n)\Delta_n + \Delta c$, eliminating qs_sqr, qc_sq, ds_sqr, and dc_sq (~43% fewer multiplications per iteration with identical floating-point accuracy).
  3. Precision Tier Threshold & SA-Aware Slice Budgeting (src/Renderer.js):

    • Extends Tier 2 (DS, Double-Single) to z < 21.0 (scale > 1.0e-21), since post-rebase squaring $Z_n^2 + \Delta c$ only requires $z/2$ digits of mantissa precision.
    • Discounts skipIter in _effectiveIter(config) for interactive preview and progressive slice budgeting, and updates opsMultiplier (12.0 for DS, 8.0 for QS) to reflect the factored shaders and avoid single-wavefront 1-row slice traps on tile-based mobile GPUs.
  4. Zero-Allocation f64 Quad Splitting & Early Phase-1 Spiral Exit (wasm/src/lib.rs):

    • Replaces split_fbig_to_4_f32 (12 heap allocations per step) and FBig escape addition in Orbit::advance with pure stack f64 splitting (split_f64_to_4_f32).
    • Exits the Phase 1 3 × 3 FBig spiral in search_anchor as soon as a candidate survives past the immediate exterior (score > 16) so Phase 2 surveys the rest of the viewport in fast hardware f64 perturbation.

@github-actions

Copy link
Copy Markdown

Visit the preview URL for this PR (updated for commit 602986a):

https://fractious-deep--pr96-perf-series-approx-a-ow4olipy.web.app

(expires Fri, 02 Oct 2026 15:35:33 GMT)

🔥 via Firebase Hosting GitHub Action 🌎

Sign: 348393a9746312ebab7a45ce5eba16e6d119af34

@rowan-m
rowan-m merged commit 7ead256 into main Sep 25, 2026
2 checks passed
@rowan-m
rowan-m deleted the perf/series-approx-and-factored-perturbation branch September 25, 2026 15:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant