perf: pack reference orbit to 16 bytes/point, fast-extend periodic orbits, and scale deep-zoom slice ceiling - #98
Merged
Conversation
…bits, and scale deep-zoom slice ceiling
|
Visit the preview URL for this PR (updated for commit fa42de4): https://fractious-deep--pr98-perf-orbit-packing-a-f1zzexlj.web.app (expires Mon, 05 Oct 2026 15:21:47 GMT) 🔥 via Firebase Hosting GitHub Action 🌎 Sign: 348393a9746312ebab7a45ce5eba16e6d119af34 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
reference_orbitbuffer size (32 bytes/point->16 bytes/point): Store each reference orbit point as a singlevec4<f32>([re_hi, re_lo, im_hi, im_lo]) usingsplit_f64_to_2_f32instead of2 × vec4<f32>. Since reference points are converted fromf64(53-bit significand), twof32s already capture the full Double-Single significand while halving WASM memory allocation, worker transfer size,writeBufferupload bandwidth, GPU storage cache footprint, andraw_xmregister pressure infs_main_qs.wasm/src/lib.rs(Orbit::advance), record the detected cycle period fromCycleDetectorand, once a full period ofpconsecutive points is stationary inqs, copy thep-point cycle directly up tolimitinstead of re-iterating in arbitrary-precisionFBig.src/renderer/shader.wgsl, extendfs_main_f32Brent periodicity detection across its full zoom range by comparing perturbationsdeltawhenevernext_m == period_m(whereX_mcancels out), and usefmainds_mul.Renderer._sliceRows(), floormaxOpsatMIN_SLICE_ROWS * 2 * rowOpsso fast exterior slices at ultra-high iteration counts (where16 * budget < 32 rows) can still scale up to2 × MIN_SLICE_ROWS(64 rows) while keeping slow interior slices floored at32 rows.