Skip to content

Prepare for Project Valhalla value classes (JDK 28+) #54

Description

@stalep

Summary

Audit and prepare jjq's JqValue type hierarchy for Project Valhalla value classes (JEP 401, targeting JDK 28 preview in March 2027). Value classes eliminate object headers, enable heap flattening (inline storage in arrays/fields), and allow scalarization (zero-allocation method parameter passing). This directly addresses jjq's measured L1 cache miss bottleneck by making arrays of JqValue types cache-friendly.

Background

Project Valhalla (JEP 401) introduces value class — classes without identity that the JVM can optimize aggressively:

  • Scalarization: value objects are decomposed into their fields when passed to/from methods — no allocation, no GC pressure
  • Heap flattening: value objects stored inline in fields and arrays — no pointer indirection, contiguous memory layout
  • No object header: eliminates 16 bytes of overhead per instance

A value class is declared with value class Foo { ... }. Constraints: all fields implicitly final, no synchronized, == becomes field-by-field comparison (substitutability), class is final by default.

Reference: JEP 401, JVM Weekly deep dive

Why This Matters for jjq

perf stat data (2026-07-02) shows L1 cache misses as the primary bottleneck:

Benchmark L1-dcache-load-misses Miss rate
parse_flat_10kb 1,240 0.79%
parse_nested_10kb 1,486 0.50%
prod_extractMetric 4,658 5.69%

Branch misprediction is negligible (0.02-0.15%, IPC 4.3-5.6). The bottleneck is pointer chasing through heap-allocated value objects when iterating arrays.

Allocation profile shows JqValue types as top allocators:

Type Samples %
byte[] 3,304 32.1%
JqString 1,493 14.5%
JqValue[] 511 5.0%
JqNumber 13 0.1%

With value classes, JqNumber and JqString instances would be scalarized (zero allocation when passed between methods) and flattened in arrays (contiguous memory, cache-friendly).

Per-Type Assessment

JqNull — Ready for value class

  • No mutable state, no fields beyond type tag
  • Singleton pattern (JqNull.NULL) becomes unnecessary — all JqNull instances are substitutable
  • == comparison: no fields → always true for any JqNull — correct behavior

JqBoolean — Ready for value class

  • One final boolean field
  • Singleton pattern (TRUE/FALSE) becomes unnecessary — value equality is free
  • == comparison: compares boolean value field — correct behavior

JqNumber — Ready with minor change

  • Current: 40 bytes (16B header + 8B long + 8B BigDecimal ref + 8B double + 1B boolean + padding)
  • As value class: ~25 bytes (no header), scalarized in method calls, flat in arrays
  • Issue: cachedDecimal is a lazy mutable field (private transient BigDecimal cachedDecimal). Value classes require all fields to be final.
  • Fix: Remove cachedDecimal. Compute BigDecimal.valueOf(longVal) on every decimalValue() call. This field is rarely accessed (only for high-precision arithmetic in add/subtract/multiply/divide with mixed long/double operands). The allocation cost of BigDecimal.valueOf() per call is small compared to the savings from eliminating JqNumber object allocation entirely.
  • Alternative: Keep JqNumber as identity class if decimalValue() is hot. Benchmark both approaches.
  • The -128..1023 cache (JqNumber[] CACHE) becomes unnecessary — value classes are inherently cheap.

JqString — Needs design decision (biggest challenge)

  • Current: 32 bytes (16B header + 8B Object source + 4B start + 4B end + 1B hasEscapes + padding)
  • Issue: The deferred string pattern uses a volatile String value field for lazy materialization. Value classes cannot have mutable fields.
  • Option A: Eagerly materialize all strings at parse time. Eliminates deferred pattern. Increases parse allocation (every string creates a Java String immediately) but eliminates JqString wrapper allocation. Net effect depends on workload: for documents where most strings are accessed, eager is better. For documents where most strings are passthrough (h5m's 93% case), deferred is better.
  • Option B: Keep JqString as an identity class. It retains deferred materialization but doesn't benefit from value class optimizations. Since JqString is the Avoid full lazy map conversion when only a subset of fields are needed #2 allocator, this leaves significant optimization on the table.
  • Option C: Two implementations — a value-class EagerJqString for when strings are known to be needed, and identity-class DeferredJqString for parse-time deferred strings. The sealed interface allows this. Complex but preserves both benefits.
  • Recommendation: Benchmark Option A (eager) vs current (deferred) on the production 14MB workload. If eager parse allocation is acceptable, make JqString a value class. If not, Option B (keep as identity class) and focus value class benefits on JqNumber/JqNull/JqBoolean.

JqArray — Stays as identity class

  • Contains List<JqValue> elements (mutable in builder pattern)
  • Has lazy caches: none currently, but the architecture assumes mutability
  • Indirect benefit: if JqNumber/JqString/JqBoolean are value classes, JqValue[] arrays inside JqArray store them flat — this is where the L1 cache improvement comes from

JqObject — Stays as identity class

  • Contains parallel arrays (String[] keys, JqValue[] values) and lazy caches (hashSlots, sortedKeysCache, mapView, etc.)
  • Multiple mutable/transient fields — incompatible with value classes
  • Indirect benefit: JqValue[] values array stores value-class elements flat

Sealed Interface Compatibility

JqValue is a sealed interface with 6 permitted subtypes. Value classes can implement interfaces. Mixing value classes and identity classes in a sealed hierarchy is explicitly supported:

public sealed interface JqValue extends Comparable<JqValue>, Serializable
    permits JqNull, JqBoolean, JqNumber, JqString, JqArray, JqObject {
    // JqNull, JqBoolean, JqNumber → value class
    // JqString → value class (if eager) or identity class (if deferred)
    // JqArray, JqObject → identity class
}

== Semantics Change

For value classes, == becomes substitutability (field-by-field comparison) instead of identity (same heap address). Impact on jjq:

  • JqNull.NULL == someNull → always true (no fields to compare) — correct
  • JqBoolean.TRUE == JqBoolean.of(true) → true (same boolean value) — correct
  • JqNumber.of(42) == JqNumber.of(42) → true (same long value) — correct, matches current equals() behavior
  • Intern cache reference equality (INTERN_TABLE[slot] == key) → still works for String (identity class), unaffected

Connection to Vector API (SIMD)

The Vector API (JEP 508) has been incubating for 10 rounds partly because it needs Valhalla — vector instances need to be value classes for SIMD register allocation. With Valhalla in JDK 28:

  • The Vector API can potentially graduate in JDK 29-30
  • ByteVector, DoubleVector become value classes mapped to SIMD registers
  • jjq's byte[] parser could upgrade from SWAR (8-byte) to Vector API (32-byte) string scanning
  • jhunter's distance computation (issue jhunter#1) could use DoubleVector for 4-8x speedup

Implementation Plan

Phase 1: Audit and experiment (JDK 28 early-access)

  • Download JDK 28 EA builds when available (expected mid-2026)
  • Create an experimental branch with --enable-preview
  • Convert JqNull and JqBoolean to value classes — verify all tests pass
  • Convert JqNumber to value class (remove cachedDecimal) — benchmark parse/query
  • Experiment with eager JqString — benchmark against deferred baseline
  • Run perf stat with value classes — measure L1 cache miss reduction on prod_extractMetric

Phase 2: Production readiness (JDK 29, likely LTS)

  • Finalize which types are value classes based on Phase 1 benchmarks
  • Remove singleton caches (JqNull.NULL, JqBoolean.TRUE/FALSE, JqNumber.CACHE) if unnecessary
  • Update Serializable support (readResolve may need changes for value classes)
  • Update documentation and README
  • Ensure all 478+ jjq-core tests pass with value classes enabled

Phase 3: Vector API integration (JDK 29-30)

  • When Vector API graduates, prototype SIMD string scanning in byte[] parser
  • Benchmark 32-byte ByteVector scanning vs current 8-byte SWAR
  • Coordinate with jhunter#1 (SIMD distance computation)

Caveats

  • JDK 28 is preview: the value modifier and its semantics may change before finalization
  • Flattening limited to ~64 bits atomically in JDK 28: JqNumber (long + double + boolean = ~17 bytes) may not flatten in arrays until null-restricted types arrive (future JEP). Scalarization would still work.
  • JDK 28 is not LTS: most production deployments will wait for JDK 29 (Sept 2027)
  • Specialized generics not in JDK 28: ArrayList<JqNumber> won't be flat until future JDK releases. Direct JqValue[] arrays will benefit immediately.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestperformancePerformance improvements and benchmarking

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions