How HotSpot intrinsics speed up the JDK's post-quantum ML-KEM and ML-DSA
An Inside Java article published on 30 September explains how the JDK's post-quantum algorithms get their speed. The JDK supports ML-KEM (FIPS 203), ML-DSA (FIPS 204) and the hash-based HSS/LMS signatures of RFC 8554, all implemented in Java for portability. For the most performance-sensitive methods, HotSpot can substitute an intrinsic: platform-specific machine code that uses CPU features such as SHA-3 and SHA-256 acceleration, vector instructions and efficient modular arithmetic.
An intrinsic does not replace the Java code. Methods marked @IntrinsicCandidate may be swapped when the current platform has an implementation, and the Java method remains the fallback everywhere else. The article's rule for choosing candidates: the method must be expensive, frequently called, and meaningfully faster in machine code.
ML-KEM fits that rule because its hot paths are regular computations over polynomials with exactly 256 coefficients: forward and inverse number-theoretic transforms, multiplication in the transform domain, polynomial addition, Barrett reduction and bit packing. The authors stress that assembly is not inherently faster than Java; these operations win because CPUs can apply the same few steps to every coefficient. ML-DSA does related work but is not interchangeable, since it uses the modulus 8380417 where ML-KEM uses 3329, so each algorithm needs its own intrinsics; ML-DSA also has specific ones such as implDilithiumDecomposePoly. For HSS/LMS verification, most of the time goes into SHA-256, so the SHA-256 intrinsic does most of the work.
The article reports throughput with intrinsics against the Java-only code on Ampere Altra and Intel Ice Lake servers, and notes how to read its charts: a 218% speed-up means 3.18 times as fast, not 2.18.

Why it matters
Post-quantum key exchange is becoming the default in TLS, and its cost lands on every handshake a Java server makes. The article shows the JDK's approach: keep one portable implementation and accelerate the few primitives that dominate, so applications get the speed without platform-specific code of their own.