Benchmarks: Publish Eigen vs Arm Performance Libraries, NVPL and OpenBLAS on NVIDIA DGX Spark

Two pages for one Arm Cortex-X925 core of an NVIDIA DGX Spark (GB10: 10 Cortex-X925 + 10 Cortex-A725 cores): Eigen master 210e0a609 built for NEON against Arm Performance Libraries 26.07, NVPL 26.5 and OpenBLAS 0.3.34, each the newest release, and the same cells with Eigen built for SVE at the CPU's 128-bit vector length against the NEON session's library measurements. Same operations, cells and method as the Grace and Apple M4 pages.

Preview: NEON, SVE

GEMM at n = 1024, GFLOP/s, double / float:

Eigen NEON Eigen SVE Arm Performance Libraries NVPL OpenBLAS
72 / 155 76 / 160 83 / 171 79 / 170 69 / 137

What the pages show

  • From n = 192 on Eigen's NEON GEMM runs at 0.85–0.90 of Arm Performance Libraries in double and 0.86–0.96 in float, and at 0.89–0.92 of NVPL from n = 768 on; it leads OpenBLAS's Neoverse V2 kernels by 4–10 % (double) and 11–17 % (float).
  • Eigen is the fastest arm for SYEV at every size (1.2–3.5x), for float GETRF from n = 24, float POTRF from n = 128 and float TRSM up to n = 2048, and for DOT in L1. GEEV falls to 0.23–0.28x of the libraries at n = 1024, as on the other pages.
  • NVPL's fixed cost per call is 15 ns (Eigen's DOT at n = 2: 0.6 ns), which sets its ratios at n ≤ 16.
  • Eigen's SVE build runs at 0.96 of its NEON build (geometric mean over 882 cells). Unlike on Grace, double GEMM is 4–9 % faster under SVE and SYRK, POTRF and GETRF in double follow it; float TRSM runs at 0.38–0.57 of NEON from n = 8 to 128 and float GETRF at 0.47–0.81 up to 128.

Found along the way, not filed yet

Each reproduces in every pass, so none is noise:

  • SVE build only: GEMM and SYRK drop at n = 192 (double GEMM 59 GFLOP/s against 78 and 80 at n = 160 and 200; double SYRK 42 against 59 and 67 at n = 128 and 200).
  • SYRK drops at n = 384 and 768 in float (103 and 121 GFLOP/s against 132–153 around them) and at n = 512 in double (52 against 70), in both builds; GEQRF drops at n = 64.
  • OpenBLAS 0.3.34's square double GEMM drops at n = 16, 32, 64, 96, 128 and 256 (13.5–56 GFLOP/s against 66–68 at n = 100, 160, 192 and 257).

Method and provenance

  • Tree measured: master 210e0a609 with the eigen!2903 harness commit and the filled-in cortex-x925-20c profile on top, 27ef2d3a5 on the fork branch benchmark-comparison-harness-x925. Apart from the two machine profiles, the harness files are byte-identical to the tree behind the Grace re-measure in !16 (merged); the profile is now in eigen!2903 as 7e242b343, the same file.
  • GCC 13.3, -O3 -DNDEBUG (NEON, default Armv8-A target; GCC 13 does not accept -mcpu=cortex-x925) and -O3 -DNDEBUG -march=armv9-a+sve2 -msve-vector-bits=128 -DEIGEN_ARM64_USE_SVE (SVE). Sequential LP64 libraries called through the Fortran symbols, 1 thread, pinned to one Cortex-X925 core.
  • Arm Performance Libraries from Arm's deb installer (--install-to), NVPL from the redistributable archives, OpenBLAS 0.3.34 built from the release tarball for its NEOVERSEV2 target, which its own CPU detection selects for the X925 (0.3.34 has no X925 target), with gfortran 13.3. Google Benchmark 1.9.5 built in Release.
  • The twelve operations were split across three DGX Spark machines to fit an 8-hour job limit, every arm of an operation (three NEON passes, two SVE passes) on one machine, so every chart and ratio comes from one machine; the pages say which operation ran where. Runs gated on measured CPU use (< 0.5 cores over 5 s); 36 NEON runs (6.6 h) and 24 SVE runs (2.2 h), none repeated, no errored cell.
  • Every number in the prose is derived from the result files with plot_comparison.py's aggregation; the reduction scripts reproduce the September Grace pages' published spreads, SVE band table and cell counts. The 60 result files sit beside the charts.
  • Independent of !16 (merged): the two merge cleanly in either order.

🤖 Generated with Claude Code

Edited by Rasmus Munk Larsen

Merge request reports

Loading
Loading