Core: Vectorize triangular solves across right-hand sides

This adds a shared SIMD kernel for TRSM, processing independent right-hand sides in packet lanes. GEBP’s register budget determines the RHS packet group, and Eigen’s configured L1 cache size controls when a solve bypasses matrix-multiply packing. Fresh measurements cover GCC and Clang × SSE2 and AVX2+FMA, using the same source revision and benchmark in all four configurations.

Measured revision: eba485c30, before the branch was rebased to 015451069. The TRSM kernel, benchmark and focused test sources are byte-identical after rebasing; the rebased tree has not been timed.

Overall speedup: pre-optimization Eigen 15f227178 → eba485c30. Geometric means over seven square sizes, 16–2048, for each scalar. Column-major lower-triangular left solve, including the RHS copy; greater than 1× is faster.

Compiler / ISA float geomean double geomean Range across 14 cases
GCC / SSE2 1.548× 1.282× 1.005–3.683×
Clang / SSE2 1.549× 1.388× 1.024–3.964×
GCC / AVX2+FMA 2.548× 1.765× 1.085–7.924×
Clang / AVX2+FMA 2.859× 1.841× 1.092–12.262×

Review-fix cost: 0c7bdb309 → eba485c30. The same 44-case grid in each configuration, including both triangle layouts, square and narrow RHS cases. The Clang/AVX2 review row retains the earlier valid run; its attempted refresh was interrupted by background work and excluded. Both rebuilt binaries are byte-identical to the earlier pair.

Compiler / ISA Column-major geomean (22) Row-major geomean (22) Lowest observed ratio
GCC / SSE2 0.990× 0.983× double 64×17: 0.878×; 20.3% variation
Clang / SSE2 0.996× 1.001× double 8×8 (row-major A): 0.972×; 1.6% variation
GCC / AVX2+FMA 1.012× 0.985× float 16×16 (row-major A): 0.897×; 1.3% variation
Clang / AVX2+FMA (earlier run) 1.012× 0.997× double 16×16 (row-major A): 0.965×; 4.3% variation

Variation in the last column is the maximum CV or run spread, not a confidence interval. The GCC/SSE2 0.878× case has 20.3% variation and is inconclusive. Tiny row-major float solves retain a cost: GCC/SSE2 8×8 takes 5.6% longer, and GCC/AVX2 16×16 takes 11.5% longer (0.0713 → 0.0795 µs), consistent with the earlier observed guard cost. Large SSE2 double gains are only about 1–3% and generally near the observed variability. Earlier localized 3–6% generalization costs against 76c9fb280 remain limitations. No additional targeted repeats were performed.

The kernel transposes a bounded RHS tile, solves register blocks of rows, and transposes the results back. Four register rows and at most two RHS packets preserve the established AVX2 schedule. WorkspaceRows = 128 bounds stack capacity separately from the cache policy: the direct-solve footprint estimate is n * (packets * PacketSize + RegisterRows + 1) * sizeof(Scalar), plus one packet transpose, against half of L1. Cache discovery and overrides use Eigen’s existing machinery; outer L2/L3 blocking remains in place. These are conservative defaults, not optimal tuning claims for every platform.

The fast path requires vectorized built-in IEEE binary float/double and compile-time unit inner stride. Partial RHS packets, complex/custom scalars and non-unit strides retain their existing paths. AVX-512 builds select their existing specialized TRSM kernel by default; unifying that implementation remains separate work.

Update: Thanks to @florian360 for the finite-result overflow example and @onalante-ebay for the expression/style suggestions. eba485c30 checks the final solved row of each row-major RHS tile before writeback, and retries a nonfinite tile using the original row-dot-product accumulation order. Every later row consumes all earlier solved rows, so nonfinite lanes propagate to the final row. Integer exponent-bit classification and Eigen’s fast-math barrier keep the check effective under Clang -ffast-math. Reciprocal initialization now uses a map of the active diagonal; the other suggestions use EIGEN_IF_CONSTEXPR and numext::round_down with an Index conversion.

The retry restores the previous diagonal-kernel behavior; it is not a general overflow-safe solver. The existing default AVX-512 specialization also overflows on the example before this MR and is unchanged.

Measurements: 2026-09-17, 13th Gen Intel(R) Core(TM) i7-13700HX under WSL2, GCC 13.3.0 / Clang 18.1.3. CPU-time medians, one process pinned to CPU 2, A/B/A/B invocations, eight repetitions per case. Windows clocks, scheduling and thermal policy were uncontrolled. Full flags, protocol and limitations are in Appendix E.

GCC and Clang correctness validation passed, including the additional SSE2 runs. The regression and prior shared-kernel AVX-512/NEON validation are detailed in Appendix E.

Appendix A: GCC / SSE2 — overall and review comparisons

Overall: 15f227178 → eba485c30 (14 fresh cases)

Case Before (µs) After (µs) Speedup CV before / after Run spread before / after
float 16×16 0.5576 0.1514 3.683× 3.1% / 1.6% 1.7% / 0.6%
float 64×64 13.7400 6.8228 2.014× 1.3% / 1.6% 0.8% / 0.5%
float 128×128 91.2179 52.3067 1.744× 1.0% / 1.2% 0.1% / 0.3%
float 256×256 561.0331 473.9017 1.184× 1.4% / 1.7% 1.7% / 0.6%
float 512×512 4234.2389 3623.6479 1.169× 1.9% / 2.1% 1.9% / 1.0%
float 1024×1024 31527.6154 27794.8264 1.134× 0.7% / 1.6% 2.4% / 2.1%
float 2048×2048 227246.9757 216860.4282 1.048× 0.9% / 1.8% 1.6% / 0.5%
double 16×16 0.5667 0.2881 1.967× 6.5% / 1.2% 16.2% / 1.2%
double 64×64 22.4007 14.1618 1.582× 2.0% / 1.2% 0.7% / 0.9%
double 128×128 181.7513 110.6166 1.643× 1.9% / 1.7% 1.9% / 0.7%
double 256×256 1158.5543 1091.9134 1.061× 1.1% / 2.9% 1.3% / 1.9%
double 512×512 8358.3884 8231.0639 1.015× 2.4% / 2.6% 1.0% / 0.4%
double 1024×1024 60585.0261 59018.6196 1.027× 2.7% / 1.6% 0.1% / 0.8%
double 2048×2048 451579.4922 449198.4952 1.005× 1.3% / 2.4% 0.5% / 0.2%

Review: 0c7bdb309 → eba485c30 (44 cases)

Case Before (µs) After (µs) Speedup CV before / after Run spread before / after
float 8×8 0.0345 0.0345 0.999× 2.3% / 1.5% 1.0% / 0.0%
float 16×16 0.1527 0.1536 0.994× 1.5% / 1.3% 2.1% / 0.2%
float 64×64 6.8934 6.9861 0.987× 2.6% / 2.8% 1.8% / 3.6%
float 128×128 52.7329 52.8021 0.999× 1.1% / 1.4% 1.1% / 0.2%
float 256×256 484.8405 486.8830 0.996× 3.8% / 1.9% 3.7% / 5.1%
float 512×512 3641.6702 3660.9862 0.995× 1.2% / 1.5% 2.0% / 2.7%
float 1024×1024 27734.8454 28237.3886 0.982× 1.3% / 1.7% 1.3% / 0.1%
float 64×8 0.8637 0.8777 0.984× 1.4% / 3.7% 0.1% / 1.0%
float 64×17 2.2002 2.2131 0.994× 3.4% / 5.2% 1.7% / 0.3%
float 128×8 3.2751 3.2834 0.997× 1.3% / 0.9% 0.8% / 1.4%
float 128×17 7.7792 7.6997 1.010× 2.4% / 1.1% 0.4% / 1.2%
double 8×8 0.0620 0.0567 1.094× 7.6% / 1.0% 19.8% / 1.2%
double 16×16 0.2920 0.2906 1.005× 2.6% / 3.2% 2.8% / 3.5%
double 64×64 14.3612 14.3880 0.998× 1.3% / 1.8% 0.8% / 2.4%
double 128×128 111.0311 112.8277 0.984× 1.1% / 1.9% 0.6% / 2.3%
double 256×256 1072.1497 1084.9934 0.988× 1.4% / 1.9% 0.1% / 0.1%
double 512×512 8000.9396 8078.9894 0.990× 1.4% / 4.2% 1.7% / 0.2%
double 1024×1024 59348.5789 59608.0896 0.996× 2.3% / 1.5% 0.5% / 1.4%
double 64×8 1.7688 1.8005 0.982× 1.2% / 2.7% 0.3% / 4.0%
double 64×17 3.8956 4.4344 0.878× 1.1% / 10.0% 0.1% / 20.3%
double 128×8 6.8641 6.9570 0.987× 1.4% / 6.1% 0.0% / 0.3%
double 128×17 15.1213 15.8217 0.956× 2.2% / 1.9% 0.2% / 0.8%
float 8×8 (row-major A) 0.0351 0.0371 0.947× 1.1% / 1.8% 0.2% / 1.4%
float 16×16 (row-major A) 0.1523 0.1549 0.984× 2.3% / 1.4% 1.7% / 1.8%
float 64×64 (row-major A) 7.3920 7.0280 1.052× 10.1% / 3.1% 15.7% / 2.5%
float 128×128 (row-major A) 52.2264 52.8342 0.988× 6.3% / 2.2% 0.3% / 1.2%
float 256×256 (row-major A) 477.2216 483.9215 0.986× 1.3% / 1.4% 0.5% / 1.2%
float 512×512 (row-major A) 3599.2099 3976.9805 0.905× 1.1% / 6.8% 0.4% / 17.1%
float 1024×1024 (row-major A) 27846.4506 28681.4812 0.971× 1.8% / 3.1% 1.0% / 3.7%
float 64×8 (row-major A) 0.8641 0.8646 0.999× 0.9% / 1.1% 1.8% / 0.5%
float 64×17 (row-major A) 2.1874 2.2039 0.993× 0.8% / 3.2% 0.5% / 1.5%
float 128×8 (row-major A) 3.2616 3.2842 0.993× 1.1% / 1.3% 0.1% / 2.0%
float 128×17 (row-major A) 8.4872 8.5835 0.989× 1.1% / 1.9% 0.4% / 1.6%
double 8×8 (row-major A) 0.0606 0.0589 1.030× 11.8% / 1.6% 14.5% / 0.9%
double 16×16 (row-major A) 0.2815 0.2823 0.997× 6.2% / 1.4% 2.9% / 0.9%
double 64×64 (row-major A) 14.3667 14.4944 0.991× 1.4% / 1.3% 1.2% / 0.5%
double 128×128 (row-major A) 111.1262 112.0957 0.991× 3.6% / 1.2% 1.4% / 0.5%
double 256×256 (row-major A) 1079.7081 1202.0414 0.898× 2.0% / 8.9% 1.2% / 12.0%
double 512×512 (row-major A) 8034.1150 8083.5608 0.994× 2.0% / 1.5% 1.1% / 0.5%
double 1024×1024 (row-major A) 59770.7063 61668.9123 0.969× 1.8% / 11.7% 0.7% / 0.7%
double 64×8 (row-major A) 1.8216 1.8725 0.973× 1.5% / 8.7% 1.2% / 5.9%
double 64×17 (row-major A) 4.0648 4.2599 0.954× 2.0% / 7.6% 0.6% / 7.5%
double 128×8 (row-major A) 7.0935 6.9709 1.018× 1.6% / 1.6% 2.7% / 0.1%
double 128×17 (row-major A) 16.1915 16.0425 1.009× 2.4% / 0.8% 1.8% / 0.5%
Appendix B: Clang / SSE2 — overall and review comparisons

Overall: 15f227178 → eba485c30 (14 fresh cases)

Case Before (µs) After (µs) Speedup CV before / after Run spread before / after
float 16×16 0.6663 0.1681 3.964× 7.3% / 3.5% 7.9% / 15.2%
float 64×64 14.0409 7.6860 1.827× 4.9% / 6.0% 0.7% / 17.1%
float 128×128 91.6284 53.4162 1.715× 4.3% / 2.8% 5.5% / 2.1%
float 256×256 592.0898 456.4255 1.297× 8.0% / 1.3% 14.8% / 0.3%
float 512×512 3923.3726 3447.7765 1.138× 3.9% / 2.7% 1.0% / 0.1%
float 1024×1024 29448.5509 27023.4574 1.090× 1.7% / 3.6% 2.0% / 3.7%
float 2048×2048 221602.5165 206758.0272 1.072× 6.8% / 3.0% 2.9% / 0.5%
double 16×16 0.6979 0.2851 2.448× 4.3% / 1.3% 18.1% / 1.1%
double 64×64 24.3558 14.2676 1.707× 3.6% / 1.2% 16.3% / 1.3%
double 128×128 187.7486 111.6513 1.682× 6.3% / 1.8% 15.1% / 1.1%
double 256×256 1236.0550 1049.8741 1.177× 5.8% / 1.3% 18.4% / 0.2%
double 512×512 9015.4704 7908.4092 1.140× 3.1% / 2.6% 21.2% / 3.4%
double 1024×1024 60647.5749 59039.0686 1.027× 1.4% / 2.4% 2.6% / 2.3%
double 2048×2048 459841.1825 448971.9905 1.024× 7.0% / 5.0% 2.4% / 0.1%

Review: 0c7bdb309 → eba485c30 (44 cases)

Case Before (µs) After (µs) Speedup CV before / after Run spread before / after
float 8×8 0.0351 0.0352 0.997× 1.2% / 1.4% 0.4% / 3.0%
float 16×16 0.1539 0.1547 0.995× 1.0% / 1.4% 0.5% / 1.4%
float 64×64 6.8761 6.9534 0.989× 1.0% / 1.2% 0.3% / 0.8%
float 128×128 52.8017 52.9650 0.997× 1.2% / 1.6% 1.2% / 0.1%
float 256×256 457.7716 459.6645 0.996× 5.4% / 6.9% 0.2% / 1.9%
float 512×512 3475.3373 3476.3559 1.000× 1.4% / 1.3% 0.8% / 0.4%
float 1024×1024 26646.3233 26526.2724 1.005× 2.2% / 6.7% 1.9% / 0.3%
float 64×8 0.8711 0.8809 0.989× 1.4% / 4.3% 0.4% / 1.7%
float 64×17 2.0853 2.0740 1.005× 3.5% / 4.9% 1.1% / 0.7%
float 128×8 3.2885 3.2787 1.003× 1.2% / 1.2% 0.6% / 2.1%
float 128×17 7.5892 7.6158 0.997× 0.9% / 1.8% 0.9% / 0.9%
double 8×8 0.0574 0.0578 0.992× 0.9% / 1.6% 1.9% / 1.5%
double 16×16 0.2865 0.2878 0.996× 1.7% / 2.4% 0.4% / 1.5%
double 64×64 14.2430 14.4612 0.985× 3.4% / 3.7% 0.4% / 3.1%
double 128×128 111.1773 114.0092 0.975× 1.6% / 1.6% 0.0% / 0.6%
double 256×256 1071.4788 1076.8173 0.995× 4.5% / 2.0% 2.8% / 3.6%
double 512×512 7981.5099 7974.0080 1.001× 1.8% / 5.4% 2.5% / 3.5%
double 1024×1024 60038.1918 59618.3565 1.007× 3.0% / 1.5% 0.6% / 1.2%
double 64×8 1.7971 1.8094 0.993× 1.4% / 2.1% 2.9% / 1.0%
double 64×17 4.0454 4.0298 1.004× 1.7% / 1.5% 4.3% / 1.2%
double 128×8 7.1518 7.1595 0.999× 3.9% / 1.5% 1.9% / 1.9%
double 128×17 15.5134 15.7093 0.988× 1.0% / 1.7% 0.3% / 1.8%
float 8×8 (row-major A) 0.0368 0.0366 1.005× 2.8% / 0.9% 4.7% / 1.3%
float 16×16 (row-major A) 0.1534 0.1539 0.997× 1.4% / 1.4% 0.6% / 0.1%
float 64×64 (row-major A) 6.9136 6.8988 1.002× 1.2% / 1.5% 0.2% / 0.5%
float 128×128 (row-major A) 52.0302 52.4365 0.992× 1.7% / 1.9% 0.2% / 0.6%
float 256×256 (row-major A) 468.0477 465.6385 1.005× 3.4% / 1.0% 1.0% / 1.5%
float 512×512 (row-major A) 3563.3720 3490.4749 1.021× 7.6% / 1.1% 2.3% / 0.1%
float 1024×1024 (row-major A) 27218.0376 26813.7949 1.015× 2.6% / 1.6% 5.2% / 0.0%
float 64×8 (row-major A) 0.8861 0.8724 1.016× 4.2% / 2.0% 3.4% / 0.7%
float 64×17 (row-major A) 2.1651 2.1744 0.996× 1.6% / 4.2% 0.2% / 1.3%
float 128×8 (row-major A) 3.2273 3.2611 0.990× 1.3% / 1.4% 0.3% / 0.8%
float 128×17 (row-major A) 8.5284 8.5458 0.998× 1.1% / 1.4% 0.6% / 0.0%
double 8×8 (row-major A) 0.0587 0.0605 0.972× 1.6% / 0.9% 0.6% / 0.1%
double 16×16 (row-major A) 0.2913 0.2930 0.994× 1.0% / 1.8% 0.9% / 0.2%
double 64×64 (row-major A) 14.3652 14.4330 0.995× 2.4% / 1.5% 0.5% / 0.3%
double 128×128 (row-major A) 110.8081 111.5462 0.993× 2.2% / 1.5% 1.5% / 1.6%
double 256×256 (row-major A) 1052.4882 1041.6074 1.010× 2.8% / 1.2% 3.5% / 0.2%
double 512×512 (row-major A) 7878.6308 7896.4791 0.998× 2.3% / 3.7% 1.7% / 0.1%
double 1024×1024 (row-major A) 59710.5261 59372.7654 1.006× 3.2% / 1.3% 0.6% / 0.4%
double 64×8 (row-major A) 1.8327 1.8034 1.016× 6.8% / 1.0% 0.9% / 0.2%
double 64×17 (row-major A) 4.0394 4.0292 1.003× 10.4% / 1.5% 2.3% / 1.3%
double 128×8 (row-major A) 6.9353 6.9864 0.993× 1.4% / 1.2% 0.6% / 0.5%
double 128×17 (row-major A) 15.8899 15.9444 0.997× 1.2% / 2.4% 1.0% / 0.9%
Appendix C: GCC / AVX2+FMA — overall and review comparisons

Overall: 15f227178 → eba485c30 (14 fresh cases)

Case Before (µs) After (µs) Speedup CV before / after Run spread before / after
float 16×16 0.5707 0.0720 7.924× 1.4% / 1.3% 0.9% / 0.3%
float 64×64 10.6330 2.4777 4.291× 1.2% / 2.6% 0.3% / 2.3%
float 128×128 52.9245 18.2985 2.892× 2.9% / 2.3% 0.5% / 1.9%
float 256×256 277.5852 143.2835 1.937× 1.1% / 8.7% 0.8% / 1.3%
float 512×512 1881.8084 1103.0460 1.706× 2.6% / 1.9% 0.3% / 1.9%
float 1024×1024 13866.7695 8406.2973 1.650× 2.4% / 1.5% 1.1% / 1.0%
float 2048×2048 85655.7687 65872.5515 1.300× 1.8% / 1.7% 1.1% / 1.2%
double 16×16 0.5137 0.1212 4.240× 1.3% / 1.7% 4.3% / 3.0%
double 64×64 11.9751 5.1286 2.335× 5.8% / 7.9% 0.6% / 0.1%
double 128×128 66.5713 36.8364 1.807× 1.7% / 1.4% 0.6% / 0.4%
double 256×256 481.8207 310.4811 1.552× 1.3% / 2.3% 0.3% / 1.9%
double 512×512 3344.5898 2314.2324 1.445× 1.2% / 1.7% 1.4% / 0.6%
double 1024×1024 21422.2124 17475.8528 1.226× 2.1% / 1.9% 0.1% / 0.6%
double 2048×2048 147848.5783 136205.5195 1.085× 2.0% / 8.6% 1.1% / 1.9%

Review: 0c7bdb309 → eba485c30 (44 cases)

Case Before (µs) After (µs) Speedup CV before / after Run spread before / after
float 8×8 0.0264 0.0252 1.045× 1.7% / 0.9% 0.1% / 0.1%
float 16×16 0.0714 0.0720 0.992× 0.9% / 0.9% 0.0% / 0.2%
float 64×64 2.4884 2.4199 1.028× 1.4% / 1.2% 0.0% / 1.2%
float 128×128 18.5123 18.1125 1.022× 1.0% / 1.1% 0.7% / 0.8%
float 256×256 143.6751 142.7290 1.007× 1.3% / 2.0% 0.4% / 0.5%
float 512×512 1095.1285 1103.2076 0.993× 1.1% / 2.1% 0.5% / 0.6%
float 1024×1024 8693.2693 8437.8220 1.030× 3.6% / 1.8% 6.9% / 0.5%
float 64×8 0.4198 0.4138 1.015× 8.3% / 1.0% 2.8% / 0.3%
float 64×17 0.9630 0.9545 1.009× 1.1% / 1.3% 0.4% / 0.1%
float 128×8 1.6742 1.6627 1.007× 1.7% / 1.0% 3.0% / 0.5%
float 128×17 3.3206 3.2641 1.017× 0.8% / 2.4% 0.1% / 0.1%
double 8×8 0.0356 0.0353 1.008× 1.8% / 1.1% 1.6% / 0.4%
double 16×16 0.1214 0.1201 1.011× 1.6% / 1.6% 3.2% / 0.6%
double 64×64 5.0954 5.1248 0.994× 0.9% / 1.4% 0.1% / 1.8%
double 128×128 36.8140 36.3208 1.014× 1.6% / 1.2% 0.9% / 0.3%
double 256×256 322.4672 316.7923 1.018× 4.7% / 1.9% 6.8% / 3.4%
double 512×512 2316.2845 2303.0798 1.006× 1.6% / 2.1% 0.2% / 2.3%
double 1024×1024 17909.0610 17630.6995 1.016× 2.3% / 2.1% 0.6% / 1.0%
double 64×8 0.5990 0.5959 1.005× 1.4% / 0.8% 0.4% / 0.6%
double 64×17 1.6419 1.5930 1.031× 2.4% / 0.6% 1.0% / 0.4%
double 128×8 2.2146 2.2286 0.994× 1.7% / 1.0% 1.1% / 0.3%
double 128×17 5.9656 5.8751 1.015× 1.9% / 1.2% 1.1% / 0.2%
float 8×8 (row-major A) 0.0276 0.0277 0.995× 1.1% / 1.5% 0.1% / 1.5%
float 16×16 (row-major A) 0.0713 0.0795 0.897× 1.3% / 0.8% 0.9% / 0.2%
float 64×64 (row-major A) 2.5523 2.5327 1.008× 10.1% / 0.8% 0.3% / 1.1%
float 128×128 (row-major A) 19.1622 19.2573 0.995× 1.2% / 1.1% 0.4% / 1.4%
float 256×256 (row-major A) 145.2083 144.6909 1.004× 1.2% / 0.8% 0.4% / 0.2%
float 512×512 (row-major A) 1104.7704 1118.2685 0.988× 1.2% / 1.8% 1.8% / 0.2%
float 1024×1024 (row-major A) 8444.9349 8436.3459 1.001× 1.6% / 0.8% 2.1% / 0.4%
float 64×8 (row-major A) 0.4072 0.4137 0.984× 0.9% / 1.8% 0.4% / 0.6%
float 64×17 (row-major A) 1.1023 1.1158 0.988× 1.1% / 2.3% 0.2% / 1.3%
float 128×8 (row-major A) 1.6629 1.6852 0.987× 1.2% / 1.9% 0.1% / 1.1%
float 128×17 (row-major A) 4.2747 4.3057 0.993× 1.0% / 1.0% 1.5% / 0.7%
double 8×8 (row-major A) 0.0353 0.0369 0.957× 10.0% / 1.0% 0.8% / 0.3%
double 16×16 (row-major A) 0.1206 0.1238 0.974× 7.2% / 0.7% 2.4% / 1.8%
double 64×64 (row-major A) 5.3403 5.4640 0.977× 2.7% / 0.9% 0.3% / 1.3%
double 128×128 (row-major A) 38.5046 39.2370 0.981× 1.4% / 2.9% 0.5% / 2.2%
double 256×256 (row-major A) 318.9094 320.1195 0.996× 1.6% / 1.8% 0.8% / 0.5%
double 512×512 (row-major A) 2319.8356 2343.3450 0.990× 1.3% / 6.3% 1.0% / 0.2%
double 1024×1024 (row-major A) 17693.2769 17645.9338 1.003× 2.0% / 1.4% 0.3% / 1.3%
double 64×8 (row-major A) 0.6260 0.6461 0.969× 1.2% / 4.4% 0.8% / 4.6%
double 64×17 (row-major A) 1.7756 1.7852 0.995× 1.3% / 0.8% 0.4% / 1.5%
double 128×8 (row-major A) 2.3421 2.3685 0.989× 1.5% / 0.9% 1.6% / 0.2%
double 128×17 (row-major A) 6.8608 6.8619 1.000× 1.2% / 2.9% 0.2% / 0.6%
Appendix D: Clang / AVX2+FMA — overall and review comparisons

Overall: 15f227178 → eba485c30 (14 fresh cases)

Case Before (µs) After (µs) Speedup CV before / after Run spread before / after
float 16×16 0.8516 0.0695 12.262× 1.7% / 1.5% 0.2% / 1.2%
float 64×64 15.6504 2.5945 6.032× 1.5% / 2.2% 0.7% / 1.5%
float 128×128 74.1042 18.8137 3.939× 0.7% / 3.7% 0.0% / 1.4%
float 256×256 378.7491 177.9588 2.128× 1.7% / 1.1% 0.4% / 0.6%
float 512×512 2286.4058 1430.5906 1.598× 1.5% / 1.3% 0.1% / 1.0%
float 1024×1024 15176.6062 11345.0487 1.338× 1.0% / 1.2% 0.1% / 0.2%
float 2048×2048 108718.5517 92347.0618 1.177× 5.2% / 1.9% 0.4% / 1.4%
double 16×16 0.6933 0.1145 6.054× 1.6% / 1.7% 3.3% / 0.4%
double 64×64 14.1777 5.3982 2.626× 1.6% / 2.6% 0.4% / 7.2%
double 128×128 77.0706 37.3305 2.065× 1.0% / 1.5% 0.6% / 1.0%
double 256×256 505.4704 370.2104 1.365× 1.7% / 1.6% 1.3% / 0.9%
double 512×512 3672.2243 2886.6571 1.272× 1.9% / 3.1% 0.7% / 0.8%
double 1024×1024 26642.6718 23104.5532 1.153× 1.3% / 5.1% 1.5% / 1.8%
double 2048×2048 200970.5788 184114.9500 1.092× 1.2% / 2.0% 0.4% / 1.2%

Review: 0c7bdb309 → eba485c30 (44 cases) — retained earlier valid run; attempted refresh excluded

Case Before (µs) After (µs) Speedup CV before / after Run spread before / after
float 8×8 0.0263 0.0265 0.990× 3.1% / 2.6% 0.8% / 2.1%
float 16×16 0.0723 0.0727 0.994× 2.8% / 2.5% 1.1% / 3.1%
float 64×64 2.6447 2.6147 1.011× 2.2% / 2.1% 1.0% / 2.0%
float 128×128 20.1068 18.9734 1.060× 1.8% / 1.9% 1.2% / 0.4%
float 256×256 178.7853 179.8428 0.994× 1.1% / 1.6% 0.2% / 1.0%
float 512×512 1418.5618 1420.6191 0.999× 0.9% / 1.7% 0.6% / 1.4%
float 1024×1024 11275.5762 11261.5353 1.001× 1.2% / 0.9% 0.3% / 1.2%
float 64×8 0.4483 0.4442 1.009× 3.0% / 2.8% 4.9% / 1.9%
float 64×17 1.1414 1.1136 1.025× 6.3% / 3.9% 0.5% / 0.3%
float 128×8 1.7242 1.7860 0.965× 3.3% / 2.9% 0.8% / 4.1%
float 128×17 3.6715 3.5161 1.044× 2.2% / 1.9% 0.4% / 2.5%
double 8×8 0.0332 0.0333 0.997× 3.3% / 3.8% 4.5% / 0.8%
double 16×16 0.1231 0.1215 1.013× 2.6% / 3.0% 5.5% / 2.3%
double 64×64 5.3873 5.2220 1.032× 2.2% / 2.1% 0.6% / 3.0%
double 128×128 39.6703 38.3379 1.035× 1.6% / 2.6% 0.0% / 2.4%
double 256×256 368.2548 368.2284 1.000× 1.7% / 0.9% 0.6% / 0.5%
double 512×512 2856.9706 2855.3293 1.001× 1.7% / 2.1% 0.5% / 0.7%
double 1024×1024 22786.2134 22784.3764 1.000× 1.7% / 1.2% 1.0% / 0.2%
double 64×8 0.6452 0.6234 1.035× 2.1% / 2.0% 2.7% / 2.3%
double 64×17 1.6986 1.7008 0.999× 3.2% / 2.7% 0.5% / 1.7%
double 128×8 2.4676 2.3751 1.039× 2.6% / 2.7% 1.7% / 0.9%
double 128×17 6.1925 6.0064 1.031× 2.5% / 2.7% 2.2% / 2.2%
float 8×8 (row-major A) 0.0269 0.0272 0.992× 4.1% / 3.1% 6.9% / 1.2%
float 16×16 (row-major A) 0.0724 0.0736 0.984× 2.6% / 2.6% 2.8% / 4.0%
float 64×64 (row-major A) 2.8109 2.6711 1.052× 1.7% / 2.2% 2.0% / 1.9%
float 128×128 (row-major A) 20.8434 20.4351 1.020× 1.6% / 2.0% 1.6% / 0.1%
float 256×256 (row-major A) 179.3322 180.0101 0.996× 2.1% / 1.4% 1.0% / 0.0%
float 512×512 (row-major A) 1442.6603 1451.8157 0.994× 1.1% / 1.5% 0.1% / 0.5%
float 1024×1024 (row-major A) 11372.3660 11425.4480 0.995× 1.4% / 0.9% 0.4% / 0.3%
float 64×8 (row-major A) 0.4380 0.4394 0.997× 1.9% / 3.0% 5.2% / 0.3%
float 64×17 (row-major A) 1.3788 1.3556 1.017× 1.2% / 2.1% 0.2% / 0.6%
float 128×8 (row-major A) 1.7504 1.7250 1.015× 1.2% / 2.7% 1.2% / 0.8%
float 128×17 (row-major A) 6.0667 6.0013 1.011× 2.4% / 0.9% 1.1% / 0.2%
double 8×8 (row-major A) 0.0340 0.0347 0.979× 2.9% / 2.9% 1.1% / 5.8%
double 16×16 (row-major A) 0.1245 0.1291 0.965× 2.5% / 4.2% 4.3% / 1.5%
double 64×64 (row-major A) 5.4711 5.5156 0.992× 2.6% / 2.1% 1.8% / 1.1%
double 128×128 (row-major A) 40.4738 40.9422 0.989× 1.5% / 1.4% 1.2% / 1.2%
double 256×256 (row-major A) 376.2471 377.5466 0.997× 2.4% / 2.4% 1.9% / 1.3%
double 512×512 (row-major A) 2894.8179 2893.7299 1.000× 2.2% / 1.6% 2.8% / 1.7%
double 1024×1024 (row-major A) 23036.7122 23123.5302 0.996× 1.6% / 2.6% 0.3% / 0.6%
double 64×8 (row-major A) 0.6588 0.6623 0.995× 2.7% / 2.5% 0.1% / 0.5%
double 64×17 (row-major A) 1.9987 2.0610 0.970× 2.1% / 4.2% 2.2% / 4.8%
double 128×8 (row-major A) 2.4868 2.5158 0.988× 2.4% / 2.6% 0.8% / 0.3%
double 128×17 (row-major A) 8.3608 8.4229 0.993× 1.0% / 1.5% 0.3% / 0.4%
Appendix E: Provenance and validation details

Measurement conditions: 2026-09-17, 13th Gen Intel(R) Core(TM) i7-13700HX under WSL2; GCC 13.3.0, Clang 18.1.3, Google Benchmark 1.9.5. Common flags: -std=c++14 -O3 -DNDEBUG -march=x86-64 -mtune=generic. SSE2 adds -msse2 -mno-avx -mno-fma; AVX2+FMA adds -mavx2 -mfma. Compiler/Eigen macros were checked. Each pair uses identical sources, flags and filters, without external BLAS or OpenMP.

One benchmark ran at a time on CPU 2, in A/B/A/B order, with eight repetitions of ≥0.08 CPU seconds per case per invocation. Compilation and correctness runs finished before timing. System-wide process monitoring rejected sustained competing work; Windows scheduling, clocks and thermal policy were uncontrolled. All included cases passed the benchmark’s residual check. An interrupted Clang/SSE2 overall attempt was discarded and restarted before the request to omit additional repeats. The later interrupted Clang/AVX2 review refresh was discarded without rerunning it.

The reported ratio uses CPU time. Elapsed-time ratios differed by up to 4.3% in the overall grid and 7.0% in the review grid; small differences near this variability are inconclusive. No additional targeted repeats were performed; noisy cases remain inconclusive.

MR head at publication: 0154510695ec7051f030808e520fd3c750c88015 (rebased during the measurements). Revisions used for benchmarking: original 15f22717866165e381d36d52209400daaf39a386, previous MR 0c7bdb309a94d4a297abacf48900decbdb3a2f01, candidate eba485c3099b2a9712c76d0201a684d578f5dbcf. The timed expression is x = b; a.triangularView<Lower>().solveInPlace(x); setup and residual checks are excluded. RHS storage is column-major. Speedup = before/after, each time being the median of two invocation medians. CV is the maximum within-invocation coefficient of variation; run spread is max(invocation medians) / min(invocation medians) - 1.

These are local before/after results motivated by the Zen 5 benchmark; they are not new Zen 5 or vendor-relative measurements. The Zen 5 host was busy and was not timed in this round.

The retained Clang/AVX2 review comparison used -std=c++14 -O3 -DNDEBUG -mavx2 -mfma with the same benchmark source, machine, affinity, A/B/A/B order and repetition settings. Its rebuilt before and after executables, with the explicit generic x86-64 flags above, have identical SHA-256 hashes to that earlier pair. All four overall comparisons and the other three review comparisons are fresh.

Fresh SSE2 validation also passed trsm_packet and trsm_packet_cancellation with GCC 13.3 and Clang 18.1, both scalar parts, r2 s42. The AVX2 correctness results below apply to the same unchanged candidate source.

The new regression explicitly exercises the shared implementation. Its controlled cache settings keep cancelling terms within a diagonal panel: the harness’s synthetic one-row AVX-512 double panels instead split cancellation across separate GEMM updates, which also fails before this MR. Existing small-block tests remain enabled.

Validation passed with both GCC and Clang: six CMake targets per compiler (trsm_packet, cancellation, and fast-math cancellation, each split by scalar), plus C++14 -O0, triangular and Cholesky integration checks including complex scalars. Additional checks cover SSE2, scalar-only, an eight-register budget, alternate layout/index settings, and ASan/UBSan. Cancellation cases cover upper/lower, unit/nonunit, left/right, packet tails, mixed lanes, poisoned unused coefficients, and orders 4–256. The regression fails at the previous MR revision. Shared-kernel AVX-512 passed natively on AMD RYZEN AI MAX+ 395 with GCC and Clang; AArch64 NEON passed under QEMU with GCC 14. No timings were taken on that loaded remote machine. Other ISAs and MSVC were not validated in this round.

Disassembly of the current AVX2 binaries confirms integer exponent checks in the final-row guard (scalar integer checks with GCC, packed integer masks/comparisons with Clang). The guard and retry stay outside the substitution loops. This adds work for row-major triangles, most visible in tiny solves; measured costs are included above.

Edited by Rasmus Munk Larsen

Merge request reports

Loading
Loading