Core: Vectorize triangular solves across right-hand sides
This adds a shared SIMD kernel for TRSM, processing independent right-hand sides in packet lanes. GEBP’s register budget determines the RHS packet group, and Eigen’s configured L1 cache size controls when a solve bypasses matrix-multiply packing. Fresh measurements cover GCC and Clang × SSE2 and AVX2+FMA, using the same source revision and benchmark in all four configurations.
Measured revision: eba485c30, before the branch was rebased to 015451069. The TRSM kernel, benchmark and focused test sources are byte-identical after rebasing; the rebased tree has not been timed.
Overall speedup: pre-optimization Eigen 15f227178 → eba485c30. Geometric means over seven square sizes, 16–2048, for each scalar. Column-major lower-triangular left solve, including the RHS copy; greater than 1× is faster.
| Compiler / ISA | float geomean | double geomean | Range across 14 cases |
|---|---|---|---|
| GCC / SSE2 | 1.548× | 1.282× | 1.005–3.683× |
| Clang / SSE2 | 1.549× | 1.388× | 1.024–3.964× |
| GCC / AVX2+FMA | 2.548× | 1.765× | 1.085–7.924× |
| Clang / AVX2+FMA | 2.859× | 1.841× | 1.092–12.262× |
Review-fix cost: 0c7bdb309 → eba485c30. The same 44-case grid in each configuration, including both triangle layouts, square and narrow RHS cases. The Clang/AVX2 review row retains the earlier valid run; its attempted refresh was interrupted by background work and excluded. Both rebuilt binaries are byte-identical to the earlier pair.
| Compiler / ISA | Column-major geomean (22) | Row-major geomean (22) | Lowest observed ratio |
|---|---|---|---|
| GCC / SSE2 | 0.990× | 0.983× | double 64×17: 0.878×; 20.3% variation |
| Clang / SSE2 | 0.996× | 1.001× | double 8×8 (row-major A): 0.972×; 1.6% variation |
| GCC / AVX2+FMA | 1.012× | 0.985× | float 16×16 (row-major A): 0.897×; 1.3% variation |
| Clang / AVX2+FMA (earlier run) | 1.012× | 0.997× | double 16×16 (row-major A): 0.965×; 4.3% variation |
Variation in the last column is the maximum CV or run spread, not a confidence interval. The GCC/SSE2 0.878× case has 20.3% variation and is inconclusive. Tiny row-major float solves retain a cost: GCC/SSE2 8×8 takes 5.6% longer, and GCC/AVX2 16×16 takes 11.5% longer (0.0713 → 0.0795 µs), consistent with the earlier observed guard cost. Large SSE2 double gains are only about 1–3% and generally near the observed variability. Earlier localized 3–6% generalization costs against 76c9fb280 remain limitations. No additional targeted repeats were performed.
The kernel transposes a bounded RHS tile, solves register blocks of rows, and transposes the results back. Four register rows and at most two RHS packets preserve the established AVX2 schedule. WorkspaceRows = 128 bounds stack capacity separately from the cache policy: the direct-solve footprint estimate is n * (packets * PacketSize + RegisterRows + 1) * sizeof(Scalar), plus one packet transpose, against half of L1. Cache discovery and overrides use Eigen’s existing machinery; outer L2/L3 blocking remains in place. These are conservative defaults, not optimal tuning claims for every platform.
The fast path requires vectorized built-in IEEE binary float/double and compile-time unit inner stride. Partial RHS packets, complex/custom scalars and non-unit strides retain their existing paths. AVX-512 builds select their existing specialized TRSM kernel by default; unifying that implementation remains separate work.
Update: Thanks to @florian360 for the finite-result overflow example and @onalante-ebay for the expression/style suggestions. eba485c30 checks the final solved row of each row-major RHS tile before writeback, and retries a nonfinite tile using the original row-dot-product accumulation order. Every later row consumes all earlier solved rows, so nonfinite lanes propagate to the final row. Integer exponent-bit classification and Eigen’s fast-math barrier keep the check effective under Clang -ffast-math. Reciprocal initialization now uses a map of the active diagonal; the other suggestions use EIGEN_IF_CONSTEXPR and numext::round_down with an Index conversion.
The retry restores the previous diagonal-kernel behavior; it is not a general overflow-safe solver. The existing default AVX-512 specialization also overflows on the example before this MR and is unchanged.
Measurements: 2026-09-17, 13th Gen Intel(R) Core(TM) i7-13700HX under WSL2, GCC 13.3.0 / Clang 18.1.3. CPU-time medians, one process pinned to CPU 2, A/B/A/B invocations, eight repetitions per case. Windows clocks, scheduling and thermal policy were uncontrolled. Full flags, protocol and limitations are in Appendix E.
GCC and Clang correctness validation passed, including the additional SSE2 runs. The regression and prior shared-kernel AVX-512/NEON validation are detailed in Appendix E.
Appendix A: GCC / SSE2 — overall and review comparisons
Overall: 15f227178 → eba485c30 (14 fresh cases)
| Case | Before (µs) | After (µs) | Speedup | CV before / after | Run spread before / after |
|---|---|---|---|---|---|
| float 16×16 | 0.5576 | 0.1514 | 3.683× | 3.1% / 1.6% | 1.7% / 0.6% |
| float 64×64 | 13.7400 | 6.8228 | 2.014× | 1.3% / 1.6% | 0.8% / 0.5% |
| float 128×128 | 91.2179 | 52.3067 | 1.744× | 1.0% / 1.2% | 0.1% / 0.3% |
| float 256×256 | 561.0331 | 473.9017 | 1.184× | 1.4% / 1.7% | 1.7% / 0.6% |
| float 512×512 | 4234.2389 | 3623.6479 | 1.169× | 1.9% / 2.1% | 1.9% / 1.0% |
| float 1024×1024 | 31527.6154 | 27794.8264 | 1.134× | 0.7% / 1.6% | 2.4% / 2.1% |
| float 2048×2048 | 227246.9757 | 216860.4282 | 1.048× | 0.9% / 1.8% | 1.6% / 0.5% |
| double 16×16 | 0.5667 | 0.2881 | 1.967× | 6.5% / 1.2% | 16.2% / 1.2% |
| double 64×64 | 22.4007 | 14.1618 | 1.582× | 2.0% / 1.2% | 0.7% / 0.9% |
| double 128×128 | 181.7513 | 110.6166 | 1.643× | 1.9% / 1.7% | 1.9% / 0.7% |
| double 256×256 | 1158.5543 | 1091.9134 | 1.061× | 1.1% / 2.9% | 1.3% / 1.9% |
| double 512×512 | 8358.3884 | 8231.0639 | 1.015× | 2.4% / 2.6% | 1.0% / 0.4% |
| double 1024×1024 | 60585.0261 | 59018.6196 | 1.027× | 2.7% / 1.6% | 0.1% / 0.8% |
| double 2048×2048 | 451579.4922 | 449198.4952 | 1.005× | 1.3% / 2.4% | 0.5% / 0.2% |
Review: 0c7bdb309 → eba485c30 (44 cases)
| Case | Before (µs) | After (µs) | Speedup | CV before / after | Run spread before / after |
|---|---|---|---|---|---|
| float 8×8 | 0.0345 | 0.0345 | 0.999× | 2.3% / 1.5% | 1.0% / 0.0% |
| float 16×16 | 0.1527 | 0.1536 | 0.994× | 1.5% / 1.3% | 2.1% / 0.2% |
| float 64×64 | 6.8934 | 6.9861 | 0.987× | 2.6% / 2.8% | 1.8% / 3.6% |
| float 128×128 | 52.7329 | 52.8021 | 0.999× | 1.1% / 1.4% | 1.1% / 0.2% |
| float 256×256 | 484.8405 | 486.8830 | 0.996× | 3.8% / 1.9% | 3.7% / 5.1% |
| float 512×512 | 3641.6702 | 3660.9862 | 0.995× | 1.2% / 1.5% | 2.0% / 2.7% |
| float 1024×1024 | 27734.8454 | 28237.3886 | 0.982× | 1.3% / 1.7% | 1.3% / 0.1% |
| float 64×8 | 0.8637 | 0.8777 | 0.984× | 1.4% / 3.7% | 0.1% / 1.0% |
| float 64×17 | 2.2002 | 2.2131 | 0.994× | 3.4% / 5.2% | 1.7% / 0.3% |
| float 128×8 | 3.2751 | 3.2834 | 0.997× | 1.3% / 0.9% | 0.8% / 1.4% |
| float 128×17 | 7.7792 | 7.6997 | 1.010× | 2.4% / 1.1% | 0.4% / 1.2% |
| double 8×8 | 0.0620 | 0.0567 | 1.094× | 7.6% / 1.0% | 19.8% / 1.2% |
| double 16×16 | 0.2920 | 0.2906 | 1.005× | 2.6% / 3.2% | 2.8% / 3.5% |
| double 64×64 | 14.3612 | 14.3880 | 0.998× | 1.3% / 1.8% | 0.8% / 2.4% |
| double 128×128 | 111.0311 | 112.8277 | 0.984× | 1.1% / 1.9% | 0.6% / 2.3% |
| double 256×256 | 1072.1497 | 1084.9934 | 0.988× | 1.4% / 1.9% | 0.1% / 0.1% |
| double 512×512 | 8000.9396 | 8078.9894 | 0.990× | 1.4% / 4.2% | 1.7% / 0.2% |
| double 1024×1024 | 59348.5789 | 59608.0896 | 0.996× | 2.3% / 1.5% | 0.5% / 1.4% |
| double 64×8 | 1.7688 | 1.8005 | 0.982× | 1.2% / 2.7% | 0.3% / 4.0% |
| double 64×17 | 3.8956 | 4.4344 | 0.878× | 1.1% / 10.0% | 0.1% / 20.3% |
| double 128×8 | 6.8641 | 6.9570 | 0.987× | 1.4% / 6.1% | 0.0% / 0.3% |
| double 128×17 | 15.1213 | 15.8217 | 0.956× | 2.2% / 1.9% | 0.2% / 0.8% |
| float 8×8 (row-major A) | 0.0351 | 0.0371 | 0.947× | 1.1% / 1.8% | 0.2% / 1.4% |
| float 16×16 (row-major A) | 0.1523 | 0.1549 | 0.984× | 2.3% / 1.4% | 1.7% / 1.8% |
| float 64×64 (row-major A) | 7.3920 | 7.0280 | 1.052× | 10.1% / 3.1% | 15.7% / 2.5% |
| float 128×128 (row-major A) | 52.2264 | 52.8342 | 0.988× | 6.3% / 2.2% | 0.3% / 1.2% |
| float 256×256 (row-major A) | 477.2216 | 483.9215 | 0.986× | 1.3% / 1.4% | 0.5% / 1.2% |
| float 512×512 (row-major A) | 3599.2099 | 3976.9805 | 0.905× | 1.1% / 6.8% | 0.4% / 17.1% |
| float 1024×1024 (row-major A) | 27846.4506 | 28681.4812 | 0.971× | 1.8% / 3.1% | 1.0% / 3.7% |
| float 64×8 (row-major A) | 0.8641 | 0.8646 | 0.999× | 0.9% / 1.1% | 1.8% / 0.5% |
| float 64×17 (row-major A) | 2.1874 | 2.2039 | 0.993× | 0.8% / 3.2% | 0.5% / 1.5% |
| float 128×8 (row-major A) | 3.2616 | 3.2842 | 0.993× | 1.1% / 1.3% | 0.1% / 2.0% |
| float 128×17 (row-major A) | 8.4872 | 8.5835 | 0.989× | 1.1% / 1.9% | 0.4% / 1.6% |
| double 8×8 (row-major A) | 0.0606 | 0.0589 | 1.030× | 11.8% / 1.6% | 14.5% / 0.9% |
| double 16×16 (row-major A) | 0.2815 | 0.2823 | 0.997× | 6.2% / 1.4% | 2.9% / 0.9% |
| double 64×64 (row-major A) | 14.3667 | 14.4944 | 0.991× | 1.4% / 1.3% | 1.2% / 0.5% |
| double 128×128 (row-major A) | 111.1262 | 112.0957 | 0.991× | 3.6% / 1.2% | 1.4% / 0.5% |
| double 256×256 (row-major A) | 1079.7081 | 1202.0414 | 0.898× | 2.0% / 8.9% | 1.2% / 12.0% |
| double 512×512 (row-major A) | 8034.1150 | 8083.5608 | 0.994× | 2.0% / 1.5% | 1.1% / 0.5% |
| double 1024×1024 (row-major A) | 59770.7063 | 61668.9123 | 0.969× | 1.8% / 11.7% | 0.7% / 0.7% |
| double 64×8 (row-major A) | 1.8216 | 1.8725 | 0.973× | 1.5% / 8.7% | 1.2% / 5.9% |
| double 64×17 (row-major A) | 4.0648 | 4.2599 | 0.954× | 2.0% / 7.6% | 0.6% / 7.5% |
| double 128×8 (row-major A) | 7.0935 | 6.9709 | 1.018× | 1.6% / 1.6% | 2.7% / 0.1% |
| double 128×17 (row-major A) | 16.1915 | 16.0425 | 1.009× | 2.4% / 0.8% | 1.8% / 0.5% |
Appendix B: Clang / SSE2 — overall and review comparisons
Overall: 15f227178 → eba485c30 (14 fresh cases)
| Case | Before (µs) | After (µs) | Speedup | CV before / after | Run spread before / after |
|---|---|---|---|---|---|
| float 16×16 | 0.6663 | 0.1681 | 3.964× | 7.3% / 3.5% | 7.9% / 15.2% |
| float 64×64 | 14.0409 | 7.6860 | 1.827× | 4.9% / 6.0% | 0.7% / 17.1% |
| float 128×128 | 91.6284 | 53.4162 | 1.715× | 4.3% / 2.8% | 5.5% / 2.1% |
| float 256×256 | 592.0898 | 456.4255 | 1.297× | 8.0% / 1.3% | 14.8% / 0.3% |
| float 512×512 | 3923.3726 | 3447.7765 | 1.138× | 3.9% / 2.7% | 1.0% / 0.1% |
| float 1024×1024 | 29448.5509 | 27023.4574 | 1.090× | 1.7% / 3.6% | 2.0% / 3.7% |
| float 2048×2048 | 221602.5165 | 206758.0272 | 1.072× | 6.8% / 3.0% | 2.9% / 0.5% |
| double 16×16 | 0.6979 | 0.2851 | 2.448× | 4.3% / 1.3% | 18.1% / 1.1% |
| double 64×64 | 24.3558 | 14.2676 | 1.707× | 3.6% / 1.2% | 16.3% / 1.3% |
| double 128×128 | 187.7486 | 111.6513 | 1.682× | 6.3% / 1.8% | 15.1% / 1.1% |
| double 256×256 | 1236.0550 | 1049.8741 | 1.177× | 5.8% / 1.3% | 18.4% / 0.2% |
| double 512×512 | 9015.4704 | 7908.4092 | 1.140× | 3.1% / 2.6% | 21.2% / 3.4% |
| double 1024×1024 | 60647.5749 | 59039.0686 | 1.027× | 1.4% / 2.4% | 2.6% / 2.3% |
| double 2048×2048 | 459841.1825 | 448971.9905 | 1.024× | 7.0% / 5.0% | 2.4% / 0.1% |
Review: 0c7bdb309 → eba485c30 (44 cases)
| Case | Before (µs) | After (µs) | Speedup | CV before / after | Run spread before / after |
|---|---|---|---|---|---|
| float 8×8 | 0.0351 | 0.0352 | 0.997× | 1.2% / 1.4% | 0.4% / 3.0% |
| float 16×16 | 0.1539 | 0.1547 | 0.995× | 1.0% / 1.4% | 0.5% / 1.4% |
| float 64×64 | 6.8761 | 6.9534 | 0.989× | 1.0% / 1.2% | 0.3% / 0.8% |
| float 128×128 | 52.8017 | 52.9650 | 0.997× | 1.2% / 1.6% | 1.2% / 0.1% |
| float 256×256 | 457.7716 | 459.6645 | 0.996× | 5.4% / 6.9% | 0.2% / 1.9% |
| float 512×512 | 3475.3373 | 3476.3559 | 1.000× | 1.4% / 1.3% | 0.8% / 0.4% |
| float 1024×1024 | 26646.3233 | 26526.2724 | 1.005× | 2.2% / 6.7% | 1.9% / 0.3% |
| float 64×8 | 0.8711 | 0.8809 | 0.989× | 1.4% / 4.3% | 0.4% / 1.7% |
| float 64×17 | 2.0853 | 2.0740 | 1.005× | 3.5% / 4.9% | 1.1% / 0.7% |
| float 128×8 | 3.2885 | 3.2787 | 1.003× | 1.2% / 1.2% | 0.6% / 2.1% |
| float 128×17 | 7.5892 | 7.6158 | 0.997× | 0.9% / 1.8% | 0.9% / 0.9% |
| double 8×8 | 0.0574 | 0.0578 | 0.992× | 0.9% / 1.6% | 1.9% / 1.5% |
| double 16×16 | 0.2865 | 0.2878 | 0.996× | 1.7% / 2.4% | 0.4% / 1.5% |
| double 64×64 | 14.2430 | 14.4612 | 0.985× | 3.4% / 3.7% | 0.4% / 3.1% |
| double 128×128 | 111.1773 | 114.0092 | 0.975× | 1.6% / 1.6% | 0.0% / 0.6% |
| double 256×256 | 1071.4788 | 1076.8173 | 0.995× | 4.5% / 2.0% | 2.8% / 3.6% |
| double 512×512 | 7981.5099 | 7974.0080 | 1.001× | 1.8% / 5.4% | 2.5% / 3.5% |
| double 1024×1024 | 60038.1918 | 59618.3565 | 1.007× | 3.0% / 1.5% | 0.6% / 1.2% |
| double 64×8 | 1.7971 | 1.8094 | 0.993× | 1.4% / 2.1% | 2.9% / 1.0% |
| double 64×17 | 4.0454 | 4.0298 | 1.004× | 1.7% / 1.5% | 4.3% / 1.2% |
| double 128×8 | 7.1518 | 7.1595 | 0.999× | 3.9% / 1.5% | 1.9% / 1.9% |
| double 128×17 | 15.5134 | 15.7093 | 0.988× | 1.0% / 1.7% | 0.3% / 1.8% |
| float 8×8 (row-major A) | 0.0368 | 0.0366 | 1.005× | 2.8% / 0.9% | 4.7% / 1.3% |
| float 16×16 (row-major A) | 0.1534 | 0.1539 | 0.997× | 1.4% / 1.4% | 0.6% / 0.1% |
| float 64×64 (row-major A) | 6.9136 | 6.8988 | 1.002× | 1.2% / 1.5% | 0.2% / 0.5% |
| float 128×128 (row-major A) | 52.0302 | 52.4365 | 0.992× | 1.7% / 1.9% | 0.2% / 0.6% |
| float 256×256 (row-major A) | 468.0477 | 465.6385 | 1.005× | 3.4% / 1.0% | 1.0% / 1.5% |
| float 512×512 (row-major A) | 3563.3720 | 3490.4749 | 1.021× | 7.6% / 1.1% | 2.3% / 0.1% |
| float 1024×1024 (row-major A) | 27218.0376 | 26813.7949 | 1.015× | 2.6% / 1.6% | 5.2% / 0.0% |
| float 64×8 (row-major A) | 0.8861 | 0.8724 | 1.016× | 4.2% / 2.0% | 3.4% / 0.7% |
| float 64×17 (row-major A) | 2.1651 | 2.1744 | 0.996× | 1.6% / 4.2% | 0.2% / 1.3% |
| float 128×8 (row-major A) | 3.2273 | 3.2611 | 0.990× | 1.3% / 1.4% | 0.3% / 0.8% |
| float 128×17 (row-major A) | 8.5284 | 8.5458 | 0.998× | 1.1% / 1.4% | 0.6% / 0.0% |
| double 8×8 (row-major A) | 0.0587 | 0.0605 | 0.972× | 1.6% / 0.9% | 0.6% / 0.1% |
| double 16×16 (row-major A) | 0.2913 | 0.2930 | 0.994× | 1.0% / 1.8% | 0.9% / 0.2% |
| double 64×64 (row-major A) | 14.3652 | 14.4330 | 0.995× | 2.4% / 1.5% | 0.5% / 0.3% |
| double 128×128 (row-major A) | 110.8081 | 111.5462 | 0.993× | 2.2% / 1.5% | 1.5% / 1.6% |
| double 256×256 (row-major A) | 1052.4882 | 1041.6074 | 1.010× | 2.8% / 1.2% | 3.5% / 0.2% |
| double 512×512 (row-major A) | 7878.6308 | 7896.4791 | 0.998× | 2.3% / 3.7% | 1.7% / 0.1% |
| double 1024×1024 (row-major A) | 59710.5261 | 59372.7654 | 1.006× | 3.2% / 1.3% | 0.6% / 0.4% |
| double 64×8 (row-major A) | 1.8327 | 1.8034 | 1.016× | 6.8% / 1.0% | 0.9% / 0.2% |
| double 64×17 (row-major A) | 4.0394 | 4.0292 | 1.003× | 10.4% / 1.5% | 2.3% / 1.3% |
| double 128×8 (row-major A) | 6.9353 | 6.9864 | 0.993× | 1.4% / 1.2% | 0.6% / 0.5% |
| double 128×17 (row-major A) | 15.8899 | 15.9444 | 0.997× | 1.2% / 2.4% | 1.0% / 0.9% |
Appendix C: GCC / AVX2+FMA — overall and review comparisons
Overall: 15f227178 → eba485c30 (14 fresh cases)
| Case | Before (µs) | After (µs) | Speedup | CV before / after | Run spread before / after |
|---|---|---|---|---|---|
| float 16×16 | 0.5707 | 0.0720 | 7.924× | 1.4% / 1.3% | 0.9% / 0.3% |
| float 64×64 | 10.6330 | 2.4777 | 4.291× | 1.2% / 2.6% | 0.3% / 2.3% |
| float 128×128 | 52.9245 | 18.2985 | 2.892× | 2.9% / 2.3% | 0.5% / 1.9% |
| float 256×256 | 277.5852 | 143.2835 | 1.937× | 1.1% / 8.7% | 0.8% / 1.3% |
| float 512×512 | 1881.8084 | 1103.0460 | 1.706× | 2.6% / 1.9% | 0.3% / 1.9% |
| float 1024×1024 | 13866.7695 | 8406.2973 | 1.650× | 2.4% / 1.5% | 1.1% / 1.0% |
| float 2048×2048 | 85655.7687 | 65872.5515 | 1.300× | 1.8% / 1.7% | 1.1% / 1.2% |
| double 16×16 | 0.5137 | 0.1212 | 4.240× | 1.3% / 1.7% | 4.3% / 3.0% |
| double 64×64 | 11.9751 | 5.1286 | 2.335× | 5.8% / 7.9% | 0.6% / 0.1% |
| double 128×128 | 66.5713 | 36.8364 | 1.807× | 1.7% / 1.4% | 0.6% / 0.4% |
| double 256×256 | 481.8207 | 310.4811 | 1.552× | 1.3% / 2.3% | 0.3% / 1.9% |
| double 512×512 | 3344.5898 | 2314.2324 | 1.445× | 1.2% / 1.7% | 1.4% / 0.6% |
| double 1024×1024 | 21422.2124 | 17475.8528 | 1.226× | 2.1% / 1.9% | 0.1% / 0.6% |
| double 2048×2048 | 147848.5783 | 136205.5195 | 1.085× | 2.0% / 8.6% | 1.1% / 1.9% |
Review: 0c7bdb309 → eba485c30 (44 cases)
| Case | Before (µs) | After (µs) | Speedup | CV before / after | Run spread before / after |
|---|---|---|---|---|---|
| float 8×8 | 0.0264 | 0.0252 | 1.045× | 1.7% / 0.9% | 0.1% / 0.1% |
| float 16×16 | 0.0714 | 0.0720 | 0.992× | 0.9% / 0.9% | 0.0% / 0.2% |
| float 64×64 | 2.4884 | 2.4199 | 1.028× | 1.4% / 1.2% | 0.0% / 1.2% |
| float 128×128 | 18.5123 | 18.1125 | 1.022× | 1.0% / 1.1% | 0.7% / 0.8% |
| float 256×256 | 143.6751 | 142.7290 | 1.007× | 1.3% / 2.0% | 0.4% / 0.5% |
| float 512×512 | 1095.1285 | 1103.2076 | 0.993× | 1.1% / 2.1% | 0.5% / 0.6% |
| float 1024×1024 | 8693.2693 | 8437.8220 | 1.030× | 3.6% / 1.8% | 6.9% / 0.5% |
| float 64×8 | 0.4198 | 0.4138 | 1.015× | 8.3% / 1.0% | 2.8% / 0.3% |
| float 64×17 | 0.9630 | 0.9545 | 1.009× | 1.1% / 1.3% | 0.4% / 0.1% |
| float 128×8 | 1.6742 | 1.6627 | 1.007× | 1.7% / 1.0% | 3.0% / 0.5% |
| float 128×17 | 3.3206 | 3.2641 | 1.017× | 0.8% / 2.4% | 0.1% / 0.1% |
| double 8×8 | 0.0356 | 0.0353 | 1.008× | 1.8% / 1.1% | 1.6% / 0.4% |
| double 16×16 | 0.1214 | 0.1201 | 1.011× | 1.6% / 1.6% | 3.2% / 0.6% |
| double 64×64 | 5.0954 | 5.1248 | 0.994× | 0.9% / 1.4% | 0.1% / 1.8% |
| double 128×128 | 36.8140 | 36.3208 | 1.014× | 1.6% / 1.2% | 0.9% / 0.3% |
| double 256×256 | 322.4672 | 316.7923 | 1.018× | 4.7% / 1.9% | 6.8% / 3.4% |
| double 512×512 | 2316.2845 | 2303.0798 | 1.006× | 1.6% / 2.1% | 0.2% / 2.3% |
| double 1024×1024 | 17909.0610 | 17630.6995 | 1.016× | 2.3% / 2.1% | 0.6% / 1.0% |
| double 64×8 | 0.5990 | 0.5959 | 1.005× | 1.4% / 0.8% | 0.4% / 0.6% |
| double 64×17 | 1.6419 | 1.5930 | 1.031× | 2.4% / 0.6% | 1.0% / 0.4% |
| double 128×8 | 2.2146 | 2.2286 | 0.994× | 1.7% / 1.0% | 1.1% / 0.3% |
| double 128×17 | 5.9656 | 5.8751 | 1.015× | 1.9% / 1.2% | 1.1% / 0.2% |
| float 8×8 (row-major A) | 0.0276 | 0.0277 | 0.995× | 1.1% / 1.5% | 0.1% / 1.5% |
| float 16×16 (row-major A) | 0.0713 | 0.0795 | 0.897× | 1.3% / 0.8% | 0.9% / 0.2% |
| float 64×64 (row-major A) | 2.5523 | 2.5327 | 1.008× | 10.1% / 0.8% | 0.3% / 1.1% |
| float 128×128 (row-major A) | 19.1622 | 19.2573 | 0.995× | 1.2% / 1.1% | 0.4% / 1.4% |
| float 256×256 (row-major A) | 145.2083 | 144.6909 | 1.004× | 1.2% / 0.8% | 0.4% / 0.2% |
| float 512×512 (row-major A) | 1104.7704 | 1118.2685 | 0.988× | 1.2% / 1.8% | 1.8% / 0.2% |
| float 1024×1024 (row-major A) | 8444.9349 | 8436.3459 | 1.001× | 1.6% / 0.8% | 2.1% / 0.4% |
| float 64×8 (row-major A) | 0.4072 | 0.4137 | 0.984× | 0.9% / 1.8% | 0.4% / 0.6% |
| float 64×17 (row-major A) | 1.1023 | 1.1158 | 0.988× | 1.1% / 2.3% | 0.2% / 1.3% |
| float 128×8 (row-major A) | 1.6629 | 1.6852 | 0.987× | 1.2% / 1.9% | 0.1% / 1.1% |
| float 128×17 (row-major A) | 4.2747 | 4.3057 | 0.993× | 1.0% / 1.0% | 1.5% / 0.7% |
| double 8×8 (row-major A) | 0.0353 | 0.0369 | 0.957× | 10.0% / 1.0% | 0.8% / 0.3% |
| double 16×16 (row-major A) | 0.1206 | 0.1238 | 0.974× | 7.2% / 0.7% | 2.4% / 1.8% |
| double 64×64 (row-major A) | 5.3403 | 5.4640 | 0.977× | 2.7% / 0.9% | 0.3% / 1.3% |
| double 128×128 (row-major A) | 38.5046 | 39.2370 | 0.981× | 1.4% / 2.9% | 0.5% / 2.2% |
| double 256×256 (row-major A) | 318.9094 | 320.1195 | 0.996× | 1.6% / 1.8% | 0.8% / 0.5% |
| double 512×512 (row-major A) | 2319.8356 | 2343.3450 | 0.990× | 1.3% / 6.3% | 1.0% / 0.2% |
| double 1024×1024 (row-major A) | 17693.2769 | 17645.9338 | 1.003× | 2.0% / 1.4% | 0.3% / 1.3% |
| double 64×8 (row-major A) | 0.6260 | 0.6461 | 0.969× | 1.2% / 4.4% | 0.8% / 4.6% |
| double 64×17 (row-major A) | 1.7756 | 1.7852 | 0.995× | 1.3% / 0.8% | 0.4% / 1.5% |
| double 128×8 (row-major A) | 2.3421 | 2.3685 | 0.989× | 1.5% / 0.9% | 1.6% / 0.2% |
| double 128×17 (row-major A) | 6.8608 | 6.8619 | 1.000× | 1.2% / 2.9% | 0.2% / 0.6% |
Appendix D: Clang / AVX2+FMA — overall and review comparisons
Overall: 15f227178 → eba485c30 (14 fresh cases)
| Case | Before (µs) | After (µs) | Speedup | CV before / after | Run spread before / after |
|---|---|---|---|---|---|
| float 16×16 | 0.8516 | 0.0695 | 12.262× | 1.7% / 1.5% | 0.2% / 1.2% |
| float 64×64 | 15.6504 | 2.5945 | 6.032× | 1.5% / 2.2% | 0.7% / 1.5% |
| float 128×128 | 74.1042 | 18.8137 | 3.939× | 0.7% / 3.7% | 0.0% / 1.4% |
| float 256×256 | 378.7491 | 177.9588 | 2.128× | 1.7% / 1.1% | 0.4% / 0.6% |
| float 512×512 | 2286.4058 | 1430.5906 | 1.598× | 1.5% / 1.3% | 0.1% / 1.0% |
| float 1024×1024 | 15176.6062 | 11345.0487 | 1.338× | 1.0% / 1.2% | 0.1% / 0.2% |
| float 2048×2048 | 108718.5517 | 92347.0618 | 1.177× | 5.2% / 1.9% | 0.4% / 1.4% |
| double 16×16 | 0.6933 | 0.1145 | 6.054× | 1.6% / 1.7% | 3.3% / 0.4% |
| double 64×64 | 14.1777 | 5.3982 | 2.626× | 1.6% / 2.6% | 0.4% / 7.2% |
| double 128×128 | 77.0706 | 37.3305 | 2.065× | 1.0% / 1.5% | 0.6% / 1.0% |
| double 256×256 | 505.4704 | 370.2104 | 1.365× | 1.7% / 1.6% | 1.3% / 0.9% |
| double 512×512 | 3672.2243 | 2886.6571 | 1.272× | 1.9% / 3.1% | 0.7% / 0.8% |
| double 1024×1024 | 26642.6718 | 23104.5532 | 1.153× | 1.3% / 5.1% | 1.5% / 1.8% |
| double 2048×2048 | 200970.5788 | 184114.9500 | 1.092× | 1.2% / 2.0% | 0.4% / 1.2% |
Review: 0c7bdb309 → eba485c30 (44 cases) — retained earlier valid run; attempted refresh excluded
| Case | Before (µs) | After (µs) | Speedup | CV before / after | Run spread before / after |
|---|---|---|---|---|---|
| float 8×8 | 0.0263 | 0.0265 | 0.990× | 3.1% / 2.6% | 0.8% / 2.1% |
| float 16×16 | 0.0723 | 0.0727 | 0.994× | 2.8% / 2.5% | 1.1% / 3.1% |
| float 64×64 | 2.6447 | 2.6147 | 1.011× | 2.2% / 2.1% | 1.0% / 2.0% |
| float 128×128 | 20.1068 | 18.9734 | 1.060× | 1.8% / 1.9% | 1.2% / 0.4% |
| float 256×256 | 178.7853 | 179.8428 | 0.994× | 1.1% / 1.6% | 0.2% / 1.0% |
| float 512×512 | 1418.5618 | 1420.6191 | 0.999× | 0.9% / 1.7% | 0.6% / 1.4% |
| float 1024×1024 | 11275.5762 | 11261.5353 | 1.001× | 1.2% / 0.9% | 0.3% / 1.2% |
| float 64×8 | 0.4483 | 0.4442 | 1.009× | 3.0% / 2.8% | 4.9% / 1.9% |
| float 64×17 | 1.1414 | 1.1136 | 1.025× | 6.3% / 3.9% | 0.5% / 0.3% |
| float 128×8 | 1.7242 | 1.7860 | 0.965× | 3.3% / 2.9% | 0.8% / 4.1% |
| float 128×17 | 3.6715 | 3.5161 | 1.044× | 2.2% / 1.9% | 0.4% / 2.5% |
| double 8×8 | 0.0332 | 0.0333 | 0.997× | 3.3% / 3.8% | 4.5% / 0.8% |
| double 16×16 | 0.1231 | 0.1215 | 1.013× | 2.6% / 3.0% | 5.5% / 2.3% |
| double 64×64 | 5.3873 | 5.2220 | 1.032× | 2.2% / 2.1% | 0.6% / 3.0% |
| double 128×128 | 39.6703 | 38.3379 | 1.035× | 1.6% / 2.6% | 0.0% / 2.4% |
| double 256×256 | 368.2548 | 368.2284 | 1.000× | 1.7% / 0.9% | 0.6% / 0.5% |
| double 512×512 | 2856.9706 | 2855.3293 | 1.001× | 1.7% / 2.1% | 0.5% / 0.7% |
| double 1024×1024 | 22786.2134 | 22784.3764 | 1.000× | 1.7% / 1.2% | 1.0% / 0.2% |
| double 64×8 | 0.6452 | 0.6234 | 1.035× | 2.1% / 2.0% | 2.7% / 2.3% |
| double 64×17 | 1.6986 | 1.7008 | 0.999× | 3.2% / 2.7% | 0.5% / 1.7% |
| double 128×8 | 2.4676 | 2.3751 | 1.039× | 2.6% / 2.7% | 1.7% / 0.9% |
| double 128×17 | 6.1925 | 6.0064 | 1.031× | 2.5% / 2.7% | 2.2% / 2.2% |
| float 8×8 (row-major A) | 0.0269 | 0.0272 | 0.992× | 4.1% / 3.1% | 6.9% / 1.2% |
| float 16×16 (row-major A) | 0.0724 | 0.0736 | 0.984× | 2.6% / 2.6% | 2.8% / 4.0% |
| float 64×64 (row-major A) | 2.8109 | 2.6711 | 1.052× | 1.7% / 2.2% | 2.0% / 1.9% |
| float 128×128 (row-major A) | 20.8434 | 20.4351 | 1.020× | 1.6% / 2.0% | 1.6% / 0.1% |
| float 256×256 (row-major A) | 179.3322 | 180.0101 | 0.996× | 2.1% / 1.4% | 1.0% / 0.0% |
| float 512×512 (row-major A) | 1442.6603 | 1451.8157 | 0.994× | 1.1% / 1.5% | 0.1% / 0.5% |
| float 1024×1024 (row-major A) | 11372.3660 | 11425.4480 | 0.995× | 1.4% / 0.9% | 0.4% / 0.3% |
| float 64×8 (row-major A) | 0.4380 | 0.4394 | 0.997× | 1.9% / 3.0% | 5.2% / 0.3% |
| float 64×17 (row-major A) | 1.3788 | 1.3556 | 1.017× | 1.2% / 2.1% | 0.2% / 0.6% |
| float 128×8 (row-major A) | 1.7504 | 1.7250 | 1.015× | 1.2% / 2.7% | 1.2% / 0.8% |
| float 128×17 (row-major A) | 6.0667 | 6.0013 | 1.011× | 2.4% / 0.9% | 1.1% / 0.2% |
| double 8×8 (row-major A) | 0.0340 | 0.0347 | 0.979× | 2.9% / 2.9% | 1.1% / 5.8% |
| double 16×16 (row-major A) | 0.1245 | 0.1291 | 0.965× | 2.5% / 4.2% | 4.3% / 1.5% |
| double 64×64 (row-major A) | 5.4711 | 5.5156 | 0.992× | 2.6% / 2.1% | 1.8% / 1.1% |
| double 128×128 (row-major A) | 40.4738 | 40.9422 | 0.989× | 1.5% / 1.4% | 1.2% / 1.2% |
| double 256×256 (row-major A) | 376.2471 | 377.5466 | 0.997× | 2.4% / 2.4% | 1.9% / 1.3% |
| double 512×512 (row-major A) | 2894.8179 | 2893.7299 | 1.000× | 2.2% / 1.6% | 2.8% / 1.7% |
| double 1024×1024 (row-major A) | 23036.7122 | 23123.5302 | 0.996× | 1.6% / 2.6% | 0.3% / 0.6% |
| double 64×8 (row-major A) | 0.6588 | 0.6623 | 0.995× | 2.7% / 2.5% | 0.1% / 0.5% |
| double 64×17 (row-major A) | 1.9987 | 2.0610 | 0.970× | 2.1% / 4.2% | 2.2% / 4.8% |
| double 128×8 (row-major A) | 2.4868 | 2.5158 | 0.988× | 2.4% / 2.6% | 0.8% / 0.3% |
| double 128×17 (row-major A) | 8.3608 | 8.4229 | 0.993× | 1.0% / 1.5% | 0.3% / 0.4% |
Appendix E: Provenance and validation details
Measurement conditions: 2026-09-17, 13th Gen Intel(R) Core(TM) i7-13700HX under WSL2; GCC 13.3.0, Clang 18.1.3, Google Benchmark 1.9.5. Common flags: -std=c++14 -O3 -DNDEBUG -march=x86-64 -mtune=generic. SSE2 adds -msse2 -mno-avx -mno-fma; AVX2+FMA adds -mavx2 -mfma. Compiler/Eigen macros were checked. Each pair uses identical sources, flags and filters, without external BLAS or OpenMP.
One benchmark ran at a time on CPU 2, in A/B/A/B order, with eight repetitions of ≥0.08 CPU seconds per case per invocation. Compilation and correctness runs finished before timing. System-wide process monitoring rejected sustained competing work; Windows scheduling, clocks and thermal policy were uncontrolled. All included cases passed the benchmark’s residual check. An interrupted Clang/SSE2 overall attempt was discarded and restarted before the request to omit additional repeats. The later interrupted Clang/AVX2 review refresh was discarded without rerunning it.
The reported ratio uses CPU time. Elapsed-time ratios differed by up to 4.3% in the overall grid and 7.0% in the review grid; small differences near this variability are inconclusive. No additional targeted repeats were performed; noisy cases remain inconclusive.
MR head at publication: 0154510695ec7051f030808e520fd3c750c88015 (rebased during the measurements). Revisions used for benchmarking: original 15f22717866165e381d36d52209400daaf39a386, previous MR 0c7bdb309a94d4a297abacf48900decbdb3a2f01, candidate eba485c3099b2a9712c76d0201a684d578f5dbcf. The timed expression is x = b; a.triangularView<Lower>().solveInPlace(x); setup and residual checks are excluded. RHS storage is column-major. Speedup = before/after, each time being the median of two invocation medians. CV is the maximum within-invocation coefficient of variation; run spread is max(invocation medians) / min(invocation medians) - 1.
These are local before/after results motivated by the Zen 5 benchmark; they are not new Zen 5 or vendor-relative measurements. The Zen 5 host was busy and was not timed in this round.
The retained Clang/AVX2 review comparison used -std=c++14 -O3 -DNDEBUG -mavx2 -mfma with the same benchmark source, machine, affinity, A/B/A/B order and repetition settings. Its rebuilt before and after executables, with the explicit generic x86-64 flags above, have identical SHA-256 hashes to that earlier pair. All four overall comparisons and the other three review comparisons are fresh.
Fresh SSE2 validation also passed trsm_packet and trsm_packet_cancellation with GCC 13.3 and Clang 18.1, both scalar parts, r2 s42. The AVX2 correctness results below apply to the same unchanged candidate source.
The new regression explicitly exercises the shared implementation. Its controlled cache settings keep cancelling terms within a diagonal panel: the harness’s synthetic one-row AVX-512 double panels instead split cancellation across separate GEMM updates, which also fails before this MR. Existing small-block tests remain enabled.
Validation passed with both GCC and Clang: six CMake targets per compiler (trsm_packet, cancellation, and fast-math cancellation, each split by scalar), plus C++14 -O0, triangular and Cholesky integration checks including complex scalars. Additional checks cover SSE2, scalar-only, an eight-register budget, alternate layout/index settings, and ASan/UBSan. Cancellation cases cover upper/lower, unit/nonunit, left/right, packet tails, mixed lanes, poisoned unused coefficients, and orders 4–256. The regression fails at the previous MR revision. Shared-kernel AVX-512 passed natively on AMD RYZEN AI MAX+ 395 with GCC and Clang; AArch64 NEON passed under QEMU with GCC 14. No timings were taken on that loaded remote machine. Other ISAs and MSVC were not validated in this round.
Disassembly of the current AVX2 binaries confirms integer exponent checks in the final-row guard (scalar integer checks with GCC, packed integer masks/comparisons with Clang). The guard and retry stay outside the substitution loops. This adds work for row-major triangles, most visible in tiny solves; measured costs are included above.