BlockSparseMatrix: add block-sparse data structure and operations

Adds BlockSparseMatrix<Scalar, Options, BR, BC>, a sparse matrix whose nonzero entries are fixed-size dense blocks. It supports triangular and selfadjoint views with the same API as SparseMatrix, and plugs into the Eigen product evaluator so expressions like y.noalias() = A * x and A.triangularView<Upper>().solveInPlace(x) work without extra glue.


Benchmarks (nB=200 blocks, sparsity=5%)

Speedup is Sm time / BSM time; values >1 mean BSM is faster. For TriMV_DiagT the baseline is Sm_TriMV; for SymmMV_DiagNSA/SymmMV_DiagSA it is Sm_SymmMV.

Sparsity level (1%, 5%, 10%) has little effect on the relative speedups: both Sm and BSM traverse the same block sparsity pattern, so the gains come from within-block compute density rather than from how many blocks there are. Results at 5% are representative.

B=2

Operation float double cf cd
SpMV 1.46x 1.87x 2.50x 2.21x
TriMV 1.32x 1.51x 1.49x 1.50x
TriMV_DiagT 1.86x 2.36x 2.74x 2.05x
SymmMV_DiagNSA 0.87x 0.85x 1.54x 1.36x
SymmMV_DiagSA 1.03x 1.36x 1.68x 1.70x
TriSolve 2.60x 3.08x 2.01x 2.71x
SpSp_Add 10.38x 11.99x 9.90x 10.33x
SpSp_Mul 1.59x 1.84x 1.79x 1.99x

B=3

Operation float double cf cd
SpMV 1.87x 2.36x 1.65x 1.48x
TriMV 1.97x 2.01x 1.39x 1.35x
TriMV_DiagT 2.64x 2.63x 1.83x 1.51x
SymmMV_DiagNSA 0.97x 1.04x 1.30x 1.40x
SymmMV_DiagSA 1.30x 1.53x 1.42x 1.38x
TriSolve 2.95x 3.22x 1.79x 2.34x
SpSp_Add 21.18x 14.45x 14.31x 10.42x
SpSp_Mul 3.28x 3.85x 3.51x 3.14x

B=4

Operation float double cf cd
SpMV 4.21x 3.69x 5.54x 2.60x
TriMV 4.31x 4.09x 3.45x 2.54x
TriMV_DiagT 6.03x 5.42x 6.32x 3.17x
SymmMV_DiagNSA 2.19x 2.13x 3.15x 2.25x
SymmMV_DiagSA 2.22x 2.17x 3.80x 2.51x
TriSolve 8.14x 7.57x 2.88x 3.03x
SpSp_Add 35.61x 18.75x 19.28x 9.92x
SpSp_Mul 7.06x 7.48x 7.55x 6.52x
Edited by Charles Schlosser

Merge request reports

Loading
Loading