Loading
Optimize dot product for contiguous data with runtime-unit stride
Description
Optimize dot product for contiguous data with runtime-unit stride
Introduce generic_dot_impl_helper to optimize dot product evaluations when the expressions have compile-time inner stride != 1 but runtime inner stride == 1. This commonly occurs in inner products resulting from matrix-vector multiplications (GEMV) where the matrix has a single row or column.
Also add SFINAE guards to TransposeImpl::data() to prevent compile errors when the nested expression does not support direct data access.
Performance improvements on related benchmarks:
- Gemv_float/1/1024: 2.24 -> 23.73 GFLOPS (~10.6x speedup, +957%)
- Gemv_float/1/256: 2.59 -> 20.95 GFLOPS (~8.1x speedup, +708%)
- Gemv_double/1/1024: 2.21 -> 12.52 GFLOPS (~5.7x speedup, +465%)
- Gemv_double/1/256: 2.47 -> 11.77 GFLOPS (~4.8x speedup, +376%)
- Gemv_cfloat/1/1024: 5.10 -> 18.05 GFLOPS (~3.5x speedup, +254%)
- Gemv_cfloat/1/256: 5.04 -> 15.46 GFLOPS (~3.1x speedup, +206%)