Optimize dot product for contiguous data with runtime-unit stride

Description

Optimize dot product for contiguous data with runtime-unit stride

Introduce generic_dot_impl_helper to optimize dot product evaluations when the expressions have compile-time inner stride != 1 but runtime inner stride == 1. This commonly occurs in inner products resulting from matrix-vector multiplications (GEMV) where the matrix has a single row or column.

Also add SFINAE guards to TransposeImpl::data() to prevent compile errors when the nested expression does not support direct data access.

Performance improvements on related benchmarks:

  • Gemv_float/1/1024: 2.24 -> 23.73 GFLOPS (~10.6x speedup, +957%)
  • Gemv_float/1/256: 2.59 -> 20.95 GFLOPS (~8.1x speedup, +708%)
  • Gemv_double/1/1024: 2.21 -> 12.52 GFLOPS (~5.7x speedup, +465%)
  • Gemv_double/1/256: 2.47 -> 11.77 GFLOPS (~4.8x speedup, +376%)
  • Gemv_cfloat/1/1024: 5.10 -> 18.05 GFLOPS (~3.5x speedup, +254%)
  • Gemv_cfloat/1/256: 5.04 -> 15.46 GFLOPS (~3.1x speedup, +206%)

Reference issue

Additional information

Merge request reports

Loading
Loading