Skip to content

Improve dynamic broadcast and transpose assignment performance - #2934

Open
wolfv wants to merge 8 commits into
xtensor-stack:masterfrom
wolfv:perf-upstream-master
Open

wolfv wants to merge 8 commits into
xtensor-stack:masterfrom
wolfv:perf-upstream-master

Conversation

@wolfv

@wolfv wolfv commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Summary

This PR improves assignment performance for dynamic-rank broadcast and permutation expressions, focusing on the stepper-heavy workloads reported in #2784, #2683, and #2859.

It adds two conservative execution paths:

  1. Runtime pointer/stride planning for supported 3-D broadcast expressions.
  2. Source-ordered permutation kernels for dense 3-D transpose/cast/scaling expressions.

Unsupported types, dimensions, layouts, views, and stride patterns retain the existing assignment path. The PR also adds comparison benchmarks for raw loops, Eigen, Armadillo (when available), dynamic shapes, fixed shapes, broadcasting, and transpose/cast expressions.

Runtime broadcast plan

For non-trivial three-dimensional row-major broadcasts with directly addressable arithmetic leaves:

  • Flatten the expression tree once into typed leaf/function plans.
  • Cache leaf pointers and aligned per-axis strides.
  • Represent broadcast axes as zero strides.
  • Distinguish zero and unit innermost strides so the compiler can vectorize the generated loop.
  • Retain the original assignment path for unsupported expressions and layouts.

This avoids rebuilding/reseeking recursive steppers for every short inner row.

Permutation kernels

For dense 3-D cast(transpose(...)) / scalar assignments:

  • Determine source-contiguous loop order once from runtime strides.
  • Use a 32x32 tiled fallback for general dense permutations.
  • Use a source-ordered RGB deinterleave path for interleaved 3-channel input and planar output.
  • Preserve exact division semantics rather than replacing division with reciprocal multiplication.
  • Fall back to the existing assignment machinery when the expression or stride pattern is not eligible.

Benchmarks

benchmark_compare.cpp covers dynamic contiguous fused arithmetic, dynamic N-D broadcasting, fixed-shape fused arithmetic, transpose/cast/scaling, raw loops, Eigen Array/Tensor, and optional Armadillo comparisons.

Measured with -O3 -march=native, xsimd 14, AVX-512, SIMD width 8:

Workload Original xtensor This PR Raw loop Eigen Improvement
Dynamic broadcast, {64,64,16} ~163 us 11.9 us 7.3 us 35.8 us 13.6x
Transpose/cast, {128,256,3} -> {3,128,256} ~165 us 7.0 us 7.4 us 42.6 us 23.6x
Dynamic contiguous fused arithmetic ~22-24 us 22.6 us 18.5 us 16.3 us roughly unchanged
Fixed contiguous fused arithmetic, 256 elements — 52.8 ns — 52.9 ns parity with Eigen

Validation

  • Full non-SIMD suite: 78/78 tests passed.
  • Full SIMD suite: 78/78 tests passed (20,213 assertions in aggregate xtest).
  • Pre-commit hooks pass.
  • Added coverage for dense transpose/cast assignment, RGB deinterleave, general tiled permutation, runtime broadcast planning, and unsupported-expression fallback.

@codspeed

codspeed Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will regress 1 benchmark

⚠️ 1 benchmark measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. This result is not comparable, so it counts as unchanged.

Preventing compiler optimizations

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
❌ 1 regressed benchmark
✅ 253 untouched benchmarks
🆕 7 new benchmarks

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Benchmark BASE HEAD Efficiency
❌ xshape_initializer[std\:\:array<std\:\:size_t, 4>] 151.7 ns 180.8 ns -16.13%
⚡ xshape_access[std\:\:array<std\:\:size_t, 4>] 154.4 ns 125.3 ns +23.28%
🆕 broadcast_dynamic_raw N/A 594.7 µs N/A
🆕 broadcast_dynamic_xtensor N/A 1.2 ms N/A
🆕 linear_dynamic_raw N/A 1.1 ms N/A
🆕 linear_dynamic_xtensor N/A 1.4 ms N/A
🆕 linear_fixed_xtensor N/A 4.8 µs N/A
🆕 transpose_cast_raw N/A 400.7 µs N/A
🆕 transpose_cast_xtensor N/A 456.8 µs N/A
⚠️ create_xview < 1 ns < 1 ns N/A

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing wolfv:perf-upstream-master (68f5541) with master (d9a57b6)

Open in CodSpeed

@wolfv

wolfv commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

CodSpeed’s remaining failure is unrelated measurement noise: this PR does not touch xshape, the report warns that the compared runtime environments differ, and the paired xshape_access result moves by a similar amount in the opposite direction. A fresh benchmark workflow rerun also passed. All correctness, sanitizer, platform, xsimd, TBB, and benchmark jobs are green.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant