-ftree-slp-transpose-vectorize
In the loop splitting phase, this option inserts temporary arrays to enhance the data flow analysis capability for loops with continuous memory access reads. In the vectorization (SLP) phase, the SLP analysis for transposing grouped_stores is added.
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 | int foo (unsigned char *oxa, int ia, unsigned char *oxb, int ib) { unsigned tmp[4][4]; unsigned a0, a1, a2, a3; int sum = 0; for (int i = 0; i < 4; i++, oxa += ia, oxb += ib) { a0 = (oxa[0] - oxb[0]) + ((oxa[4] - oxb[4]) << 16); a1 = (oxa[1] - oxb[1]) + ((oxa[5] - oxb[5]) << 16); a2 = (oxa[2] - oxb[2]) + ((oxa[6] - oxb[6]) << 16); a3 = (oxa[3] - oxb[3]) + ((oxa[7] - oxb[7]) << 16); int t0 = a0 + a1; int t1 = a0 - a1; int t2 = a2 + a3; int t3 = a2 - a3; tmp[i][0] = t0 + t2; tmp[i][2] = t0 - t2; tmp[i][1] = t1 + t3; tmp[i][3] = t1 - t3; } for (int i = 0; i < 4; i++) { int t0 = tmp[0][i] + tmp[1][i]; int t1 = tmp[0][i] - tmp[1][i]; int t2 = tmp[2][i] + tmp[3][i]; int t3 = tmp[2][i] - tmp[3][i]; a0 = t0 + t2; a2 = t0 - t2; a1 = t1 + t3; a3 = t1 - t3; sum += a0 + a1 + a2 + a3; } return sum; } |
In the preceding test case, the first for loop can be split into the following formats:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 | for (int i = 0; i < 4; i++, oxa += ia, oxb += ib) { a00[i] = (oxa[0] - oxb[0]) + ((oxa[4] - oxb[4]) << 16); a11[i] = (oxa[1] - oxb[1]) + ((oxa[5] - oxb[5]) << 16); a22[i] = (oxa[2] - oxb[2]) + ((oxa[6] - oxb[6]) << 16); a33[i] = (oxa[3] - oxb[3]) + ((oxa[7] - oxb[7]) << 16); } for (int i = 0; i < 4; i++) { int t0 = a00[i] + a11[i]; int t1 = a00[i] - a11[i]; int t2 = a22[i] + a33[i]; int t3 = a22[i] - a33[i]; tmp[i][0] = t0 + t2; tmp[i][2] = t0 - t2; tmp[i][1] = t1 + t3; tmp[i][3] = t1 - t3; } |
In the first loop obtained through splitting, the calculation on the right of the equal sign (=) is isomorphic, and the loads are continuous, allowing for vectorization. However, the memory addresses of a00[i], a11[i], a22[i] and a33[i] on the left are discontinuous and cannot serve as the root node of the vectored SLP tree. The vectorization is not available in this scenario. When a00[i], a11[i], a22[i] and a33[i] are written to the memory, the expected register content is as follows.
Register |
Value |
|---|---|
vec0 |
a00[0] a00[1] a00[2] a00[3] |
vec1 |
a11[0] a11[1] a11[2] a11[3] |
vec2 |
a22[0] a22[1] a22[2] a22[3] |
vec3 |
a33[0] a33[1] a33[2] a33[3] |
In each iteration, the content of the register can be calculated as follows:
Register |
Value |
|---|---|
vec0 |
a00[0] a11[0] a22[0] a33[0] |
vec1 |
a00[1] a11[1] a22[1] a33[1] |
vec2 |
a00[2] a11[2] a22[2] a33[2] |
vec3 |
a00[3] a11[3] a22[3] a33[3] |
The grouped_stores is transposed to obtain the expected root node of the SLP tree. Then, the SLP capability is used to perform subsequent vectorization analysis.
In addition, in the second loop obtained through splitting and the last loop in the test case, the tmp two-dimensional array is read immediately after being written to the memory. In this scenario, the memory access behavior is optimized to the permutation behavior between registers. This memory access optimization is enabled by default.
Usage
Add the following content to the option:
1 | -O3 -ftree-slp-transpose-vectorize
|
Note: This option takes effect only after the -O3 optimization level is enabled.