-fuaddsub-overflow-match-all, -fif-conversion-gimple
Description
Arm-related instruction combination and optimization: Low-bit multiplication algorithms are identified and output the result using efficient high-bit multiplication instructions.
Usage
Use the -fuaddsub-overflow-match-all and -fif-conversion-gimple options to enable optimization.
Note:
- This optimization requires the -O2 or higher optimization level and must be used together with the -ftree-fold-phiopt option.
- This optimization is similar to -fmerge-mull, but it can be used in more comprehensive scenarios, such as 8-digit by 8-digit multiplication and 16-digit by 16-digit multiplication.
Result
The test case is as follows:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 | #include <stdint.h> typedef unsigned __int128 uint128_t; uint128_t mul128_perm (uint64_t a, uint64_t b) { uint64_t a_lo = a & 0xFFFFFFFF; uint64_t b_lo = b & 0xFFFFFFFF; uint64_t a_hi = a >> 32; uint64_t b_hi = b >> 32; uint64_t lolo = a_lo * b_lo; uint64_t lohi = a_lo * b_hi; uint64_t hilo = a_hi * b_lo; uint64_t hihi = a_hi * b_hi; uint64_t middle = hilo + lohi; uint64_t middle_hi = middle >> 32; uint64_t middle_lo = middle << 32; uint64_t res_lo = lolo + middle_lo; uint64_t res_hi = hihi + middle_hi; res_hi = res_lo < middle_lo ? res_hi + 1 : res_hi; res_hi = middle < hilo ? res_hi + 0x100000000 : res_hi; uint128_t res = ((uint128_t) res_hi) << 64; res += res_lo; return res; } |
Test command:
1 | gcc -O2 -ftree-fold-phiopt -fif-conversion-gimple -fuaddsub-overflow-match-all -S test.c -o test.s |
Figure 1 Option disabled


Figure 2 Option enabled


After this option is enabled, the generated assembly code uses the umulh and mul instructions to optimize the 64-bit multiplication calculation.
Parent topic: Static Compilation Optimization