我要评分
获取效率
正确性
完整性
易理解

flamegraph

Command Function

You can extract hotspot functions from a flame graph SVG file or a Performance Boundary Analyzer result JSON file and analyze them to quickly identify functions or code that have optimization opportunities and receive rewriting suggestions.

Syntax

./tiancheng flamegraph [options]

Parameter Description

Table 1 Parameter description

Parameter

Option

Description

-h/--help

-

(Optional) Obtains help information.

-o

-

(Mandatory) Output directory for report files.

-I

-

(Optional) Header file directory.

--clang-resource-dir

-

(Optional) Directory containing the Clang built-in header file (for example, stddef.h). By default, the header file directory in the lib directory of the tool is used.

NOTE:

The default value of this parameter is install_path/tiancheng/lib/clang/17.

--gcc-toolchain

-

(Optional) GCC Compiler installation directory, which is used as the toolchain path for code analysis. By default, the GCC Compiler directory of the system is used.

-r

-

(Optional) Root directory of the project to which the source file to be analyzed belongs. This directory is used to locate function symbols and search for header files.

--simd-target

neon/sve

(Optional) Arm instruction set to convert to. The value is NEON or SVE and defaults to neon.

  • NEON: SIMD instruction set in the Arm architecture. It uses 128-bit vector registers for parallel data processing and supports 8-, 16-, 32-, and 64-bit integer operations, as well as single-precision floating-point operations. It is supported on servers based on the Kunpeng 920 series and Kunpeng 950 processors.
  • SVE: It is a vector-length-agnostic SIMD instruction set that allows the same code to automatically achieve optimal performance on Arm processors with different vector widths. It is supported on some of servers based on the new Kunpeng 920 models and Kunpeng 950 processors.
NOTE:

When x86 Intel intrinsics analysis is enabled, the --simd-target parameter must be set to sve.

-v

-

(Optional) Enables the detailed output mode to display additional debugging information and the analysis process. By default, this parameter is disabled.

--flamegraph

-

(Optional) Path to the flame graph SVG file or the Performance Boundary Analyzer result JSON file in flame graph scanning mode. It is required when using flame graph scanning mode.

-top

-

(Optional) Top N hotspot functions. The default value is top 10. It is required when using flame graph scanning mode.

--enable-ifstmt-scan

-

(Optional) Enables the scanning and analysis of if statements.

--enable-intel-intrinsics

-

(Optional) Enables x86 Intel intrinsics analysis to analyze x86 vectorized instructions and rewrite x86 vectorized code. By default, this parameter is disabled.

--features

fma/dq/bw/vl/bf16/vnni/fp16

(Optional) x86 instruction set features. The default value is empty. Multiple options are supported, separated by commas (,). This parameter is only effective when x86 Intel intrinsics analysis is enabled.

  • FMA: Fused multiply-add operation, which performs multiplication and addition in a single instruction. It does not depend on the AVX512 instruction set and can be used in combination with AVX, AVX2, and AVX512 instruction sets.
  • DQ: Doubleword/Quadword support, an AVX512 instruction set feature that supports 32-bit and 64-bit integer operations and mask operations.
  • BW: Byte and word extension, an AVX512 instruction set feature that supports 8-bit and 16-bit integer operations.
  • VL: Vector length extension, an AVX512 instruction set feature that allows AVX512 instructions to operate on 128-bit and 256-bit vector registers.
  • BF16: 16-bit brain floating point format, an AVX512 instruction set feature that supports 16-bit brain floating point operations.
  • VNNI: Vector neural network instructions, an AVX512 instruction set feature that supports low-precision integer vector dot product and multiply-accumulate operations.
  • FP16: Half-precision floating point, an AVX512 instruction set feature that supports vector arithmetic operations on half-precision floating-point numbers.

--isa

sse2/avx/avx2/avx512

(Optional) x86 instruction set to be identified. The default value is avx2. This parameter is only effective when x86 Intel intrinsics analysis is enabled.

  • SSE2: SIMD instruction set in the x86 architecture. It uses 128-bit vector registers for parallel data processing and supports 8-, 16-, 32-, and 64-bit integer operations, as well as double-precision floating-point operations.
  • AVX: SIMD instruction set in the x86 architecture. It enables parallel data processing using 256-bit vector registers, supporting 256-bit single-precision and double-precision floating-point operations.
  • AVX2: SIMD instruction set in the x86 architecture. It enables parallel data processing using 256-bit vector registers, supporting 256-bit integer and floating-point operations.
  • AVX512: SIMD instruction set in the x86 architecture. It enables parallel data processing using 512-bit vector registers, supporting 512-bit integer operations and floating-point operations.
NOTE:

If the identified instruction set is AVX512, the optimized source code must run on a device that supports SVE with a 512-bit vector length.

Example

View information about the flamegraph command.
./tiancheng flamegraph -h

Command output:

OVERVIEW: Use an imported flamegraph/hotspot result as the analysis scope

USAGE:
  tiancheng flamegraph -r <project-root> -o <path> --flamegraph <file> [--top N] [options]

OPTIONS:
  -h/--help                                   Display available options
  -o <pathname>                               Specify output report file path
  -I <pathname>                               Include the extra header files
  --clang-resource-dir=<path>                 Clang resource directory (path up to lib/clang/<version>, excluding 'include')
  --gcc-toolchain=<path>                      Path to the GCC toolchain root directory used by Clang
  -r <pathname>                               Inputfile root path
  --simd-target=<value>                       Convert simd type selection
    =neon                                     Convert to ARM NEON intrinsics
    =sve                                      Convert to ARM SVE intrinsics
  -v                                          Enable verbose output
  --flamegraph=<file>                         Path to the flamegraph SVG file
  --top=<N>                                   Analyze top n hotspot functions
  --enable-ifstmt-scan                        Enable scanning of if statements
  --enable-intel-intrinsics                   Enable analysis of x86 Intel intrinsics
  --features=<fma,dq,bw,vl,bf16,vnni,fp16>
                                              x86 feature flags, comma-separated
                                              Layered on top of --isa. Only effective with --enable-intel-intrinsics
    fma                                       Fused Multiply-Add: enables _mm256_fmadd_ps etc.  -> -mfma
                                              (works with avx/avx2/avx512, not AVX-512-specific)
    dq                                        Doubleword/Quadword: enables 64-bit mask ops,     -> -mavx512dq
                                              _mm512_reduce_add_ps, _mm512_mullo_epi64 etc.
    bw                                        Byte/Word: enables byte/word mask and compare ops -> -mavx512bw
    vl                                        Vector Length: enables 128/256-bit AVX-512 ops   -> -mavx512vl
    bf16                                      BFloat16: enables _mm512_dpbf16_ps etc.           -> -mavx512bf16
                                              (auto-implies dq, bw, vl)
    vnni                                      Vector Neural Net Instructions: int8 dot-product  -> -mavx512vnni
    fp16                                      FP16 arithmetic: enables _mm512_add_ph etc.       -> -mavx512fp16
  --isa=<level>                               x86 ISA level for Intel intrinsics analysis
                                              Only effective with --enable-intel-intrinsics
    =sse2                                     SSE2:    128-bit SIMD, baseline x86-64
    =avx                                      AVX:     256-bit SIMD, float only
    =avx2                                     AVX2:    AVX + 256-bit integer ops (default)
    =avx512                                   AVX-512: AVX2 + 512-bit SIMD foundation

EXAMPLES:
  Note: [option] denotes an optional argument.

  Flamegraph hotspot scan
    tiancheng flamegraph -o ./output -r /path/to/project --flamegraph hotspot.svg [--top 10]