Sample Usage Guide
【免费下载链接】geGE(Graph Engine)是面向昇腾的图编译器和执行器,提供了计算图优化、多流并行、内存复用和模型下沉等技术手段,加速模型执行效率,减少模型内存占用。 GE 提供对 PyTorch、TensorFlow 前端的友好接入能力,并同时支持 onnx、pb 等主流模型格式的解析与编译。项目地址: https://gitcode.com/cann/ge
1. Function Description
This sample demonstrates graph construction using operator overload, aimed at helping graph developers quickly understand operator overload definition. Unlikeoperator_overload, this directory usesSession.run_graph_with_stream_asyncinterface during execution phase. For more information about asynchronous execution, please refer to Asynchronous Execution.
2. Directory Structure
python/ ├── src/ | ├── common.py // Common logic: constants, graph construction, input tensors, GE lifecycle | ├── make_add_graph.py // Asynchronous execution sample | └── make_add_graph_custom_allocator.py // Register custom SamplePoolAllocator asynchronous execution sample ├── README.md // README file ├── run_sample.sh // Execution script3. Usage Instructions
3.1. Prepare CANN Package
- This sample requires installing two sets of CANN simultaneously: latest development package for graph compilation, officially released package from website provides
pyACLmodule for graph execution. "Compilation" and "execution" in this text specifically refer to graph compilation and graph execution, not GE engineering source code compilation. - Installation please refer to:
- Latest development package, for graph compilation, provides latest GE/Python capabilities required by this sample. Installation please refer to Environment Preparation "Method 3: Manual installation of software packages > Scenario 1: Experience master version capabilities or develop based on master version", install latest version of
toolkitandopspackages - Officially released CANN
toolkitandopspackages, for graph execution, providepyACL. Installation please refer to Environment Preparation "Method 3: Manual installation of software packages > Scenario 2: Experience released version capabilities or develop based on released version", install officially released software packages
- Latest development package, for graph compilation, provides latest GE/Python capabilities required by this sample. Installation please refer to Environment Preparation "Method 3: Manual installation of software packages > Scenario 1: Experience master version capabilities or develop based on master version", install latest version of
- Set environment variables (assuming latest development package is installed at /usr/local/Ascend/, officially released package is installed at /usr/local/Ascend-release/)
source /usr/local/Ascend/cann/set_env.sh export PYTHONPATH="$PYTHONPATH:/usr/local/Ascend-release/cann/python/site-packages"3.2. Execute
This directory provides two optional targets:
3.2.1. Default allocator sample
bash run_sample.sh -t sample_and_run_pythonThis command will:
- Generate dump graph and use GE built-in allocator to asynchronously execute the graph
After successful execution, you will see:
[Success] sample execution successful, pbtxt dump generated in current directory. The file starts with ge_onnx_ and can be opened in netron for display3.2.2. Custom allocator sample
bash run_sample.sh -t sample_and_run_python_custom_allocatorThis command on basis of default sample, throughSession.register_external_allocatorregisters customSamplePoolAllocatorto GE. This class inheritsge.allocator.Allocatorbase class: maintains free block list byexact byte count, hits then reuse, otherwiseacl.rt.mallocallocate;freetime puts block back to pool (not immediatelyacl.rt.free). When object is recycled, in__del__uniformlyacl.rt.freecached blocks.
After successful execution, terminal will show output in following format (specific line count depends on GE allocation/release for output Tensor; address andsizebased on actual execution):
[Info] SamplePoolAllocator registered to stream [SamplePool] new : addr=0x<device memory address> size=<byte count> B [Info] Asynchronous Graph execution successful! [SamplePool] cache : addr=0x<device memory address> size=<byte count> B [Info] SamplePoolAllocator unregistered [SamplePool] destroy: freed <N> cached blocks [Info] Running environment cleaned [Success] sample execution successful, pbtxt dump generated in current directory. The file starts with ge_onnx_ and can be opened in netron for displayOutput File Description
After successful execution, the following files will be generated in current directory:
ge_onnx_*.pbtxt- protobuf text format of graph structure, can be viewed with netron
3.3. Log Printing
If you need log printing to assist debugging during executable program execution, you can set the following environment variables beforebash run_sample.sh -t sample_and_run_pythonto print logs to screen:
export ASCEND_SLOG_PRINT_TO_STDOUT=1 # Print logs to screen export ASCEND_GLOBAL_LOG_LEVEL=0 # Log level set to debug level3.4. DUMP Graph During Graph Compilation Process
If you need to DUMP graph to assist debugging graph compilation process during executable program execution, you can set the following environment variables beforebash run_sample.sh -t sample_and_run_pythonto DUMP graph to execution path:
export DUMP_GE_GRAPH=24. Core Concepts Introduction
4.1. Graph Construction Steps
- Create graph builder (provides context, workspace and construction-related methods needed for graph construction)
- Add starting nodes (starting nodes refer to nodes without input dependencies, usually including graph inputs (like Data nodes) and weight constants (like Const nodes))
- Add intermediate nodes (intermediate nodes are computation nodes with input dependencies, usually generated by user graph construction logic, and connected using existing nodes as inputs)
- Set graph output (explicitly specify graph output nodes as computation result endpoints)
4.2. Operator Overload
Concept Explanation:Operator overload is a syntax sugar provided by ES API, with syntax encapsulation for AI operators, making graph construction code more concise and intuitive.
Graph Construction API Features:
- Operator overload maintains the same type checking and constraints as function calls
【免费下载链接】geGE(Graph Engine)是面向昇腾的图编译器和执行器,提供了计算图优化、多流并行、内存复用和模型下沉等技术手段,加速模型执行效率,减少模型内存占用。 GE 提供对 PyTorch、TensorFlow 前端的友好接入能力,并同时支持 onnx、pb 等主流模型格式的解析与编译。项目地址: https://gitcode.com/cann/ge
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考