Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Performance Concepts and Metrics
- Latency, throughput, power consumption, and resource utilization
- Differentiating system-level vs. model-level bottlenecks
- Profiling techniques for inference versus training
Profiling on Huawei Ascend
- Utilizing CANN Profiler and MindInsight
- Diagnostics for kernels and operators
- Understanding offload patterns and memory mapping
Profiling on Biren GPU
- Performance monitoring via the Biren SDK
- Kernel fusion, memory alignment, and execution queues
- Power and temperature-aware profiling
Profiling on Cambricon MLU
- Using BANGPy and Neuware performance tools
- Kernel-level visibility and log interpretation
- Integrating the MLU profiler with deployment frameworks
Graph and Model-Level Optimization
- Strategies for graph pruning and quantization
- Operator fusion and restructuring the computational graph
- Standardizing input sizes and tuning batch parameters
Memory and Kernel Optimization
- Optimizing memory layout and reuse
- Efficient buffer management across different chipsets
- Platform-specific kernel tuning techniques
Cross-Platform Best Practices
- Performance portability: Abstraction strategies
- Creating shared tuning pipelines for multi-chip environments
- Example: Tuning an object detection model across Ascend, Biren, and MLU
Summary and Next Steps
Requirements
- Practical experience with AI model training or deployment pipelines
- Foundational understanding of GPU/MLU compute principles and model optimization
- Basic proficiency with performance profiling tools and key metrics
Target Audience
- Performance engineers
- Machine learning infrastructure teams
- AI system architects
21 Hours