Delivered either online or onsite, our instructor-led live GPU (Graphics Processing Unit) training programs use interactive discussion and practical exercises to demonstrate the core principles of GPU technology and the essential skills needed to program them.
GPU training is offered as "online live training" or "onsite live training". Online live training (also known as "remote live training") is conducted through an interactive remote desktop. Onsite live training takes place locally at customer premises in Varna or at NobleProg corporate training centers in Varna.
The "Central Point" complex offers quick access to main roads leading to the airport, the northern and southern resorts and the Varna - Sofia and Varna - Burgas highways.
This instructor-led course in Varna assists intermediate AI engineers in constructing and optimizing neural network models via the Huawei Ascend platform and CANN toolkit. Learners will set up environments, build applications using MindSpore, and manage deployment to edge or cloud settings.
This live, instructor-led training in Varna delves into Huawei's AI stack, covering the spectrum from the CANN SDK to the MindSpore framework. It is designed to assist beginner and intermediate professionals in understanding how these components synergize on Ascend hardware to streamline lifecycle management and deployment.
This instructor-led, live training in Varna (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use OpenACC to program heterogeneous devices and exploit their parallelism.
By the end of this training, participants will be able to:
Set up an OpenACC development environment.
Write and run a basic OpenACC program.
Annotate code with OpenACC directives and clauses.
This instructor-led training on Varna focuses on deploying and optimizing CV and NLP models using the CANN SDK for Ascend hardware. Learners will master model conversion, integration into live pipelines, and enhancing inference performance for real-time detection and analysis.
This instructor-led live training in Varna (online or onsite) is tailored for beginner to intermediate-level developers who wish to understand the basics of GPU programming and the key frameworks and tools used for developing GPU applications.
By the end of this training, participants will be able to: Understand the difference between CPU and GPU computing and the benefits and challenges of GPU programming.
Choose the right framework and tool for their GPU application.
Create a basic GPU program that performs vector addition using one or more of the frameworks and tools.
Use the respective APIs, languages, and libraries to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use the respective memory spaces, such as global, local, constant, and private, to optimize data transfers and memory accesses.
Use the respective execution models, such as work-items, work-groups, threads, blocks, and grids, to control the parallelism.
Debug and test GPU programs using tools such as CodeXL, CUDA-GDB, CUDA-MEMCHECK, and NVIDIA Nsight.
Optimize GPU programs using techniques such as coalescing, caching, prefetching, and profiling.
This instructor-led live training in Varna empowers advanced developers to build, deploy, and optimize custom AI operators. Learners will master the integration of CANN TIK and Apache TVM, enabling sophisticated optimization and scheduling on Huawei Ascend hardware for superior real-world performance.
This instructor-led, live training in Varna (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use different frameworks for GPU programming and compare their features, performance, and compatibility.
By the end of this training, participants will be able to:
Set up a development environment that includes OpenCL SDK, CUDA Toolkit, ROCm Platform, a device that supports OpenCL, CUDA, or ROCm, and Visual Studio Code.
Create a basic GPU program that performs vector addition using OpenCL, CUDA, and ROCm, and compare the syntax, structure, and execution of each framework.
Use the respective APIs to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use the respective languages to write kernels that execute on the device and manipulate data.
Use the respective built-in functions, variables, and libraries to perform common tasks and operations.
Use the respective memory spaces, such as global, local, constant, and private, to optimize data transfers and memory accesses.
Use the respective execution models to control the threads, blocks, and grids that define the parallelism.
Debug and test GPU programs using tools such as CodeXL, CUDA-GDB, CUDA-MEMCHECK, and NVIDIA Nsight.
Optimize GPU programs using techniques such as coalescing, caching, prefetching, and profiling.
This live training session in Varna provides a comprehensive overview of CloudMatrix for scalable AI inference. Discover how to deploy, optimize, and monitor models using CANN and MindSpore. Practical exercises will guide you through packaging, conversion, serving, and performance tuning for both real-time and batch workloads.
This live, instructor-led training in Varna delves into the core principles and practical fundamentals of deploying AI models on Ascend edge devices via the CANN toolkit, enabling participants to develop essential skills for compiling, optimizing, and managing resource-constrained environments.
This instructor-led, live training in Varna (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to install and use ROCm on Windows to program AMD GPUs and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes ROCm Platform, a AMD GPU, and Visual Studio Code on Windows.
Create a basic ROCm program that performs vector addition on the GPU and retrieves the results from the GPU memory.
Use ROCm API to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use HIP language to write kernels that execute on the GPU and manipulate data.
Use HIP built-in functions, variables, and libraries to perform common tasks and operations.
Use ROCm and HIP memory spaces, such as global, shared, constant, and local, to optimize data transfers and memory accesses.
Use ROCm and HIP execution models to control the threads, blocks, and grids that define the parallelism.
Debug and test ROCm and HIP programs using tools such as ROCm Debugger and ROCm Profiler.
Optimize ROCm and HIP programs using techniques such as coalescing, caching, prefetching, and profiling.
This instructor-led, live training in Varna (online or onsite) is aimed at beginner to intermediate developers who wish to use ROCm and HIP to program AMD GPUs and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes ROCm Platform, an AMD GPU, and Visual Studio Code.
Create a basic ROCm program that performs vector addition on the GPU and retrieves the results from the GPU memory.
Use ROCm API to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use HIP language to write kernels that execute on the GPU and manipulate data.
Use HIP built-in functions, variables, and libraries to perform common tasks and operations.
Use ROCm and HIP memory spaces, such as global, shared, constant, and local, to optimize data transfers and memory accesses.
Use ROCm and HIP execution models to control the threads, blocks, and grids that define the parallelism.
Debug and test ROCm and HIP programs using tools such as ROCm Debugger and ROCm Profiler.
Optimize ROCm and HIP programs using techniques such as coalescing, caching, prefetching, and profiling.
This live training in Varna introduces the CANN toolkit to AI framework developers. Learn to set up environments, convert models, and deploy applications on Ascend hardware using MindSpore, TensorFlow, or PyTorch, covering the full workflow from training to inference.
Enhance AI workload performance on Ascend, Biren, and Cambricon through this practical training in Varna. Master model benchmarking, pinpoint bottlenecks, and deploy graph, kernel, and operator-level optimizations. Refine deployment pipelines to optimize throughput and latency across these top-tier platforms.
Maximize neural network inference efficiency on Ascend AI processors through this advanced, instructor-led training in Varna. Delve into CANN's runtime architecture, utilizing the Graph Engine, TIK, and TVM for profiling, custom operator creation, and resolving memory bottlenecks.
Transition CUDA applications to Chinese GPU ecosystems, such as Huawei Ascend and Biren, within Varna. This instructor-led course assists advanced programmers in code translation and performance refinement, featuring practical labs for porting CUDA codebases to emerging SDKs.
This instructor-led, live training in Varna (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use CUDA to program NVIDIA GPUs and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes CUDA Toolkit, an NVIDIA GPU, and Visual Studio Code.
Create a basic CUDA program that performs vector addition on the GPU and retrieves the results from the GPU memory.
Use the CUDA API to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use the CUDA C/C++ language to write kernels that execute on the GPU and manipulate data.
Use CUDA built-in functions, variables, and libraries to perform common tasks and operations.
Use CUDA memory spaces, such as global, shared, constant, and local, to optimize data transfers and memory accesses.
Use the CUDA execution model to control the threads, blocks, and grids that define the parallelism.
Debug and test CUDA programs using tools such as CUDA-GDB, CUDA-MEMCHECK, and NVIDIA Nsight.
Optimize CUDA programs using techniques such as coalescing, caching, prefetching, and profiling.
This live training in Varna assists intermediate AI developers in deploying models on Ascend processors through the CANN toolkit. You will learn to convert frameworks like PyTorch and TensorFlow, optimize performance, and troubleshoot issues to achieve efficient edge and cloud inference.
This live training on Varna empowers developers to program and optimize applications on Biren AI accelerators. Learners will explore GPU architecture, configure the SDK, and adapt CUDA code for Biren. The course emphasizes performance tuning and debugging techniques.
This live, instructor-led training in Varna provides developers with the competencies needed to construct and deploy AI models utilizing BANGPy and Neuware on Cambricon MLUs. Learners will manage environment configuration, build optimized models, and integrate MLU acceleration into both edge and data center applications.
This instructor-led, live training in Varna (online or onsite) is designed for beginner-level system administrators and IT professionals who want to learn how to install, configure, manage, and troubleshoot CUDA environments.
By the end of this training, participants will be able to:
Understand the architecture, components, and capabilities of CUDA.
This instructor-led, live training in Varna (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use OpenCL to program heterogeneous devices and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes OpenCL SDK, a device that supports OpenCL, and Visual Studio Code.
Create a basic OpenCL program that performs vector addition on the device and retrieves the results from the device memory.
Use OpenCL API to query device information, create contexts, command queues, buffers, kernels, and events.
Use OpenCL C language to write kernels that execute on the device and manipulate data.
Use OpenCL built-in functions, extensions, and libraries to perform common tasks and operations.
Use OpenCL host and device memory models to optimize data transfers and memory accesses.
Use OpenCL execution model to control the work-items, work-groups, and ND-ranges.
Debug and test OpenCL programs using tools such as CodeXL, Intel VTune, and NVIDIA Nsight.
Optimize OpenCL programs using techniques such as vectorization, loop unrolling, local memory, and profiling.
This instructor-led live training in Varna (online or onsite) is designed for C++ developers who want to use CUDA to accelerate applications, write high-performance GPU kernels, and leverage parallel algorithm libraries for scientific computing, data processing, and machine learning workloads.
This instructor-led, live training in Varna (online or onsite) is aimed at C/C++ developers who wish to use CUDA to accelerate compute-intensive applications, including data processing, scientific simulations, machine learning workloads, and image processing pipelines.
This instructor-led live training in Varna (offered online or onsite) is intended for software developers, data analysts, and technical experts who aim to use TensorFlow 2.x and Keras to build, train, and deploy deep learning models for computer vision, natural language processing, and multimodal applications.
This instructor-led, live training course in Varna covers how to program GPUs for parallel computing, how to use various platforms, how to work with the CUDA platform and its features, and how to perform various optimization techniques using CUDA. Some of the applications include deep learning, analytics, image processing and engineering applications.
Read more...
Last Updated:
Testimonials (1)
Trainers energy and humor.
Tadeusz Kaluba - Nokia Solutions and Networks Sp. z o.o.
Online GPU (Graphics Processing Unit) training in Varna, GPU (Graphics Processing Unit) training courses in Varna, Weekend Graphics Processing Unit (GPU) courses in Varna, Evening GPU training in Varna, GPU instructor-led in Varna, GPU instructor-led in Varna, Graphics Processing Unit trainer in Varna, GPU classes in Varna, GPU (Graphics Processing Unit) one on one training in Varna, GPU (Graphics Processing Unit) coaching in Varna, Graphics Processing Unit instructor in Varna, Online Graphics Processing Unit training in Varna, Weekend GPU training in Varna, Evening Graphics Processing Unit courses in Varna, GPU (Graphics Processing Unit) boot camp in Varna, Graphics Processing Unit (GPU) on-site in Varna, GPU private courses in Varna