Delivered either online or onsite, our instructor-led live GPU (Graphics Processing Unit) training programs use interactive discussion and practical exercises to demonstrate the core principles of GPU technology and the essential skills needed to program them.
GPU training is offered as "online live training" or "onsite live training". Online live training (also known as "remote live training") is conducted through an interactive remote desktop. Onsite live training takes place locally at customer premises in Sofia or at NobleProg corporate training centers in Sofia.
NobleProg -- Your Local Training Provider
Crystal Business Center
ул. "Осогово" 40, Sofia, Bulgaria, 1303
Crystal Business Center is located in the central part of Sofia, on the corner of "Osogovo" street. and "Todor Aleksandrov" blvd. The building is easily accessible by metro (only 50 m from Opalchenska station) and other public transport. Its total area is 8000 sq.m. The office area is 6171 sq.m.
This instructor-led course in Sofia assists intermediate AI engineers in constructing and optimizing neural network models via the Huawei Ascend platform and CANN toolkit. Learners will set up environments, build applications using MindSpore, and manage deployment to edge or cloud settings.
This live, instructor-led training in Sofia delves into Huawei's AI stack, covering the spectrum from the CANN SDK to the MindSpore framework. It is designed to assist beginner and intermediate professionals in understanding how these components synergize on Ascend hardware to streamline lifecycle management and deployment.
This instructor-led, live training in Sofia (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use OpenACC to program heterogeneous devices and exploit their parallelism.
By the end of this training, participants will be able to:
Set up an OpenACC development environment.
Write and run a basic OpenACC program.
Annotate code with OpenACC directives and clauses.
This instructor-led training on Sofia focuses on deploying and optimizing CV and NLP models using the CANN SDK for Ascend hardware. Learners will master model conversion, integration into live pipelines, and enhancing inference performance for real-time detection and analysis.
This instructor-led live training in Sofia (online or onsite) is tailored for beginner to intermediate-level developers who wish to understand the basics of GPU programming and the key frameworks and tools used for developing GPU applications.
By the end of this training, participants will be able to: Understand the difference between CPU and GPU computing and the benefits and challenges of GPU programming.
Choose the right framework and tool for their GPU application.
Create a basic GPU program that performs vector addition using one or more of the frameworks and tools.
Use the respective APIs, languages, and libraries to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use the respective memory spaces, such as global, local, constant, and private, to optimize data transfers and memory accesses.
Use the respective execution models, such as work-items, work-groups, threads, blocks, and grids, to control the parallelism.
Debug and test GPU programs using tools such as CodeXL, CUDA-GDB, CUDA-MEMCHECK, and NVIDIA Nsight.
Optimize GPU programs using techniques such as coalescing, caching, prefetching, and profiling.
This instructor-led, live training in Sofia equips advanced developers with the skills to build, deploy, and tune custom AI operators. Participants will master CANN TIK and Apache TVM integration, enabling advanced optimization and scheduling on Huawei Ascend hardware for real-world performance.
This instructor-led, live training in Sofia (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use different frameworks for GPU programming and compare their features, performance, and compatibility.
By the end of this training, participants will be able to:
Set up a development environment that includes OpenCL SDK, CUDA Toolkit, ROCm Platform, a device that supports OpenCL, CUDA, or ROCm, and Visual Studio Code.
Create a basic GPU program that performs vector addition using OpenCL, CUDA, and ROCm, and compare the syntax, structure, and execution of each framework.
Use the respective APIs to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use the respective languages to write kernels that execute on the device and manipulate data.
Use the respective built-in functions, variables, and libraries to perform common tasks and operations.
Use the respective memory spaces, such as global, local, constant, and private, to optimize data transfers and memory accesses.
Use the respective execution models to control the threads, blocks, and grids that define the parallelism.
Debug and test GPU programs using tools such as CodeXL, CUDA-GDB, CUDA-MEMCHECK, and NVIDIA Nsight.
Optimize GPU programs using techniques such as coalescing, caching, prefetching, and profiling.
This live training session in Sofia provides a comprehensive overview of CloudMatrix for scalable AI inference. Discover how to deploy, optimize, and monitor models using CANN and MindSpore. Practical exercises will guide you through packaging, conversion, serving, and performance tuning for both real-time and batch workloads.
This instructor-led live training in Sofia explores the fundamental concepts and practical skills required for deploying AI models on Ascend edge devices using the CANN toolkit, equipping participants with the ability to compile, optimize, and manage performance in constrained environments.
This instructor-led, live training in Sofia (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to install and use ROCm on Windows to program AMD GPUs and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes ROCm Platform, a AMD GPU, and Visual Studio Code on Windows.
Create a basic ROCm program that performs vector addition on the GPU and retrieves the results from the GPU memory.
Use ROCm API to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use HIP language to write kernels that execute on the GPU and manipulate data.
Use HIP built-in functions, variables, and libraries to perform common tasks and operations.
Use ROCm and HIP memory spaces, such as global, shared, constant, and local, to optimize data transfers and memory accesses.
Use ROCm and HIP execution models to control the threads, blocks, and grids that define the parallelism.
Debug and test ROCm and HIP programs using tools such as ROCm Debugger and ROCm Profiler.
Optimize ROCm and HIP programs using techniques such as coalescing, caching, prefetching, and profiling.
This instructor-led, live training in Sofia (online or onsite) is aimed at beginner to intermediate developers who wish to use ROCm and HIP to program AMD GPUs and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes ROCm Platform, an AMD GPU, and Visual Studio Code.
Create a basic ROCm program that performs vector addition on the GPU and retrieves the results from the GPU memory.
Use ROCm API to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use HIP language to write kernels that execute on the GPU and manipulate data.
Use HIP built-in functions, variables, and libraries to perform common tasks and operations.
Use ROCm and HIP memory spaces, such as global, shared, constant, and local, to optimize data transfers and memory accesses.
Use ROCm and HIP execution models to control the threads, blocks, and grids that define the parallelism.
Debug and test ROCm and HIP programs using tools such as ROCm Debugger and ROCm Profiler.
Optimize ROCm and HIP programs using techniques such as coalescing, caching, prefetching, and profiling.
This live training in Sofia introduces the CANN toolkit to AI framework developers. Learn to set up environments, convert models, and deploy applications on Ascend hardware using MindSpore, TensorFlow, or PyTorch, covering the full workflow from training to inference.
Enhance AI workload performance on Ascend, Biren, and Cambricon through this practical training in Sofia. Master model benchmarking, pinpoint bottlenecks, and deploy graph, kernel, and operator-level optimizations. Refine deployment pipelines to optimize throughput and latency across these top-tier platforms.
Maximize neural network inference efficiency on Ascend AI processors through this advanced, instructor-led training in Sofia. Delve into CANN's runtime architecture, utilizing the Graph Engine, TIK, and TVM for profiling, custom operator creation, and resolving memory bottlenecks.
Transition CUDA applications to Chinese GPU ecosystems, such as Huawei Ascend and Biren, within Sofia. This instructor-led course assists advanced programmers in code translation and performance refinement, featuring practical labs for porting CUDA codebases to emerging SDKs.
This instructor-led, live training in Sofia (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use CUDA to program NVIDIA GPUs and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes CUDA Toolkit, an NVIDIA GPU, and Visual Studio Code.
Create a basic CUDA program that performs vector addition on the GPU and retrieves the results from the GPU memory.
Use the CUDA API to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use the CUDA C/C++ language to write kernels that execute on the GPU and manipulate data.
Use CUDA built-in functions, variables, and libraries to perform common tasks and operations.
Use CUDA memory spaces, such as global, shared, constant, and local, to optimize data transfers and memory accesses.
Use the CUDA execution model to control the threads, blocks, and grids that define the parallelism.
Debug and test CUDA programs using tools such as CUDA-GDB, CUDA-MEMCHECK, and NVIDIA Nsight.
Optimize CUDA programs using techniques such as coalescing, caching, prefetching, and profiling.
This live training in Sofia assists intermediate AI developers in deploying models on Ascend processors through the CANN toolkit. You will learn to convert frameworks like PyTorch and TensorFlow, optimize performance, and troubleshoot issues to achieve efficient edge and cloud inference.
This live training on Sofia empowers developers to program and optimize applications on Biren AI accelerators. Learners will explore GPU architecture, configure the SDK, and adapt CUDA code for Biren. The course emphasizes performance tuning and debugging techniques.
This live, instructor-led training in Sofia provides developers with the competencies needed to construct and deploy AI models utilizing BANGPy and Neuware on Cambricon MLUs. Learners will manage environment configuration, build optimized models, and integrate MLU acceleration into both edge and data center applications.
This instructor-led, live training in Sofia (online or onsite) is designed for beginner-level system administrators and IT professionals who want to learn how to install, configure, manage, and troubleshoot CUDA environments.
By the end of this training, participants will be able to:
Understand the architecture, components, and capabilities of CUDA.
This instructor-led, live training in Sofia (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use OpenCL to program heterogeneous devices and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes OpenCL SDK, a device that supports OpenCL, and Visual Studio Code.
Create a basic OpenCL program that performs vector addition on the device and retrieves the results from the device memory.
Use OpenCL API to query device information, create contexts, command queues, buffers, kernels, and events.
Use OpenCL C language to write kernels that execute on the device and manipulate data.
Use OpenCL built-in functions, extensions, and libraries to perform common tasks and operations.
Use OpenCL host and device memory models to optimize data transfers and memory accesses.
Use OpenCL execution model to control the work-items, work-groups, and ND-ranges.
Debug and test OpenCL programs using tools such as CodeXL, Intel VTune, and NVIDIA Nsight.
Optimize OpenCL programs using techniques such as vectorization, loop unrolling, local memory, and profiling.
This instructor-led live training in Sofia (online or onsite) is designed for C++ developers who want to use CUDA to accelerate applications, write high-performance GPU kernels, and leverage parallel algorithm libraries for scientific computing, data processing, and machine learning workloads.
This instructor-led, live training in Sofia (online or onsite) is aimed at C/C++ developers who wish to use CUDA to accelerate compute-intensive applications, including data processing, scientific simulations, machine learning workloads, and image processing pipelines.
This instructor-led live training in Sofia (offered online or onsite) is intended for software developers, data analysts, and technical experts who aim to use TensorFlow 2.x and Keras to build, train, and deploy deep learning models for computer vision, natural language processing, and multimodal applications.
This instructor-led, live training course in Sofia covers how to program GPUs for parallel computing, how to use various platforms, how to work with the CUDA platform and its features, and how to perform various optimization techniques using CUDA. Some of the applications include deep learning, analytics, image processing and engineering applications.
Read more...
Last Updated:
Testimonials (1)
Trainers energy and humor.
Tadeusz Kaluba - Nokia Solutions and Networks Sp. z o.o.
Online GPU training in Sofia, Graphics Processing Unit training courses in Sofia, Weekend Graphics Processing Unit courses in Sofia, Evening Graphics Processing Unit (GPU) training in Sofia, GPU instructor-led in Sofia, Graphics Processing Unit private courses in Sofia, Graphics Processing Unit (GPU) trainer in Sofia, GPU instructor in Sofia, Evening Graphics Processing Unit courses in Sofia, Graphics Processing Unit (GPU) instructor-led in Sofia, Graphics Processing Unit classes in Sofia, Online Graphics Processing Unit training in Sofia, GPU on-site in Sofia, Graphics Processing Unit one on one training in Sofia, GPU (Graphics Processing Unit) coaching in Sofia, Graphics Processing Unit boot camp in Sofia, Weekend Graphics Processing Unit training in Sofia