Whether delivered online or on-site, our instructor-led live Big Data training programs begin by establishing the foundational concepts of Big Data. The curriculum then advances to the programming languages and methodologies essential for conducting Data Analysis. We explore, compare, and apply the tools and infrastructure required for Big Data storage, Distributed Processing, and Scalability through practical demonstration sessions.
Big Data training can be delivered as either "online live training" or "onsite live training". Online live training, also referred to as "remote live training", is conducted via an interactive remote desktop. Onsite live training can take place at your premises in Sofia or at NobleProg corporate training centers in Sofia.
NobleProg -- Your Local Training Provider
Crystal Business Center
ул. "Осогово" 40, Sofia, Bulgaria, 1303
Crystal Business Center is located in the central part of Sofia, on the corner of "Osogovo" street. and "Todor Aleksandrov" blvd. The building is easily accessible by metro (only 50 m from Opalchenska station) and other public transport. Its total area is 8000 sq.m. The office area is 6171 sq.m.
This instructor-led, live training in Sofia (online or onsite) is aimed at intermediate-level data scientists and engineers who wish to use Google Colab and Apache Spark for big data processing and analytics.
By the end of this training, participants will be able to:
Set up a big data environment using Google Colab and Spark.
Process and analyze large datasets efficiently with Apache Spark.
Visualize big data in a collaborative environment.
This instructor-led live training in Sofia covers the Stratio platform, focusing on the Rocket and Intelligence modules with PySpark. Participants will master data ingestion, transformation, and advanced analytics, gaining practical skills in loops, UDFs, and enterprise data workflows.
This live training in Sofia covers data warehousing concepts, dimensional modeling, and ETL pipeline design. Participants will build star schemas, optimize OLAP workloads, and implement governance, gaining practical skills for building robust analytical systems.
Upon completion of this instructor-led, live training in Sofia, participants will acquire a practical, real-world understanding of Big Data, along with its associated technologies, methodologies, and tools.
Through hands-on exercises, participants will have the opportunity to apply this knowledge directly. Group interaction and feedback from the instructor are key components of the class experience.
The course begins with an introduction to the fundamental concepts of Big Data, then moves on to the programming languages and methodologies employed for Data Analysis. Finally, we explore the tools and infrastructure that support Big Data storage, Distributed Processing, and Scalability.
This live training in Sofia helps intermediate to advanced users master Greenplum architecture and data modeling. Participants will learn to design distributed tables, apply high-performance SQL, and interpret EXPLAIN plans for optimal query execution in large-scale analytic environments.
This instructor-led training in Sofia covers Greenplum installation, updates, and library management. Participants learn to configure clusters, apply safe patches, and handle extensions for advanced analytics, building practical skills for real-world environments.
This instructor-led, live training in Sofia (online or onsite) is aimed at advanced-level data professionals who wish to optimize data processing workflows, ensure data integrity, and implement robust data lakehouse solutions that can handle the complexities of modern big data applications.
By the end of this training, participants will be able to:
Gain an in-depth understanding of Iceberg’s architecture, including metadata management and file layout.
Configure Iceberg for optimal performance in various environments and integrate it with multiple data processing engines.
This instructor-led live training in Sofia (online or onsite) targets beginner-level data professionals seeking to acquire the knowledge and skills necessary to effectively utilize Apache Iceberg for managing large-scale datasets, ensuring data integrity, and optimizing data processing workflows.
Upon completing this training, participants will be able to:
Develop a comprehensive understanding of Apache Iceberg's architecture, features, and benefits.
Explore table formats, partitioning strategies, schema evolution, and time travel capabilities.
Install and configure Apache Iceberg across various environments.
Create, manage, and manipulate Iceberg tables.
Comprehend the process of migrating data from other table formats to Iceberg.
This instructor-led live training in Sofia (online or onsite) targets intermediate-level IT professionals seeking to enhance their skills in data architecture, governance, cloud computing, and big data technologies to effectively manage and analyze large datasets for organizational data migration.
By the end of this training, participants will be able to:
Understand the foundational concepts and components of various data architectures.
Gain a comprehensive understanding of data governance principles and their importance in regulatory environments.
Implement and manage data governance frameworks such as DAMA and TOGAF.
Leverage cloud platforms for efficient data storage, processing, and management.
This instructor-led live training in Sofia (online or onsite) is aimed at intermediate-level data engineers who wish to learn how to use Azure Data Lake Storage Gen2 for effective data analytics solutions.
By the end of this training, participants will be able to:
Understand the architecture and key features of Azure Data Lake Storage Gen2.
Optimize data storage and access for cost and performance.
Integrate Azure Data Lake Storage Gen2 with other Azure services for analytics and data processing.
Develop solutions using the Azure Data Lake Storage Gen2 API.
Troubleshoot common issues and optimize storage strategies.
This instructor-led, live training in Sofia (online or onsite) is aimed at developers who wish to use and integrate Spark, Hadoop, and Python to process, analyze, and transform large and complex data sets.
By the end of this training, participants will be able to:
Set up the necessary environment to start processing big data with Spark, Hadoop, and Python.
Understand the features, core components, and architecture of Spark and Hadoop.
Learn how to integrate Spark, Hadoop, and Python for big data processing.
Explore the tools in the Spark ecosystem (Spark MlLib, Spark Streaming, Kafka, Sqoop, Kafka, and Flume).
Build collaborative filtering recommendation systems similar to Netflix, YouTube, Amazon, Spotify, and Google.
Use Apache Mahout to scale machine learning algorithms.
This instructor-led live training in Sofia covers advanced big data techniques, distributed computing, and machine learning at scale. It is designed for advanced data professionals aiming to master real-time analytics, deep learning integration, and robust data governance strategies.
This instructor-led live training in Sofia (online or onsite) targets intermediate-level IT professionals seeking a comprehensive understanding of IBM DataStage from both administrative and development perspectives. This knowledge enables them to manage and utilize the tool effectively in their respective workplaces.
By the end of this training, participants will be able to:
Understand the core concepts of DataStage.
Learn how to effectively install, configure, and manage DataStage environments.
Connect to various data sources and extract data efficiently from databases, flat files, and external sources.
Conducted by a live instructor, this training on Sofia (available online or on-site) is designed for system administrators with beginner to intermediate expertise who aim to deploy, maintain, and optimize Spark clusters.
By the conclusion of this training, participants will be empowered to:
Install and configure Apache Spark in various environments.
Manage cluster resources and monitor Spark applications.
Optimize the performance of Spark clusters.
Implement security measures and ensure high availability.
In this live, instructor-led training in Sofia, participants will explore how to harness the combined power of Python and Spark for big data analysis. The course emphasizes practical application through a sequence of hands-on exercises.
Upon completing this training, participants will be able to:
Utilize Spark alongside Python to conduct comprehensive Big Data analysis.
Complete exercises that reflect authentic, real-world data scenarios.
Employ diverse tools and techniques for big data analysis using the PySpark framework.
Explore Big Data BI for government agencies in Sofia. This course covers Hadoop, NoSQL, predictive analytics, and real-time tools to manage vast, diverse data streams. Learn fraud detection, cybersecurity, and ROI strategies to transform unstructured data into strategic assets for mission success.
In this instructor-led live training in Sofia, participants will learn the mindset with which to approach Big Data technologies, assess their impact on existing processes and policies, and implement these technologies for the purpose of identifying criminal activity and preventing crime. Case studies from law enforcement organizations around the world will be examined to gain insights on their adoption approaches, challenges and results.
By the end of this training, participants will be able to:
Combine Big Data technology with traditional data gathering processes to piece together a story during an investigation.
Implement industrial big data storage and processing solutions for data analysis.
Prepare a proposal for the adoption of the most adequate tools and processes for enabling a data-driven approach to criminal investigation.
Explore the fundamentals of programming with Big Data in R, a language popular in finance. This course in Sofia covers environment setup, MPI for parallel processing, and distributed matrices. Learn to handle large datasets, perform distributed regression, and apply Monte Carlo methods efficiently.
This instructor-led live training (online or onsite) is designed for data engineers, data analysts, and data professionals who wish to use Databricks and PySpark to build scalable data pipelines and migrate existing SQL workflows.
In Sofia, this five-day training introduces real time data streaming systems. It covers core concepts, architecture patterns, and industry tools for processing continuous data at scale. Learners design, implement, and optimize scalable streaming pipelines using modern frameworks.
This instructor-led live training in Sofia is designed for intermediate administrators aiming to deploy and manage Apache NiFi in production. Participants will learn to configure clusters, design dataflows, and optimize performance through hands-on labs and scenario-based exercises.
This course offers a hands-on introduction to developing scalable data processing and Machine Learning workflows with PySpark. Attendees will gain insight into how Apache Spark functions within contemporary Big Data ecosystems and how to effectively manage large datasets by leveraging distributed computing principles.
This intensive three-day workshop is designed to help you build and refine high-performance data-processing workloads by leveraging PySpark, Pandas and Polars within Kubernetes-based environments.
Attendees will gain a functional grasp of how Spark applications operate on Kubernetes, and how specific configuration choices at the application level directly impact performance, scalability, resource utilisation and overall costs. Key optimisation topics covered include executor sizing, memory allocation strategies, dynamic allocation, partitioning logic, shuffle mechanics, resolving small-file issues and enhancing Parquet processing efficiency.
Additionally, the course tackles frequent challenges encountered with Pandas, such as memory constraints and out-of-memory errors, while introducing Polars as a high-performance solution for specific data-processing tasks. Through practical, hands-on labs, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration approaches and implement optimisation techniques in realistic ETL and machine learning contexts.
Throughout the training, the focus remains on practical decision-making: learning how to pinpoint bottlenecks, choose the right tools, configure Spark for maximum efficiency and strike a balance between performance and infrastructure costs.
This instructor-led live training in Sofia (online or onsite) is designed for engineers aiming to set up and deploy an Apache Spark system for processing massive data volumes.
By the end of this training, participants will be able to:
Install and configure Apache Spark.
Efficiently process and analyze extensive data sets.
Distinguish between Apache Spark and Hadoop MapReduce and understand their respective use cases.
Integrate Apache Spark with external machine learning tools.
This hands-on training in Sofia demystifies Apache Spark, covering RDDs, DataFrames, and Python/Scala APIs. Participants will master cloud deployment with Databricks, AWS EMR, and Glue, building practical skills for real-world data engineering and DevOps tasks effectively.
This instructor-led live training in Sofia (online or onsite) is aimed at technical professionals who wish to deploy Talend Open Studio for Big Data to simplify the process of reading and analyzing big data.
By the end of this training, participants will be able to:
Install and configure Talend Open Studio for Big Data.
Connect with big data systems such as Cloudera, Hortonworks, MapR, Amazon EMR, and Apache.
Understand and set up Open Studio's big data components and connectors.
Configure parameters to automatically generate MapReduce code.
Use Open Studio's drag-and-drop interface to run Hadoop jobs.
Prototype big data pipelines.
Automate big data integration projects.
Read more...
Last Updated:
Testimonials (7)
A journey through the Spark world: a very intense course. DSL, spark sql, partitioning vs bucketing for me.
Georgiana Elisabeta
Course - Apache Spark Fundamentals
the practices
Liliana Padilla - Hipodromo de Agua Caliente
Course - Greenplum Architecture and Data Modeling
Hands on exercises. Class should have been 5 days, but the 3 days helped to clear up a lot of questions that I had from working with NiFi already
James - BHG Financial
Course - Apache NiFi for Administrators
The ability of the trainer to align the course with the requirements of the organization other than just providing the course for the sake of delivering it.
Masilonyane - Revenue Services Lesotho
Course - Big Data Business Intelligence for Govt. Agencies
The fact that we were able to take with us most of the information/course/presentation/exercises done, so that we can look over them and perhaps redo what we didint understand first time or improve what we already did.
Raul Mihail Rat - Accenture Industrial SS
Course - Python, Spark, and Hadoop for Big Data
Having hands on session / assignments
Poornima Chenthamarakshan - Intelligent Medical Objects
Course - Apache Spark in the Cloud
The subject matter and the pace were perfect.
Tim - Ottawa Research and Development Center, Science Technology Branch, Agriculture and Agri-Food Canada
Online Big Data training in Sofia, Big Data training courses in Sofia, Weekend Big Data courses in Sofia, Evening Big Data training in Sofia, Big Data instructor-led in Sofia, Big Data instructor-led in Sofia, Big Data private courses in Sofia, Evening Big Data courses in Sofia, Big Data trainer in Sofia, Big Data one on one training in Sofia, Big Data on-site in Sofia, Big Data boot camp in Sofia, Big Data instructor in Sofia, Big Data coaching in Sofia, Big Data classes in Sofia, Weekend Big Data training in Sofia, Online Big Data training in Sofia