AI-Driven Observability: From Logs to LLM-Powered Insights Training Course
Conventional observability methods depend on dashboards, threshold-based alerts, and manual log investigation. AI-driven observability revolutionizes this approach by enabling natural language querying of telemetry data, utilizing LLMs for root cause analysis, employing foundation models for anomaly detection, and producing context-aware automated incident summaries.
This instructor-led live training, available online or onsite, targets observability and SRE engineers looking to incorporate LLMs and AI into their monitoring, alerting, and incident analysis workflows.
Upon completion of this training, participants will be equipped to:
- Create natural language interfaces for querying Prometheus, Elasticsearch, and SQL-based observability repositories.
- Develop pipelines for LLM-powered log analysis and anomaly detection.
- Produce automated incident summaries and draft postmortems from raw telemetry data.
- Design AI-assisted root cause analysis workflows featuring evidence chaining.
- Integrate foundation models for time-series anomaly detection and forecasting.
- Deploy an AI-enhanced on-call experience with intelligent alert enrichment.
Course Format
- Interactive lectures and discussions.
- Extensive exercises and practical practice.
- Hands-on implementation within a live lab environment.
Customization Options
- To request customized training, please contact us to make arrangements.
Course Outline
The AI Observability Landscape
- Transitioning from dashboards to conversations: the shift toward AI-augmented observability.
- Relevant LLM capabilities for observability: summarization, reasoning, and pattern matching.
- Architecture patterns for embedding AI into existing observability stacks.
Natural Language Telemetry Querying
- Text-to-PromQL: translating natural language into monitoring queries.
- NL querying for Elasticsearch, OpenSearch, and Loki log repositories.
- SQL generation from natural language for structured telemetry.
- Building a query assistant agent with tool use and context awareness.
LLM-Powered Log Analysis
- Automated log parsing and structuring using LLMs.
- Anomaly detection in log streams via embedding similarity.
- Log clustering and pattern discovery at scale.
- Generating human-readable explanations from raw log sequences.
Intelligent Alerting and Incident Enrichment
- Alert correlation and deduplication with semantic understanding.
- Automated incident context gathering from runbooks, historical incidents, and documentation.
- Smart alert routing based on content understanding and team expertise.
- Reducing alert fatigue through AI-driven noise reduction.
AI-Assisted Root Cause Analysis
- Hypothesis generation from multi-source telemetry correlation.
- Evidence chaining: linking symptoms across metrics, logs, and traces.
- Guided troubleshooting with interactive AI diagnosis sessions.
- Building a root cause analysis agent with progressive investigation capabilities.
Automated Incident Response and Communication
- Generating incident summaries and status updates from telemetry data.
- Automated postmortem drafting with timeline reconstruction.
- Tailored stakeholder communication for both technical and executive audiences.
- Runbook suggestions and automated remediation recommendations.
ML for Observability
- Time-series forecasting for capacity planning and anomaly prediction.
- Foundation models for zero-shot anomaly detection on metrics.
- Embedding-based service dependency mapping and topology discovery.
- Training and deploying lightweight ML models alongside observability pipelines.
Production Deployment and Ethics
- Latency and cost considerations for real-time AI observability.
- Data privacy: ensuring LLMs do not leak sensitive telemetry data.
- Human oversight: determining when AI diagnosis requires operator validation.
- Measuring impact: tracking MTTD, MTTR, and on-call experience metrics.
Requirements
- Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Familiarity with log management and metrics concepts.
- Basic Python scripting skills for data processing.
Target Audience
- SRE and observability engineers adopting AI-enhanced tooling.
- Platform engineers developing next-generation monitoring pipelines.
- DevOps leads evaluating the integration of LLMs into incident workflows.
Open Training Courses require 5+ participants.
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Booking
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Enquiry
AI-Driven Observability: From Logs to LLM-Powered Insights - Consultancy Enquiry
Upcoming Courses
Related Courses
Agentic Development with Gemini 3 and Google Antigravity
21 HoursGoogle Antigravity is an agentic development environment built to create autonomous agents that can plan, reason, code, and execute actions leveraging Gemini 3’s multimodal capabilities.
This live, instructor-led training session, available online or onsite, targets advanced technical professionals aiming to design, construct, and deploy autonomous agents using Gemini 3 and the Antigravity platform.
By the end of this program, participants will be equipped to:
- Construct autonomous workflows that leverage Gemini 3 for reasoning, planning, and execution.
- Create agents within Antigravity capable of analyzing tasks, generating code, and interacting with external tools.
- Seamlessly integrate Gemini-powered agents into enterprise systems and APIs.
- Refine agent behavior, ensuring safety and reliability within complex environments.
Course Format
- Expert-led demonstrations paired with interactive discussions.
- Hands-on experimentation in autonomous agent development.
- Practical implementation utilizing Antigravity, Gemini 3, and relevant cloud tools.
Customization Options
- If your team requires specific domain behaviors or custom integrations, please reach out to us to tailor the program to your needs.
Advanced Antigravity: Feedback Loops, Learning & Long-Term Agent Memory
14 HoursGoogle Antigravity serves as a sophisticated framework for exploring long-lived agents and the emergent interactive behaviors they exhibit.
This live, instructor-led training, available both online and on-site, is designed for advanced professionals seeking to design, analyze, and optimize agents that can retain memory, refine performance through feedback, and evolve across extended operational periods.
By the end of this course, participants will be equipped to:
- Architect long-term memory structures to ensure agent persistence.
- Establish effective feedback loops that guide agent behavior.
- Assess learning trajectories and monitor model drift.
- Embed memory mechanisms within complex multi-agent ecosystems.
Course Format
- Expert-facilitated discussions accompanied by technical demonstrations.
- Hands-on practice through structured design challenges.
- Practical application of concepts in simulated agent environments.
Customization Options
- Should your organization require specific content or case studies, please reach out to tailor this training to your needs.
Advanced Mastra Integrations: APIs, Tools, Enterprise Data & External Systems
21 HoursMastra is a framework that facilitates deep integration between AI agents, APIs, enterprise applications, and external data systems.
This instructor-led, live training (online or onsite) is aimed at intermediate-level engineers who wish to build reliable, secure, and scalable integrations between Mastra agents and the broader enterprise ecosystem.
Upon completing this training, participants will be prepared to:
- Implement API-driven integrations between Mastra agents and external services.
- Connect enterprise data systems and tools to automated agent workflows.
- Apply secure data exchange and authentication best practices.
- Design integration layers that are scalable, maintainable, and production ready.
Format of the Course
- Interactive lecture and discussion.
- Hands-on integration engineering and API exercises.
- Live-lab implementation using real-world enterprise scenarios.
Course Customization Options
- Custom API scenarios, enterprise system mappings, or data-integration workshops are available upon request.
Interactive AI Agents: AgentCore Memory, Code Interpreter & Browser Tool in Action
14 HoursAgentCore empowers AI agents to deliver highly interactive, dynamic, and context-aware experiences by providing robust memory persistence, a secure code interpreter, and advanced browser capabilities.
This instructor-led live training, available online or onsite, is designed for intermediate to advanced technical professionals who want to build and deploy AI agents that retain long-term context, perform real-time computations, and interact directly with web user interfaces.
Upon completion, participants will have the ability to:
- Deploy AgentCore memory to create stateful, context-aware workflows.
- Utilize the secure code interpreter for dynamic calculations and data transformations.
- Integrate the browser tool to facilitate real-time data acquisition and UI engagement.
- Develop interactive agents tailored for analytics, customer support, and research applications.
Training Delivery Format
- Engaging lectures combined with open discussion sessions.
- Practical lab exercises focused on AgentCore memory and tool utilization.
- In-depth case studies covering analytics, automation, and customer support scenarios.
Customization Options
- Reach out to our team to arrange a tailored training experience for this course.
Accelerating AI Agent Deployment with AgentCore Runtime & Gateway
14 HoursThe AgentCore Runtime & Gateway is an AWS service combination designed for packaging, deploying, and securely exposing AI agents, featuring streamlined integrations with external systems.
This instructor-led live training (available online or onsite) targets intermediate engineering teams looking to transition from agent prototypes to production by mastering the AgentCore Runtime for deployment and the Gateway for secure connectivity and API integration.
Upon completion of this training, participants will be able to:
- Establish AgentCore Runtime environments and package agents for deployment.
- Expose agents via Gateway using authenticated, rate-limited endpoints.
- Integrate external tools and APIs into agent workflows using stable contracts.
- Implement observability, logging, and usage monitoring for production operations.
Course Format
- Interactive lectures and discussions.
- Hands-on labs focused on Runtime deployments and Gateway integrations.
- Practical exercises emphasizing reliability, security, and rollout strategies.
Course Customization Options
- To request customized training for this course, please contact us to make arrangements.
Antigravity for Developers: Building Agent-First Applications
21 HoursAntigravity serves as a specialized development environment tailored for engineering AI-powered, agent-centric applications.
Delivered by expert instructors in either a live online or onsite setting, this course is tailored for intermediate developers seeking to build practical solutions leveraging autonomous AI agents within the Antigravity ecosystem.
Upon successful completion, participants will possess the capability to:
- Engineer applications that depend on autonomous and synchronized AI agents.
- Leverage the Antigravity IDE, editor, terminal, and browser for comprehensive end-to-end development.
- Orchestrate complex multi-agent workflows using the Agent Manager.
- Embed agent capabilities into robust, production-ready software architectures.
Instructional Approach
- Combined theoretical presentations with detailed practical demonstrations.
- Substantial hands-on practice through guided laboratory exercises.
- Direct implementation work conducted within the live Antigravity environment.
Customization Availability
- To align content specifically with your development stack, please reach out to us for a tailored training version.
Getting Started with Antigravity: An Introduction to Agent-First IDEs
14 HoursGoogle Antigravity represents a new class of agent-first development environments, engineered to optimize engineering workflows via intelligent automation.
This live, instructor-led training session—available either online or onsite—is tailored for beginners seeking to grasp the core principles of Antigravity and comprehend how agent-driven coding environments can boost productivity.
By the end of this course, participants will gain the ability to:
- Set up and configure Google Antigravity.
- Navigate and interpret both the Editor View and Manager View interfaces.
- Collaborate efficiently with agents to automate routine development tasks.
- Leverage Antigravity for generating, refining, and overseeing project files.
Course Delivery Method
- Instructor-led explanations complemented by live, real-time demonstrations.
- Supervised practical exercises emphasizing hands-on agent interaction.
- Guided exploration of key Antigravity features within a secure lab environment.
Customization Availability
- If you need a customized version of this training, please reach out to us to discuss a tailored program.
Antigravity for Web Automation & Browser-Based Tasks
21 HoursGoogle Antigravity is a platform designed to develop agents that interact with web applications, browser environments, and workflows spanning multiple surfaces.
This instructor-led live training, available online or on-site, targets intermediate-level professionals aiming to build, automate, and test browser-based workflows using Google Antigravity.
By the end of the training, participants will have the ability to:
- Develop agents that engage with web applications within a browser surface.
- Automate end-to-end workflows across various browser contexts.
- Validate and troubleshoot agent behavior in UI-driven environments.
- Implement cross-surface automation strategies utilizing Antigravity.
Course Format
- Guided instruction complemented by practical demonstrations.
- Hands-on activities and scenario-based exercises.
- Implementation of agent workflows within an interactive lab setting.
Customization Options
- To align the training with specific objectives, please contact us for tailored course customization.
Building Fully Managed AI Agents with AgentCore: From Concept to Production
14 HoursAgentCore streamlines the creation, optimization, and oversight of fully managed AI agents by offering a comprehensive suite of services designed for large-scale deployment.
This live, instructor-led session (available online or in-person) is designed for practitioners ranging from beginners to intermediates who are eager to acquire practical skills in building production-grade AI agents using AgentCore.
Upon completing this training, participants will be equipped to:
- Grasp the fundamental capabilities of AgentCore for developing AI agents.
- Architect and set up basic AI agents utilizing managed services.
- Combine workflows to bolster agent functionality.
- Launch and monitor AI agents within production settings.
Course Format
- Engaging lectures paired with interactive discussions.
- Practical labs focused on AgentCore services.
- Structured exercises guiding you from agent conception to deployment.
Customization Availability
- To request a tailored training experience for this course, please reach out to us to make arrangements.
AI Agent Development with Mastra
14 HoursThis live, instructor-led session—available both online and on-site—is specifically designed for mid-level software developers and engineering teams aiming to construct scalable, observable AI systems utilizing Mastra.
Upon completion of this program, attendees will possess the ability to:
- Grasp Mastra’s architectural design and its integration capabilities with LLMs and external APIs.
- Architect and develop AI agents and workflows using TypeScript.
- Leverage Mastra’s observability and memory utilities to track and enhance agent performance.
- Release production-grade AI applications by utilizing Mastra’s core framework capabilities.
Mastra Debugging, Evaluation & Quality Assurance for AI Agents
21 HoursMastra is a framework that delivers structured tools for evaluating, debugging, and ensuring the reliability of AI agents operating within complex workflows.
This instructor-led live training (available online or onsite) is designed for intermediate-level practitioners who want to rigorously test agent behavior, enhance reliability, and implement measurable evaluation processes.
Upon completion of this training, participants will be able to confidently:
- Apply debugging techniques to identify and correct issues in agent behavior.
- Evaluate agents using structured metrics, benchmarks, and quality scores.
- Implement tooling and workflows to monitor reliability, drift, and hallucinations.
- Design QA strategies to ensure consistent and predictable agent performance.
Course Format
- Interactive lectures and discussions.
- Hands-on exercises in debugging and evaluation.
- Live-lab analysis of agent behaviors using observability tools.
Customization Options
- Customized reliability testing scenarios and industry-specific QA methods can be arranged upon request.
Mastra Ops & Production Engineering: Deploying and Scaling AI Agents
21 HoursMastra is an operational framework designed to streamline the deployment, scaling, and lifecycle management of AI agents in production environments.
This instructor-led, live training (online or onsite) is aimed at intermediate-level to advanced-level technical professionals who need to operationalize AI agents reliably and efficiently across production systems.
Upon completion of this training, attendees will be equipped to:
- Deploy Mastra-based AI agents into controlled, production-grade environments.
- Scale agents horizontally and vertically using platform-native primitives.
- Implement observability pipelines to track agent behaviour and performance.
- Optimize runtime configurations to reduce latency, costs, and operational risks.
Format of the Course
- Interactive lecture and discussion.
- Hands-on exercises focused on real deployment scenarios.
- Live-lab implementation using containerized and orchestrated environments.
Course Customization Options
- Customization of topics, hands-on labs, or industry-specific scenarios is available upon request.
Mastra Workflow Automation & Multi-Agent Orchestration
21 HoursMastra is a framework that empowers sophisticated workflow automation and coordination across multiple AI agents within distributed systems.
This instructor-led live training, available both online and onsite, targets intermediate-level professionals seeking to design, orchestrate, and manage multi-agent workflows at scale.
Upon completion, participants will acquire the skills to:
- Design intricate workflows leveraging Mastra's orchestration capabilities.
- Coordinate multiple agents executing parallel or dependent tasks.
- Implement monitoring and debugging tools for effective workflow execution.
- Optimize orchestration logic to enhance reliability, throughput, and automation efficiency.
Course Format
- Interactive lectures and discussions.
- Hands-on exercises focused on workflow design and automation.
- Practical implementation within a containerized live-lab environment.
Customization Options
- Customized automation scenarios, enterprise integrations, or workflow patterns can be provided upon request.
Managing Agent Workflows in Google Antigravity: Orchestration, Planning and Artifacts
14 HoursGoogle Antigravity functions as an agent-centric development platform, designed to orchestrate, monitor, and coordinate AI-powered coding and automation tasks.
This live, instructor-led training session, available online or on-site, is tailored for intermediate-level professionals seeking to design, oversee, and enhance multi-agent workflows within the Google Antigravity ecosystem.
By the end of this program, participants will be equipped to:
- Define agent responsibilities and set up orchestration pipelines through the Manager interface.
- Create and analyze Antigravity artifacts, such as task lists, plans, logs, and browser recordings.
- Apply verification strategies to maintain transparency and auditability of agent actions.
- Enhance multi-agent collaboration for intricate development and operational requirements.
Course Format
- Structured presentations accompanied by practical demonstrations.
- Scenario-based exercises addressing real-world workflow challenges.
- Direct experimentation within an active Antigravity workspace.
Customization Options
- For a customized version of this course, please reach out to discuss specific requirements.
Testing & Verifying Agent-Driven Code: Quality Assurance in Antigravity
14 HoursAntigravity is a framework designed to support sophisticated agent-driven development processes.
This live, instructor-led training session, available online or onsite, is tailored for intermediate to advanced professionals seeking to validate, verify, and secure the outputs generated by AI agents operating within Antigravity-based environments.
By the end of this training, participants will be equipped to:
- Evaluate the precision and safety of code artifacts created by agents.
- Employ structured methods to verify tasks executed by agents.
- Effectively analyze browser recordings and trace agent activities.
- Implement QA and security standards to guarantee the reliability of agent workflows.
Course Delivery Format
- Technical briefings and discussions guided by an instructor.
- Practical exercises centered on verifying real-world agent workflows.
- Hands-on testing and validation conducted in a controlled lab setting.
Customization Options
- Scenarios, workflows, and testing examples can be tailored to specific needs upon request.