Get in Touch
 Duration 21 hours

Course Outline

Introduction to Ollama Scaling

  • Architectural overview and key scaling factors for Ollama
  • Typical bottlenecks encountered in multi-user setups
  • Recommended practices for preparing infrastructure

Resource Allocation and GPU Tuning

  • Strategies for maximizing CPU and GPU usage
  • Key aspects of memory and bandwidth management
  • Setting resource limits at the container level

Container and Kubernetes Deployment

  • Encapsulating Ollama using Docker
  • Operating Ollama within Kubernetes clusters
  • Implementing load balancing and service discovery

Autoscaling and Batching Techniques

  • Crafting autoscaling rules for Ollama
  • Using batch inference to boost throughput
  • Balancing latency against throughput requirements

Latency Improvement

  • Analyzing inference performance
  • Implementing caching and model warm-up methods
  • Minimizing I/O and communication delays

Monitoring and Observability

  • Connecting Prometheus for metric collection
  • Creating dashboards using Grafana
  • Setting up alerts and incident response for Ollama infrastructure

Cost Control and Scaling Approaches

  • Allocating GPUs with cost in mind
  • Comparing cloud versus on-premises deployment options
  • Methods for achieving sustainable scale

Conclusion and Future Directions

Requirements

  • Proficiency in Linux system administration
  • Knowledge of containerization and orchestration concepts
  • Experience with deploying machine learning models

Target Audience

  • DevOps Engineers
  • ML Infrastructure Teams
  • Site Reliability Engineers

Number of participants


Price per participant

Upcoming Courses

Related Categories