Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours
Course Outline
Introduction to Ollama Scaling
- Architectural overview and key scaling factors for Ollama
- Typical bottlenecks encountered in multi-user setups
- Recommended practices for preparing infrastructure
Resource Allocation and GPU Tuning
- Strategies for maximizing CPU and GPU usage
- Key aspects of memory and bandwidth management
- Setting resource limits at the container level
Container and Kubernetes Deployment
- Encapsulating Ollama using Docker
- Operating Ollama within Kubernetes clusters
- Implementing load balancing and service discovery
Autoscaling and Batching Techniques
- Crafting autoscaling rules for Ollama
- Using batch inference to boost throughput
- Balancing latency against throughput requirements
Latency Improvement
- Analyzing inference performance
- Implementing caching and model warm-up methods
- Minimizing I/O and communication delays
Monitoring and Observability
- Connecting Prometheus for metric collection
- Creating dashboards using Grafana
- Setting up alerts and incident response for Ollama infrastructure
Cost Control and Scaling Approaches
- Allocating GPUs with cost in mind
- Comparing cloud versus on-premises deployment options
- Methods for achieving sustainable scale
Conclusion and Future Directions
Requirements
- Proficiency in Linux system administration
- Knowledge of containerization and orchestration concepts
- Experience with deploying machine learning models
Target Audience
- DevOps Engineers
- ML Infrastructure Teams
- Site Reliability Engineers