Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Tencent Hunyuan Production Fundamentals
- Overview of Tencent Hunyuan model serving scenarios.
- Production characteristics of large and MoE models.
- Common bottlenecks related to latency, throughput, and cost.
- Defining service-level objectives for inference workloads.
Deployment Architecture and Serving Flow
- Core components of a production inference stack.
- Evaluating containerized, on-premise, and cloud deployment models.
- Fundamentals of model loading, request routing, and GPU allocation.
- Designing for reliability and operational simplicity.
Latency Optimization in Practice
- Leveraging optimized inference engines such as TensorRT where applicable.
- KV-cache concepts and practical cache tuning techniques.
- Reducing startup, warmup, and response overhead.
- Measuring time to first token (TTFT) and token generation speed.
Throughput, Batching, and GPU Efficiency
- Strategies for continuous batching and request batching.
- Managing concurrency and queue behavior effectively.
- Improving GPU utilization without compromising user experience.
- Handling long-context and mixed-workload requests.
Quantization and Cost Control
- Understanding the importance of quantization for production serving.
- Practical trade-offs of FP16, INT8, and other common precision options.
- Balancing model quality, latency, and infrastructure costs.
- Creating a simple checklist for cost optimization.
Operations, Monitoring, and Readiness Review
- Autoscaling triggers for inference services.
- Monitoring latency, throughput, cache usage, and GPU health.
- Fundamentals of logging, alerting, and incident response.
- Reviewing a reference deployment and developing an improvement plan.
Requirements
- Fundamental understanding of large language model (LLM) deployment and inference workflows.
- Experience with containers, cloud or on-premise infrastructure, and API-based services.
- Working knowledge of Python or system engineering tasks.
Audience
- ML engineers deploying LLMs into production environments.
- Platform engineers responsible for GPU-based inference services.
- Solution architects designing scalable AI serving platforms.