Get in Touch
 Duration 14 hours

Course Outline

Preparing Machine Learning Models for Production

  • Encapsulating models using Docker
  • Exporting models from TensorFlow and PyTorch
  • Considerations for version control and storage

Serving Models on Kubernetes

  • Introduction to inference servers
  • Implementing TensorFlow Serving and TorchServe
  • Configuring model endpoints

Optimizing Inference Performance

  • Implementing batching methods
  • Managing concurrent request processing
  • Tuning for latency and throughput

Autoscaling ML Workloads

  • Horizontal Pod Autoscaler (HPA)
  • Vertical Pod Autoscaler (VPA)
  • Kubernetes Event-Driven Autoscaling (KEDA)

GPU Allocation and Resource Control

  • Setting up GPU-enabled nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML tasks

Model Release and Deployment Strategies

  • Blue/green deployment patterns
  • Canary release techniques
  • Conducting A/B tests for model validation

Monitoring and Observability for Production ML

  • Key metrics for inference operations
  • Best practices for logging and tracing
  • Creating dashboards and setting up alerts

Security and Reliability in Production

  • Hardening model endpoints
  • Applying network policies and access controls
  • Maintaining high availability

Wrap-Up and Future Directions

Requirements

  • A solid grasp of containerized application lifecycles
  • Practical experience with Python-based machine learning models
  • A working knowledge of core Kubernetes concepts

Target Audience

  • ML engineers
  • DevOps engineers
  • Platform engineering teams

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories