Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
AI Sovereignty and Local LLM Deployment
- Risks associated with cloud LLMs: data retention, training on user inputs, and foreign legal jurisdictions.
- Ollama architecture overview: model server, registry, and OpenAI-compatible API.
- Comparison with vLLM, llama.cpp, and Text Generation Inference.
- Model licensing terms for Llama, Mistral, Qwen, and Gemma.
Installation and Hardware Configuration
- Installing Ollama on Linux with CUDA and ROCm support.
- CPU-only fallback options and AVX/AVX2 optimization techniques.
- Docker deployment methods and persistent volume mapping.
- Multi-GPU setup and VRAM allocation strategies.
Model Management
- Pulling models from the Ollama registry: ollama pull llama3.
- Importing GGUF models from HuggingFace and TheBloke.
- Understanding quantization levels: Q4_K_M, Q5_K_M, and Q8_0 trade-offs.
- Switching models and understanding limits for concurrent model loading.
Custom Modelfiles
- Syntax for writing Modelfiles: FROM, PARAMETER, SYSTEM, and TEMPLATE directives.
- Tuning temperature, top_p, and repeat_penalty parameters.
- Engineering system prompts for role-specific behaviors.
- Creating and publishing custom models to the local registry.
API Integration
- Utilizing the OpenAI-compatible /v1/chat/completions endpoint.
- Handling streaming responses and JSON mode.
- Integrating with LangChain, LlamaIndex, and custom applications.
- Implementing authentication and rate limiting via reverse proxies.
Performance Optimization
- Managing context window sizing and KV cache allocation.
- Conducting batch inference and handling parallel requests.
- Allocating CPU threads and ensuring NUMA awareness.
- Monitoring GPU utilization and memory pressure.
Security and Compliance
- Establishing network isolation for model serving endpoints.
- Implementing input filtering and output moderation pipelines.
- Audit logging of prompts and generated completions.
- Verifying model provenance through hash checks.
Requirements
- Intermediate proficiency in Linux and container administration.
- High-level understanding of machine learning concepts and transformer models.
- Familiarity with REST APIs and JSON data formats.
Audience
- AI engineers and developers migrating away from cloud LLM APIs.
- Organizations handling sensitive data that restricts the use of cloud models.
- Government and defense teams requiring air-gapped language model solutions.
14 Hours