Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours (3 days)
Course Outline
Basics of Audio Classification
- Categories of sound events: environmental, mechanical, and human-made
- Summary of applications: surveillance, monitoring, and automation
- Distinguishing between audio classification, detection, and segmentation
Audio Data and Feature Extraction
- Various audio file types and formats
- Considerations regarding sampling rates, windowing, and frame sizes
- Extraction of MFCCs, chroma features, and mel-spectrograms
Data Preparation and Annotation
- Utilizing UrbanSound8K, ESC-50, and proprietary datasets
- Labeling sound events and defining temporal boundaries
- Dataset balancing and audio augmentation techniques
Constructing Audio Classification Models
- Application of convolutional neural networks (CNNs) to audio data
- Input strategies: raw waveforms versus extracted features
- Loss functions, evaluation metrics, and managing overfitting
Event Detection and Temporal Localization
- Frame-based and segment-based detection approaches
- Refining detections through thresholding and smoothing
- Mapping predictions onto audio timelines for visualization
Advanced Concepts and Real-Time Processing
- Leveraging transfer learning for scenarios with limited data
- Model deployment using TensorFlow Lite or ONNX
- Streaming audio handling and latency management
Project Development and Real-World Scenarios
- Architecting a complete pipeline from data ingestion to classification
- Creating proof-of-concept solutions for surveillance, quality control, or monitoring
- Integrating logging, alerting, and dashboard or API connections
Conclusion and Future Directions
Requirements
- Working knowledge of machine learning principles and model training processes
- Proficiency in Python programming and data preprocessing workflows
- Basic understanding of digital audio fundamentals
Target Audience
- Data scientists
- Machine learning engineers
- Researchers and developers specializing in audio signal processing