深度視覺運算
Deep Visual Computing
| 節 | 週四 |
|---|---|
2 09:00–09:50 | 深度視覺運算 A212(光復) 3 節連堂 |
3 10:10–11:00 | |
4 11:10–12:00 |
* 根據陽明交大上課時間表所列
Visual computing transitioned from traditional image processing that relied on manual feature engineering to representation learning driven by deep neural networks. This transition has allowed machines to achieve unprecedented accuracy in visual recognition tasks, while simultaneously introducing challenges regarding computational efficiency, model interpretability, and multimodal reasoning. This course bridges foundational deep learning algorithms with forward-looking real-world deployments such as autonomous systems and digital healthcare. The primary objective of this course is understanding deep learning across diverse core vision tasks, including image classification, semantic segmentation, image retrieval, and generative synthesis. We will explore limitations such as the heavy reliance on massive annotated datasets, bias amplification, and the inherent "black box" opacity of neural network architectures. Furthermore, introducing the frontier of AI, covering the integration of Large Language Models (LLMs) and Vision-Language Models (VLMs). The synthesis of visual and textual modalities enables systems to transition from mere pattern recognition to complex compositional reasoning, spatial logic, and agentic interactions. To ensure practical viability, the course examines edge computing optimization such as model pruning for deploying highly parameterized models onto resource-constrained hardware.
Familiar with computer operation, programming, and machine learning/deep learning
無備註
Lectures: Systematic explanation of core concepts Paper Discussions: Discussion of selected key papers Hands-on Practice: Practical assignments and projects
Homework: 30%, Midterm: 30%, Final: 40%
教師未提供此項資料
| 週次 | 主題 |
|---|---|
| 第 1 週 | Fundamentals of Computer Vision & Machine Learning 時數:[2026-09-10]羅崇銘(3.00) |
| 第 2 週 | Image Classification Architectures and Evolution 時數:[2026-09-17]羅崇銘(3.00) |
| 第 3 週 | Advanced Classification & Attention Mechanisms 時數:[2026-09-24]羅崇銘(3.00) |
| 第 4 週 | Semantic Image Segmentation 時數:[2026-10-01]羅崇銘(3.00) |
| 第 5 週 | Object Detection and Localization 時數:[2026-10-08]羅崇銘(3.00) |
| 第 6 週 | Image Retrieval and Metric Learning 時數:[2026-10-15]羅崇銘(3.00) |
| 第 7 週 | Paper Discussions 1 時數:[2026-10-22]羅崇銘(3.00) |
| 第 8 週 | Paper Discussions 2 時數:[2026-10-29]羅崇銘(3.00) |
| 第 9 週 | Generative Models: Latent Spaces & Adversarial Networks 時數:[2026-11-05]羅崇銘(3.00) |
| 第 10 週 | Vision-Language Models (VLMs) & Contrastive Pre-training 時數:[2026-11-12]羅崇銘(3.00) |
| 第 11 週 | Multimodal Large Language Models (MLLMs) 時數:[2026-11-19]羅崇銘(3.00) |
| 第 12 週 | Edge Computing, Model Compression, & Quantization 時數:[2026-11-26]羅崇銘(3.00) |
| 第 13 週 | Medical Image Analysis: Diagnostic Modalities 時數:[2026-12-03]羅崇銘(3.00) |
| 第 14 週 | Robotic Vision and Visual Servoing 時數:[2026-12-10]羅崇銘(3.00) |
| 第 15 週 | Final Project Report 1 時數:[2026-12-17]羅崇銘(3.00) |
| 第 16 週 | Final Project Report 2 時數:[2026-12-24]羅崇銘(3.00) |
1. Richard Szeliski, Computer Vision: Algorithms and Applications 2nd Edition, Springer, 2022 2. Simon J.D. Prince, Computer vision: models, learning and inference, Cambridge University Press, 2012 3. Peter Corke, Robotics, Vision and Control: Fundamental Algorithms in Python (3rd Edition), Springer Nature Switzerland AG, 2023 4. François Fleuret, The Little Book of Deep Learning (Version 1.2), Université de Genève, 2024
- 地點
- Classroom
- 時間
- After class
- 聯絡方式
