人類感知導向的視覺與多模態表徵學習
Human Perceptual-Centric Vision and Multimodal Representation Learning
| 節 | 週二 |
|---|---|
A 18:30–19:20 | 人類感知導向的視覺與多模態表徵學習 EC115(光復) 3 節連堂 |
B 19:30–20:20 | |
C 20:30–21:20 |
* 根據陽明交大上課時間表所列
本課程以人類感知為核心,探討視覺與多模態表徵學習中的重要問題,內容涵蓋視覺心理學、神經科學啟發模型、感知量測方法、品質評估、可解釋性與多模態對齊。課程兼顧理論基礎、論文閱讀、案例分析與專題實作,強調人類感知與機器表示之間的連結。
This course emphasizes human perceptual-centric vision and multimodal representation learning, and is distinct from conventional computer vision or image processing courses. It highlights perceptual mechanisms, psychophysical methods, perceptual evaluation, and the relationship between multimodal techniques and human perception. There is no strict prerequisite, but students are expected to have basic knowledge of machine learning, linear algebra, probability/statistics, and Python programming. Prior coursework in artificial intelligence, image processing, computer vision, psychology, cognitive science, or neuroscience is beneficial. The course will accommodate students from different backgrounds by providing supplementary guidance when needed. Students should also be comfortable reading English technical materials and research papers. 本課程著重人類感知導向之視覺與多模態表徵學習,與一般電腦視覺或影像處理課程不同,特別強調感知機制、心理物理方法、感知評估,以及多模態技術與人類感知之關聯。課程無硬性先修要求,但建議修課學生具備基本機器學習、線性代數、機率統計與 Python 程式設計能力;曾修習人工智慧、影像處理、電腦視覺、心理學、認知科學或神經科學相關課程者尤佳。授課時將兼顧不同背景學生之學習需求,並適度提供基礎補充與閱讀指引。修課學生亦應具備閱讀英文技術資料與研究論文之能力。
無備註
本課程採用講授、論文導讀、案例分析、課堂討論、實驗實作與專題導向學習並行之方式進行。除教師講解人類視覺感知、心理物理方法、感知品質評估、視覺與多模態表徵學習、可解釋性及人機感知對齊等基礎理論與研究方法外,亦將搭配代表性研究論文、公開資料集、感知品質評估工具、預訓練視覺與多模態模型,以及模型解釋與注意力分析方法,使學生能將課堂理論實際應用於研究問題。 四次作業將依課程進度逐步訓練學生的研究能力,由閱讀與批判既有研究、設計人類感知實驗、評估感知模型,到分析人類與機器感知差異,並作為期末研究專題的能力準備。期末專題則要求學生進一步整合文獻、研究問題、可驗證假說、實驗設計、量化分析與批判性討論,完成一項小型但具完整研究邏輯的 human-centric AI study。 課程網站將提供投影片、教材、作業說明、指定與補充論文、程式與資料資源及相關參考資料,並配置助教協助課程作業、實驗及專題相關問題。課程亦將視進度安排課堂提問、分組討論、論文評論與研究案例分析,以提升學生之主動學習、研究思考與批判分析能力。 生成式人工智慧工具可作為 brainstorming、程式協助、除錯、視覺化及文字表達等輔助工具,但學生仍須自行理解、驗證並能說明所有提交內容,不得以生成內容取代研究判斷,亦不得捏造文獻、實驗結果或人類受試資料。
1.學期作業 本課程共安排四次作業,依序由文獻閱讀、人類感知量測、感知模型評估,逐步延伸至人類與機器感知之比較,培養學生從研究問題形成、實驗設計、量化分析到批判性解釋的完整研究能力。另須完成一項期末研究專題,可由個人或 2–3 人組隊完成。專題應與人類感知及 human-centric AI 相關,形式可包括心理物理或主觀感知研究、既有研究之重現與延伸、人類與機器感知比較,或感知/偏好建模。學生須由相關文獻出發,提出明確研究問題與可驗證假說,透過實作及控制實驗取得量化證據,並進行失效案例與限制分析。期末須完成研究論文、程式與實驗結果,並進行口頭報告與問答。 2.考試狀況 本課程不採傳統筆試。評量以作業、課堂參與與討論,以及期末研究專題為主,強調學生是否能理解研究問題、設計適當的實驗或分析方法、運用量化證據支持結論,並對研究結果、失效案例及限制進行批判性解釋。 3.評量方法 • 出席與課堂討論:20% • 四次作業(HW1–HW4):40% • 期末研究專題:40%
教師未提供此項資料
| 週次 | 主題 |
|---|---|
| 第 1 週 | Week 1: Course Introduction – Human-Centric AI and Multi-modal Perception • Motivation for perceptual modeling in machine learning • Human vs machine perception: differences and conver-gence • Overview of course structure, project expectations • What is a representation? • Why modern AI needs pretraining and multimodal align-ment? 第 1 週:課程介紹-以人為中心的人工智慧與多模態感知 • 機器學習中感知建模的動機 • 人類感知與機器感知:差異與可能的匯流 • 課程架構與專題要求總覽 • 什麼是「表徵」? • 為何現代人工智慧需要預訓練與多模態對齊? |
| 第 2 週 | Week 2: Anatomy of the Human Visual System • Overview of the eye, retina (cones/rods), fovea • Visual pathway: retina to visual cortex (V1, V2, MT) • Spatial receptive fields and basic neurophysiology 第 2 週:人類視覺系統的解剖基礎 • 眼球、視網膜(錐狀細胞/桿狀細胞)、中央凹概述 • 視覺路徑:從視網膜到視覺皮質(V1、V2、MT) • 空間感受野與基本神經生理機制 |
| 第 3 週 | Week 3: Perceptual Foundations of Digital Video Systems • Visual processing as a digital system: from light to pixels to brain • Spatial and temporal contrast sensitivity (CSF) • Frequency response of the eye-brain system and its appli-cations in display and compression 第 3 週:數位影像與視訊系統中的感知基礎 • 將視覺處理視為一個數位系統:從光線到像素再到大腦 • 空間與時間對比敏感度(CSF) • 眼腦系統的頻率響應及其在顯示與壓縮中的應用 |
| 第 4 週 | Week 4: Color Vision and Chromatic Processing • Color spaces: RGB, YUV, YCbCr • Chroma subsampling and Chroma CSF • Cone-opponent channels and retinal processing of color • Applications in compression and display calibration 第 4 週:色彩視覺與色度處理 • 色彩空間:RGB、YUV、YCbCr • 色度子取樣與色度對比敏感函數(Chroma CSF) • 錐體對抗通道與視網膜的色彩處理 • 在壓縮與顯示校正中的應用 |
| 第 5 週 | Week 5: Neural Selectivity and Feature Tuning in Visual Cortex • V1 neuron tuning to spatial frequency, orientation, and mo-tion • Connection to CNN feature maps and filter visualization • Gabor filters and energy models 第 5 週:視覺皮質中的神經選擇性與特徵調諧 • V1 神經元對空間頻率、方向與運動的調諧 • 與 CNN 特徵圖及濾波器視覺化的關聯 • Gabor 濾波器與能量模型 |
| 第 6 週 | Week 6: Computational and Representation Models of Per-ceptual Learning • Ideal observer theory • Bayesian learning models of perception • Self-supervised representation learning • Human learning vs machine pretraining 第 6 週:感知學習的計算模型 • 理想觀察者理論 • 知覺的貝葉斯學習模型 • 自監督表示學習 • 人類學習與機器預訓練之比較 |
| 第 7 週 | Week 7: Psychophysical Methods and Human Perceptual Measurement • Staircase method, 2AFC, JND estimation • ROC curves and Signal Detection Theory (SDT) • Measuring perceptual thresholds and their significance in VQA/IQA 第 7 週:心理物理方法與人類感知量測 • Staircase method(階梯法)、2AFC(兩選強迫選擇)、JND 估計 • ROC 曲線與訊號偵測理論(Signal Detection Theory, SDT) • 感知閾值的量測及其在 VQA/IQA 中的重要性 |
| 第 8 週 | Week 8: Perceptual Metrics and Image/Video Quality Assess-ment • Objective vs subjective evaluation • SSIM, LPIPS, DISTS, VMAF • Incorporating perceptual thresholds and masking in model design 第 8 週:感知指標與影像/視訊品質評估 • 客觀評估與主觀評估 • SSIM、LPIPS、DISTS、VMAF • 在模型設計中納入感知閾值與 masking 效應 |
| 第 9 週 | Week 9: Multimodal Representations and Alignment • From self-supervised learning to multimodal pretraining • Vision-language models (CLIP, BLIP, LLaVA) • Shared embedding spaces and cross-modal alignment • Semantic vs perceptual alignment 第 9 週:多模態表徵與對齊 • 從自監督學習到多模態預訓練 • 視覺-語言模型(CLIP、BLIP、LLaVA) • 共享嵌入空間與跨模態對齊 • 語意對齊與感知對齊 |
| 第 10 週 | Week 10: Multimodal Preference Modeling and Subjectivity • Modeling human preferences in multimodal systems • Bias, fairness, and personalization • User studies in evaluating multimodal output • Prompting, adaptation, and human preference elicitation in multimodal systems 第 10 週:多模態偏好建模與主觀性 • 多模態系統中的人類偏好建模 • 偏差、公平性與個人化 • 用於評估多模態輸出的使用者研究方法 • 多模態系統中的 prompting、模型調適與人類偏好蒐集 |
| 第 11 週 | Week 11: Explainability and Visual Attention in Human and Machine Vision • Eye-tracking and saliency • Grad-CAM, attention rollout, relevance maps • Human-aligned explanations and interpretability 第 11 週:人類與機器視覺中的可解釋性與視覺注意力 • 眼動追蹤與顯著圖 • Grad-CAM、attention rollout、relevance maps • 與人類感知一致的解釋與可解釋性 |
| 第 12 週 | Week 12: Motion Perception and Temporal Processing • Temporal CSF and flicker fusion • Motion blur, frame rate perception • Applications in VR/AR and immersive media 第 12 週:運動知覺與時間處理 • 時間對比敏感函數(Temporal CSF)與閃爍融合 • 動態模糊與幀率知覺 • 在 VR/AR 與沉浸式媒體中的應用 |
| 第 13 週 | Week 13: Comparing Human and Machine Perception • Similarities and differences in representation • Why AI doesn’t "see" like humans yet • Bridging the gap with hybrid perceptual models • Perception, representation, and reasoning: where current multimodal AI still falls short 第 13 週:比較人類與機器感知 • 表示層面的相似與差異 • 為何 AI 仍未真正像人類一樣「看見」 • 以混合式感知模型彌合理論與實作之間的落差 • 感知、表徵與推理之間的落差:當前多模態 AI 的限制 |
| 第 14 週 | Week 14: Case Study: Perception in AR/VR, Digital Human Avatars • Motion, depth, stereo, and embodiment • Perceptual quality and realism in immersive media • Multimodal interaction and human preference in AR/VR systems • Case studies on digital humans, avatars, and perceptual alignment 第 14 週:案例研究:AR/VR 與數位人類虛擬替身中的感知議題 • 運動、深度、立體視覺與具身感 • 沉浸式媒體中的感知品質與真實感 • AR/VR 系統中的多模態互動與人類偏好 • 數位人類、虛擬替身與感知對齊之案例分析 |
| 第 15 週 | Week 15–16: Final Project Presentations • Student-led presentations on perception-based modeling projects • Peer feedback and synthesis discussion 第 15-16 週:期末專題報告 • 由學生主導的感知導向建模專題報告 • 同儕回饋與整體綜合討論 |
| 第 16 週 | Week 15–16: Final Project Presentations • Student-led presentations on perception-based modeling projects • Peer feedback and synthesis discussion 第 15-16 週:期末專題報告 • 由學生主導的感知導向建模專題報告 • 同儕回饋與整體綜合討論 |
1. 人類感知基礎:Goldstein & Cacciamani, Sensation and Perception, 11th Edition, Cengage Learning, 2021 2. 實驗方法與 threshold/JND:Kingdom & Prins, Psychophysics: A Practical Introduction (2nd ed.), Academic Press, 2016 3. 機器視覺與 representation 背景:Richard Szeliski, Computer Vision: Algorithms and Applications (2nd ed.), Springer Cham, 2022
- 地點
- EC241B
- 時間
- Tuesday 4:20 – 5:20 pm
- 聯絡方式
- Email: berriechen@nycu.edu.tw
