強化學習專論
Selected Topics in Reinforcement Learning
| 節 | 週二 |
|---|---|
A 18:30–19:20 | 強化學習專論 EC114(光復) 3 節連堂 |
B 19:30–20:20 | |
C 20:30–21:20 |
* 根據陽明交大上課時間表所列
This course is designed to teach students how to develop reinforcement learning (RL) algorithms for a wide range of applications, such as computer games, video games, intelligent traffic management, manufacturing scheduling, autonomous driving/racing, and robotics. The objectives are summarized as follows. (1) To understand the basic core concepts of reinforcement learning (RL) (2) To understand many latest RL techniques for applications (3) To familiarize with tools for developing RL, such as PyTorch, Gazebo, etc. (4) To develop practical working systems via projects such as DeepRacer.
Machine Learning/Deep Learning (suggested)
無備註
教師未提供此項資料
Projects (done individually) 50% Paper presentation (done in groups of 2 members) 20% Final exam 30%
Core of RL
1. Fundamentals of RL 2. Value Based Reinforcement Learning 3. Policy-based Reinforcement Learning
- 講授:
- 12
Advanced Topics of RL
1. Applications 2. Exploration vs. Exploitation 3. Planning 4. Advanced Exploration 5. Experience Reply
- 講授:
- 21
Presentation
State-of-the-art Research Works (TBA)
- 講授:
- 9
Introduction to RL
1. Introduction to Reinforcement Learning 2. Case studies: Lightweight Model
- 講授:
- 6
| 週次 | 主題 |
|---|---|
| 第 1 週 | Introduction to Reinforcement Learning 2025-09-02(二) |
| 第 2 週 | Case studies of lightweight model applications: 2048 and Go 2025-09-09(二) |
| 第 3 週 | Fundamentals: Markov Decision Process (MDP), Dynamic Programming (Tabular RL), Q-Learning, Function Approximation 2025-09-16(二) |
| 第 4 週 | Value-Based Reinforcement Learning: DQN, DDQN (Double DQN), Dueling Network (with Advantage), Distributional DQN 2025-09-23(二) |
| 第 5 週 | Policy-based Reinforcement Learning: Policy Gradient, Actor-Critic (Discrete actions), A2C and A3C (Asynchronous Advantage Actor-Critic) 2025-09-30(二) |
| 第 6 週 | Policy-based Reinforcement Learning: TRPO & amp
 PPO, GAE, DDPG, TD3 (Continuous Actions), SAC (Soft Actor-Critic) 2025-10-07(二) |
| 第 7 週 | Applications: DeepRacer: Augmentation, RL-cycleGAN, DrQ 2025-10-14(二) |
| 第 8 週 | Applications: Solving Rubik Cube, RL for optimization (JSP/TSP) 2025-10-21(二) |
| 第 9 週 | Exploration vs. Exploitation: Multi-Arm Bandits, UCB, Sequential Halving 2025-10-28(二) |
| 第 10 週 | Planning: Dyna, Monte-Carlo Tree Search (MCTS), AlphaGo, AlphaZero, MuZero, Path Consistency, Abstraction 2025-11-04(二) |
| 第 11 週 | Advanced Exploration: ICM, RND
 Experience Replay: PER, Ape-X 2025-11-11(二) |
| 第 12 週 | Model-based RL: DQfD, R2D3
 Multi-Agents RL (MARL) Q-mix, COMA 2025-11-18(二) |
| 第 13 週 | Presentation 2025-11-25(二) |
| 第 14 週 | Presentation 2025-12-02(二) |
| 第 15 週 | Presentation 2025-12-09(二) |
| 第 16 週 | Final exam 2025-12-16(二) |
| 第 17 週 | Final competition for DeepRacer 2025-12-23(二) |
| 第 18 週 | 2025-12-30(二) |
1. R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, Nov. 2017 2. David Silver, Online Course for Deep Reinforcement Learning. http://www.cs.ucl.ac.uk/staff/D.Silver/web/Teaching.html 3. Papers and slides.
- 地點
- TBA
- 時間
- TBA
- 聯絡方式
- TBA
