強化學習專論
Selected Topics in Reinforcement Learning
| 節 | 週二 |
|---|---|
A 18:30–19:20 | 強化學習專論 EC114(光復) 3 節連堂 |
B 19:30–20:20 | |
C 20:30–21:20 |
* 根據陽明交大上課時間表所列
This course is designed to teach students how to develop reinforcement learning (RL) algorithms for a wide range of applications, such as computer games, video games, intelligent traffic management, manufacturing scheduling, autonomous driving/racing, and robotics. The objectives are summarized as follows. (1) To understand the basic core concepts of reinforcement learning (RL) (2) To understand many latest RL techniques for applications (3) To familiarize with tools for developing RL, such as PyTorch, Gazebo, etc. (4) To develop practical working systems via projects such as DeepRacer.
Machine Learning/Deep Learning (suggested)
無備註
教師未提供此項資料
Projects (done individually) 50% Paper presentation (done in groups of 2 members) 20% Final exam 30%
Core of RL
1. Fundamentals of RL 2. Value Based Reinforcement Learning 3. Policy-based Reinforcement Learning
- 講授:
- 12
Advanced Topics of RL
1. Applications 2. Exploration vs. Exploitation 3. Planning 4. Advanced Exploration 5. Experience Reply
- 講授:
- 21
Presentation
State-of-the-art Research Works (TBA)
- 講授:
- 9
Introduction to RL
1. Introduction to Reinforcement Learning 2. Case studies: Lightweight Model
- 講授:
- 6
| 週次 | 主題 |
|---|---|
| 第 1 週 | Introduction to Reinforcement Learning 2023-09-12(二) |
| 第 2 週 | Case studies of lightweight model applications: 2048 and Go 2023-09-19(二) |
| 第 3 週 | Fundamentals: Markov Decision Process (MDP), Dynamic Programming (Tabular RL), Q-Learning, Function Approximation 2023-09-26(二) |
| 第 4 週 | Value-Based Reinforcement Learning: DQN, DDQN (Double DQN), Dueling Network (with Advantage), Distributional DQN 2023-10-03(二) |
| 第 5 週 | Policy-based Reinforcement Learning: Policy Gradient, Actor-Critic (Discrete actions), A2C and A3C (Asynchronous Advantage Actor-Critic) 2023-10-10(二) |
| 第 6 週 | Policy-based Reinforcement Learning: TRPO & PPO, GAE, DDPG, TD3 (Continuous Actions), SAC (Soft Actor-Critic) 2023-10-17(二) |
| 第 7 週 | Applications: DeepRacer: Augmentation, RL-cycleGAN, DrQ 2023-10-24(二) |
| 第 8 週 | Applications: Solving Rubik Cube, RL for optimization (JSP/TSP) 2023-10-31(二) |
| 第 9 週 | Exploration vs. Exploitation: Multi-Arm Bandits, UCB, Sequential Halving 2023-11-07(二) |
| 第 10 週 | Planning: Dyna, Monte-Carlo Tree Search (MCTS), AlphaGo, AlphaZero, MuZero, Path Consistency, Abstraction 2023-11-14(二) |
| 第 11 週 | Advanced Exploration: ICM, RND; Experience Replay: PER, Ape-X 2023-11-21(二) |
| 第 12 週 | Model-based RL: DQfD, R2D3; Multi-Agents RL (MARL) Q-mix, COMA 2023-11-28(二) |
| 第 13 週 | Presentation 2023-12-05(二) |
| 第 14 週 | Presentation 2023-12-12(二) |
| 第 15 週 | Presentation 2023-12-19(二) |
| 第 16 週 | Final exam 2023-12-26(二) |
| 第 17 週 | Final competition for DeepRacer 2024-01-02(二) |
| 第 18 週 | 2024-01-09(二) |
1. R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, Nov. 2017 2. David Silver, Online Course for Deep Reinforcement Learning. http://www.cs.ucl.ac.uk/staff/D.Silver/web/Teaching.html 3. Papers and slides.
- 地點
- TBA
- 時間
- TBA
- 聯絡方式
- TBA
