強化學習原理
Reinforcement Learning
| 節 | 週一 | 週四 |
|---|---|---|
3 10:10–11:00 | 強化學習原理 ED202(光復) 2 節連堂 | |
4 11:10–12:00 | ||
7 15:30–16:20 | 強化學習原理 ED202(光復) |
* 根據陽明交大上課時間表所列
- Learn how to model tasks as RL problems. - Understand RL from a theoretical viewpoint - Learn how to systematically solve RL problems by using various RL algorithms and perform analysis of these algorithms - Learn how to implement deep RL algorithms using software packages (e.g. Tensorflow and Pytorch) through team projects
- Some math maturity: Familiarity with calculus and probability (basic understanding of numerical optimization would help) - Programming language: Python (familiarity with Tensorflow/Pytorch would help)
無備註
教師未提供此項資料
Homework: 35% Theory Project: 30% Team Implementation Project: 35% (including 10% for presentation)
教師未提供此項資料
| 週次 | 主題 |
|---|---|
| 第 1 週 | - Course Logistics - Markov Decision Process (MDP) 2024-02-19(一),2024-02-22(四) |
| 第 2 週 | - Planning in MDPs - Bellman Equations - Value Iteration - Policy Iteration - Regularized MDPs 2024-02-26(一),2024-02-29(四) |
| 第 3 週 | - Policy Optimization - Introduction to Optimization (Convexity, Smoothness, Gradient Descent, and Mirror Descent) - Policy Gradient (PG) 2024-03-04(一),2024-03-07(四) |
| 第 4 週 | - Stochastic PG (REINFORCE, A2C, and Natural PG) - Variance Reduction 2024-03-11(一),2024-03-14(四) |
| 第 5 週 | - Model-Free Prediction - Generalized Advantage Estimation 2024-03-18(一),2024-03-21(四) |
| 第 6 週 | - Global Convergence of Policy Gradient - Global Convergence of Natural PG 2024-03-25(一),2024-03-28(四) |
| 第 7 週 | - Value Function Approximation 2024-04-01(一),2024-04-04(四) |
| 第 8 週 | - Deterministic PG, DDPG, TD3 - Off-Policy Learning via Deterministic and Stochastic Policy Gradients 2024-04-08(一),2024-04-11(四) |
| 第 9 週 | - Trust Region Policy Optimization (TRPO) - Global Convergence of TRPO - Proximal Policy Optimization (PPO) 2024-04-15(一),2024-04-18(四) |
| 第 10 週 | - Value-Based Methods and Stochastic Approximation - Sarsa, Expected Sarsa, Q-Learning, and Double Q-Learning 2024-04-22(一),2024-04-25(四) |
| 第 11 週 | - Distributional Perspective of MDPs - Distributional RL (C51, QR-DQN, and IQN) 2024-04-29(一),2024-05-02(四) |
| 第 12 週 | - Entropy-Regularized RL - Soft Q-learning - Soft Actor-Critic 2024-05-06(一),2024-05-09(四) |
| 第 13 週 | - Reinforcement Learning from Human Feedback (RLHF) - Recent Theoretical Results on RLHF - Dueling Bandits 2024-05-13(一),2024-05-16(四) |
| 第 14 週 | - Imitation Learning - Inverse Reinforcement Learning (GAIL, WAIL, AIL, and IQ-Learn) 2024-05-20(一),2024-05-23(四) |
| 第 15 週 | - Upside-Down RL - Sequence-to-Sequence Modeling for RL 2024-05-27(一),2024-05-30(四) |
| 第 16 週 | - No class (exam week) 2024-06-03(一),2024-06-06(四) |
| 第 17 週 | - Final Presentations 2024-06-10(一),2024-06-13(四) |
| 第 18 週 | 2024-06-17(一),2024-06-20(四) |
- Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, MIT Press, 2nd edition, 2018 - Alekh Agarwal, Nan Jiang, and Sham M. Kakade, Reinforcement Learning: Theory and Algorithms, 2020 - Nocedal, Jorge, and Stephen Wright. Numerical optimization. Springer Science & Business Media, 2006 - Léon Bottou, Frank E. Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning. arXiv 2016 - Tor Lattimore and Csaba Szepesvari, Bandit Algorithms. 2019
- 地點
- EC418
- 時間
- 1pm-1:30pm on Mondays
- 聯絡方式
- By email: pinghsieh@nycu.edu.tw
