2 項進行中

115-1 選課時程

進行中

  • 初選第一階段 6/15 – 6/18
  • 初選第二階段 6/22 – 6/25
  • 校際選修 進行中 8/24 – 9/18
  • 初選第三階段 8/31 – 9/3
  • 開學後加退選 進行中 9/7 – 9/21
  • 逾期加退選 9/21 – 9/24
選課資源

加入行事曆

選擇訂閱 Google Calendar,或下載通用的 ICS 檔案。

使用 Google Calendar 時,Google 會收到這份課表的公開連結。

強化學習原理

Reinforcement Learning

學期
111-2
學分
0 學分
當期課號
535515
永久課號
CSIC30046
開課單位
資訊科學與工程研究所
授課教師
謝秉均
校區
光復
類別
選修
上課時間表
週二
週五
3
10:10–11:00
強化學習原理
EC115(光復)
2 節連堂
4
11:10–12:00
7
15:30–16:20
強化學習原理
EC115(光復)

* 根據陽明交大上課時間表所列

概述

- Learn how to model tasks as RL problems. - Understand RL from a theoretical viewpoint - Learn how to systematically solve RL problems by using various RL algorithms and perform analysis of these algorithms - Learn how to implement deep RL algorithms using software packages (e.g. Tensorflow and Pytorch) through team project

先修科目

- Some math maturity: Familiarity with calculus and probability (basic understanding of optimization would help) - Programming language: Python (familiarity with Tensorflow/Pytorch would help)

備註

無備註

教學方式

**Important** The first lecture on 2/14 (Tuesday) will be delivered via Webex at https://nycu.webex.com/nycu/j.php?MTID=m4c8d72b8f0c204e24785b52c2d9511cf

評分方式

Homework: 30% Theory Project: 30% Team Implementation Project: 40% (including 10% for presentation)

課程大綱

教師未提供此項資料

週次計畫
週次主題
第 1 週

We will try our best to discuss all of the following topics:1. Markov decision process (MDP) and planning in MDPs2. Distributional Perspective of MDPs3. Policy Optimization and Gradient Descent4. Stochastic Policy Gradient Methods (REINFORCE, A2C, NPG)5. Variance Reduction and Model-Free Prediction6. Global Convergence of Policy Gradient7. Value Function Approximation8. Trust Region Policy Optimization: TRPO, PPO, and CPO9. Deterministic Policy Gradient for Continuous Control10. Off-Policy Learning via Deterministic and Stochastic Policy Gradients11. Value-Based Methods and Stochastic Approximation (Expected Sarsa, Q-Learning, and Double Q-Learning)12. Distributional RL (C51, QR-DQN, and IQN)13. Soft Actor Critic14. Imitation Learning and Inverse Reinforcement Learning

2023-02-14(二),2023-02-17(五)
第 2 週

2023-02-21(二),2023-02-24(五)
第 3 週

2023-02-28(二),2023-03-03(五)
第 4 週

2023-03-07(二),2023-03-10(五)
第 5 週

2023-03-14(二),2023-03-17(五)
第 6 週

2023-03-21(二),2023-03-24(五)
第 7 週

2023-03-28(二),2023-03-31(五)
第 8 週

2023-04-04(二),2023-04-07(五)
第 9 週

2023-04-11(二),2023-04-14(五)
第 10 週

2023-04-18(二),2023-04-21(五)
第 11 週

2023-04-25(二),2023-04-28(五)
第 12 週

2023-05-02(二),2023-05-05(五)
第 13 週

2023-05-09(二),2023-05-12(五)
第 14 週

2023-05-16(二),2023-05-19(五)
第 15 週

2023-05-23(二),2023-05-26(五)
第 16 週

2023-05-30(二),2023-06-02(五)
第 17 週

2023-06-06(二),2023-06-09(五)
第 18 週

2023-06-13(二),2023-06-16(五)
教科書

Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, MIT Press, 2nd edition, 2018 Alekh Agarwal, Nan Jiang, and Sham M. Kakade, Reinforcement Learning: Theory and Algorithms, 2020 Nocedal, Jorge, and Stephen Wright. Numerical optimization. Springer Science & Business Media, 2006 Léon Bottou, Frank E. Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning. arXiv 2016 Tor Lattimore and Csaba Szepesvari, Bandit Algorithms. 2019

Office Hours
地點
EC418
時間
4:30pm-5pm on Tuesdays
聯絡方式
By email: pinghsieh@nycu.edu.tw