2 項進行中

115-1 選課時程

進行中

  • 初選第一階段 6/15 – 6/18
  • 初選第二階段 6/22 – 6/25
  • 校際選修 進行中 8/24 – 9/18
  • 初選第三階段 8/31 – 9/3
  • 開學後加退選 進行中 9/7 – 9/21
  • 逾期加退選 9/21 – 9/24
選課資源

加入行事曆

選擇訂閱 Google Calendar,或下載通用的 ICS 檔案。

使用 Google Calendar 時,Google 會收到這份課表的公開連結。

強化學習專論

Selected Topics in Reinforcement Learning

學期
113-1
學分
0 學分
當期課號
535518
永久課號
CSIC30163
開課單位
資訊科學與工程研究所
授課教師
吳毅成
校區
光復
類別
選修
上課時間表
週二
A
18:30–19:20
強化學習專論
EC114(光復)
3 節連堂
B
19:30–20:20
C
20:30–21:20

* 根據陽明交大上課時間表所列

概述

This course is designed to teach students how to develop reinforcement learning (RL) algorithms for a wide range of applications, such as computer games, video games, intelligent traffic management, manufacturing scheduling, autonomous driving/racing, and robotics. The objectives are summarized as follows. (1) To understand the basic core concepts of reinforcement learning (RL) (2) To understand many latest RL techniques for applications (3) To familiarize with tools for developing RL, such as PyTorch, Gazebo, etc. (4) To develop practical working systems via projects such as DeepRacer.

先修科目

Machine Learning/Deep Learning (suggested)

備註

無備註

教學方式

教師未提供此項資料

評分方式

Projects (done individually) 50% Paper presentation (done in groups of 2 members) 20% Final exam 30%

課程大綱
  • Core of RL

    1. Fundamentals of RL 2. Value Based Reinforcement Learning 3. Policy-based Reinforcement Learning

    講授:
    12
  • Advanced Topics of RL

    1. Applications 2. Exploration vs. Exploitation 3. Planning 4. Advanced Exploration 5. Experience Reply

    講授:
    21
  • Presentation

    State-of-the-art Research Works (TBA)

    講授:
    9
  • Introduction to RL

    1. Introduction to Reinforcement Learning 2. Case studies: Lightweight Model

    講授:
    6
週次計畫
週次主題
第 1 週

Introduction to Reinforcement Learning

2024-09-03(二)
第 2 週

Case studies of lightweight model applications: 2048 and Go

2024-09-10(二)
第 3 週

Fundamentals: Markov Decision Process (MDP), Dynamic Programming (Tabular RL), Q-Learning, Function Approximation

2024-09-17(二)
第 4 週

Value-Based Reinforcement Learning: DQN, DDQN (Double DQN), Dueling Network (with Advantage), Distributional DQN

2024-09-24(二)
第 5 週

Policy-based Reinforcement Learning: Policy Gradient, Actor-Critic (Discrete actions), A2C and A3C (Asynchronous Advantage Actor-Critic)

2024-10-01(二)
第 6 週

Policy-based Reinforcement Learning: TRPO &amp PPO, GAE, DDPG, TD3 (Continuous Actions), SAC (Soft Actor-Critic)

2024-10-08(二)
第 7 週

Applications: DeepRacer: Augmentation, RL-cycleGAN, DrQ

2024-10-15(二)
第 8 週

Applications: Solving Rubik Cube, RL for optimization (JSP/TSP)

2024-10-22(二)
第 9 週

Exploration vs. Exploitation: Multi-Arm Bandits, UCB, Sequential Halving

2024-10-29(二)
第 10 週

Planning: Dyna, Monte-Carlo Tree Search (MCTS), AlphaGo, AlphaZero, MuZero, Path Consistency, Abstraction

2024-11-05(二)
第 11 週

Advanced Exploration: ICM, RND Experience Replay: PER, Ape-X

2024-11-12(二)
第 12 週

Model-based RL: DQfD, R2D3 Multi-Agents RL (MARL) Q-mix, COMA

2024-11-19(二)
第 13 週

Presentation

2024-11-26(二)
第 14 週

Presentation

2024-12-03(二)
第 15 週

Presentation

2024-12-10(二)
第 16 週

Final exam

2024-12-17(二)
第 17 週

Final competition for DeepRacer

2024-12-24(二)
第 18 週

2024-12-31(二)
教科書

1. R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, Nov. 2017 2. David Silver, Online Course for Deep Reinforcement Learning. http://www.cs.ucl.ac.uk/staff/D.Silver/web/Teaching.html 3. Papers and slides.

Office Hours
地點
TBA
時間
TBA
聯絡方式
TBA