3.5SpecializationSession 3 · Deep Learning & Applied Machine Learning

Reinforcement learning algorithms

15classes45htotalTH2 · PR1 · PNW2weighting per class

Class-by-class breakdown

15 classes

TH = theory · PR = practical · PNW = personal work. These are ministry weighting codes (not hours) used to split each class's minutes. Class length = course hours ÷ class count; the three time-boxes below always sum to that length.

0%

Quiz progress

0 of 0 classes attempted

First 2 classes free with a referral 14 more with enrollment180 min per classTheory 72mPractical 36mPersonal Work 72m
1

Why Reinforcement Learning Matters for Modern AI Agents

FreeTH2PR1PNW2

An AI Integration Manager who understands RL can explain how today's LLMs were aligned with human feedback — it's the foundation of safe GenAI.

Theory72m
  • RL basics: agents, actions, rewards, environments
  • How RLHF shaped modern LLMs to be helpful and safe
  • Where RL appears in agentic AI: optimization, alignment, evaluation
Practical36m
  • Map one agent behavior to a reward signal (what's being maximized?)
  • Identify where RLHF likely influenced a GenAI product you use
Personal Work72m
  • Write a short note on one RL concept that explains a GenAI behavior
  • Micro-task: submit your note
2

The RL Loop — How Agents Learn From Consequences

LockedTH2PR1PNW2

Locked class preview

A Data/Business Analyst who frames decisions as loops understands agents — each action updates the world and the next choice

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
3

The Reward Hypothesis — Designing Signals for Agent Behavior

LockedTH2PR1PNW2

Locked class preview

A Project Manager translating 'do better' into a number understands RL — every agent goal becomes a reward signal

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
4

Exploration vs Exploitation — The Agent's Core Trade-Off

LockedTH2PR1PNW2

Locked class preview

A Digital Marketing Strategist balancing new audiences vs proven ones faces the explore/exploit dilemma — it's the same math agents use

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
5

RLHF — Reinforcement Learning From Human Feedback

LockedTH2PR1PNW2

Locked class preview

An AI Integration Manager who can explain RLHF demystifies how LLMs became helpful assistants — it's the core alignment technique

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
6

RLAIF — AI Feedback as a Scalable Alternative

LockedTH2PR1PNW2

Locked class preview

A Data/Business Analyst who knows RLAIF understands how AI-generated feedback scales alignment when human labels are expensive

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
7

Reward Modeling — Learning What Humans Prefer

LockedTH2PR1PNW2

Locked class preview

A Project Manager who understands reward models knows alignment quality depends on how well the model captures human preferences

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
8

Policy Optimization — How the Model Actually Improves

LockedTH2PR1PNW2

Locked class preview

An AI Integration Manager who references policy optimization speaks the language of how LLMs are fine-tuned toward better behavior

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
9

Multi-Step Agent Optimization — Rewards Over a Whole Task

LockedTH2PR1PNW2

Locked class preview

A Data/Business Analyst who rewards an agent for end-to-end task success (not just one step) gets agents that solve real problems

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
10

Feedback Loops for Tool-Calling Agents

LockedTH2PR1PNW2

Locked class preview

A Project Manager who builds feedback loops into tool-calling agents lets them improve which tools they pick and how they use them

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
11

Evaluating Autonomous Agent Behavior — Beyond Accuracy

LockedTH2PR1PNW2

Locked class preview

An AI Integration Manager who evaluates agents on safety, cost and task success — not just accuracy — ships agents people can trust

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
12

Safety and Alignment — Preventing Harmful Agent Behavior

LockedTH2PR1PNW2

Locked class preview

A Data/Business Analyst who understands alignment risks knows why unconstrained RL on agents can produce dangerous behavior

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
13

Mini-Project — Design a Reward and Eval for an Agent

LockedTH2PR1PNW2

Locked class preview

This capstone rehearsal proves a Data/Business Analyst can design the reward signal and evaluation that would align a real agent

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
14

Inverse RL and Preference Learning — Inferring What Users Want

LockedTH2PR1PNW2

Locked class preview

A Digital Marketing Strategist who infers user preferences from behavior understands inverse RL — it learns the reward from what people actually choose

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions
15

Present and Reflect — RL for Agentic AI

LockedTH2PR1PNW2

Locked class preview

An AI Integration Manager presenting an RL-for-agents design proves they understand how modern AI is aligned and optimized

Enroll to unlock every class in this course — the full agenda, the AI tutor, and the quiz.

Talk to admissions

Unlock all 15 classes

Get the full agenda, AI tutor, and quiz for every class in 3.5 · Reinforcement learning algorithms.

Apply via admissions

Online checkout not yet live — contact admissions to enroll.