Skip to main content

← Back to Reinforcement learning

Reinforcement learning quiz

All 11 questions. After you submit, each question shows the right answer and why. Back to the quick check

ExplorerQuestion 1 of 11

A game-playing AI keeps using a move that usually wins a few points, but never tries other moves that might win more. Which balance is it getting wrong?

ExplorerQuestion 2 of 11

In reinforcement learning, nobody tells the computer the right move: it has to try moves and find out which ones earn rewards.

InvestigatorQuestion 3 of 11

What two features do Sutton and Barto call the most important distinguishing features of reinforcement learning?

InvestigatorQuestion 4 of 11

With discount rate γ = 0.9, an agent will receive a reward of 1 on each of the next three steps. What is the discounted return?

ScientistQuestion 5 of 11

Q-learning update: Q(s, a) = 0, step size α = 0.5, reward 1, γ = 0.9, and the best action in the next state has Q = 2. What is the new Q(s, a)?

ScientistQuestion 6 of 11

What inputs did the 2015 deep Q-network receive when learning to play Atari games?

AdvancedQuestion 7 of 11

On how many of the 49 Atari games did the 2015 deep Q-network achieve more than 75% of the professional human tester's score?

AdvancedQuestion 8 of 11

AlphaGo (2016) learned purely from self-play, without using any human expert games.

AdvancedQuestion 9 of 11

What is meant by the 'reward hypothesis'?

ExpertQuestion 10 of 11

In InstructGPT's training, what provided the signal for the reinforcement-learning step?

ExpertQuestion 11 of 11

Name the two AI-safety problems that Amodei et al. (2016) attribute to having the wrong objective function, and give an example of how one could arise in an RL agent.