Question: Exercise 11.6 Explain how Q-learning fits in with the agent architecture of Section 2.2.1 (page 46). Suppose that the Q-learning agent has discount factor ,
Exercise 11.6 Explain how Q-learning fits in with the agent architecture of Section 2.2.1 (page 46). Suppose that the Q-learning agent has discount factor γ, a step size of α, and is carrying out an -greedy exploration strategy.
(a) What are the components of the belief state of the Q-learning agent?
(b) What are the percepts?
(c) What is the command function of the Q-learning agent?
(d) What is the belief-state transition function of the Q-learning agent?
Step by Step Solution
There are 3 Steps involved in it
Get step-by-step solutions from verified subject matter experts
