Q-learning updates toward the best action available in the next state — even if the agent has no intention of taking it. SARSA updates toward the action the agent actually took:
The name is the five things in that formula: state, action, reward, next state, next action.
Task: write sarsa_update(q, state, action, reward, next_state, next_action, alpha, gamma) returning the full updated Q-table, every entry rounded to 4 decimal places.
q is a list of rows, one per state, each holding that state's action values: q[s][a].alpha is the step size, gamma the discount.(state, action) and (next_state, next_action) may be the same cell. Read the target value before you overwrite anything.Swapping a max for a plain lookup sounds trivial and changes the agent's character completely. Q-learning learns the value of the optimal policy regardless of how recklessly it explores; SARSA learns the value of the policy it is actually following, exploration mistakes included. On a cliff-edge task, Q-learning walks the optimal route along the edge and occasionally falls off, while SARSA — which has felt those falls in its own updates — learns a slower, safer path.