Before trusting any model, compare it with the laziest one possible: always predict the average of the training answers, whatever the row.
Write lazy_baseline(train_answers, test_answers).
train_answers.test_answers.test_answers.Return a dict with:
"rmse": the root mean squared error, rounded to 2 decimal places."r2": the R² score, rounded to 3 decimal places.Use plain Python floats. test_answers always holds at least two different values.