A voice-command model for a banking app recognises which of 20 commands a caller has spoken. Before launch it was evaluated on labelled recordings and reported one headline figure: 93% accuracy.
Later the team broke that same test set down by where the speakers were from:
| group | recordings in the test set | accuracy |
|---|---|---|
| Region A | ||
| Region B | not reported |
The app then launches properly in Region B. Nobody labels production recordings, but the input monitoring shows that callers from Region B now make up 40% of traffic, with Region A the other 60%. The model has not been retrained, and within each region it is exactly as accurate as it was on the test set.
What overall accuracy should the team now expect in production?
Give the answer as a percentage, rounded to 1 decimal place.