In Choosing K you judged clusterings by eye: a good one keeps every point close to its own centroid and keeps the centroids far apart. Two standard scores turn that judgement into numbers, so you can compare groupings of the same data without drawing a picture.
All distances here are ordinary straight-line (Euclidean) distances. Cluster holds points, and its centroid is the average of those points. The centre of all the data, , is the average of every point. There are points and clusters.
Davies–Bouldin score (lower is better)
Each cluster's spread is the average distance from its points to its centroid. For any two different clusters and , their ratio is
This ratio is large when two clusters are wide and close together, which is bad.
Calinski–Harabasz score (higher is better)
measures how far the centroids sit from the middle of the data, with each centroid counted once for every point it holds. is the error left inside the clusters, the same quantity the elbow plot tracks.
Task: write cluster_scores(points, labels). It returns the tuple (davies_bouldin, calinski_harabasz), each rounded to 4 decimal places.
points is a list of points. Each point is a list of numbers, and all points have the same length (2, 3 or more coordinates).labels[i] is the cluster of points[i]. Labels can be any values, numbers or strings, and only the grouping they describe matters.A coffee shop's customers, plotted by visits per week and average spend, can be grouped in many ways. Score two groupings with this function, and the better one usually wins on both numbers: a lower Davies–Bouldin score and a higher Calinski–Harabasz score. The two scores measure the idea differently, so now and then they disagree.