We discuss the results of a comparison between human judgments of visualization readability and an automatic assessment based on empirically-derived design guidelines. The result of this comparison is important for understanding how automated applications of empirically derived guidelines might complement subjective human judgments.
Specifically, we compare scores from Draco 2, a comprehensive test suite that includes an underlying model of visualization guidance, against subjective participant scores from the PREVis Layout readability scale. Our results show that the automatic scores were not a good predictor of human judgment of layout readability.

Whether this discrepancy stems from visual design features that Draco 2 cannot currently represent, from reader-related factors such as familiarity with visualization idioms, or from limitations of the cost model itself remains an open question.
Future work should investigate which aspects of perceived readability are currently absent from formal recommendation models. In particular, extending cost models with fine-tunable, gradual measures of visual density and validating these additions against human perception could help bridge the gap between heuristic-based recommendations and perceived readability.