,
Miikka Kuutila
,
Paul Ralph
Creative Commons Attribution 4.0 International license
Background. Code quality metrics are intended to measure latent properties of software source code. Although numerous code metrics have been proposed and used, their construct validity is rarely evaluated. Thus, the extent to which code metrics actually measure what they claim to measure is often unclear. Aim. Drawing from modern measurement theory, we investigate the construct validity of common class-level, object-oriented code quality metrics. Method. As code quality metrics are intended to reflect latent attributes, such as cohesion and coupling, we identified the factor structure of code quality metrics using Exploratory Factor Analysis (EFA). The metrics were extracted from the Apache Maven project by three software tools: Designite, JHawk, and Understand. The factor structure was later verified using Confirmatory Factor Analysis (CFA) on 22 randomly selected open source projects meeting a predetermined eligibility criteria. Results. 24 code quality metrics that correspond to six constructs: Cohesion, In-Coupling, Out-Coupling, Size, Sub-Inheritance (related to subclasses), and Sup-Inheritance (related to superclasses) were revealed in the underlying factor structure. Ten metrics did not correspond to any known dimension of software quality and were removed in the exploratory analysis. Ten additional metrics exhibited low loadings in the confirmatory analysis, suggesting their removal from the final measurement model. Size, Cohesion, Inheritance, and Coupling were the constructs retained, with subcategories identified for Inheritance and Coupling. Conclusions. Our results strongly support the construct validity of 24 code quality metrics. Coupling and Inheritance are revealed as multidimensional constructs, since they require measuring two different concepts, revealed as sub-categories in our analysis, and Complexity may be better explored in a multilevel model. Our results also corroborate the relationship between Cohesion, Size, and Out-Coupling which can be further explored in a structural model. Some metrics from the Chidamber & Kemerer metrics suite are found to perhaps be measuring different constructs than intended. Our results reveal the need for creating or integrating metrics that reflect existing constructs but measure fundamentally different properties to improve the content validity of the measurement model. Additionally, we provide useful recommendations for researchers, developers, and tool providers which stem from theoretical and empirical justification. Overall, our study demonstrates the value of applying modern measurement theory and latent variable modeling in validating software code quality metrics.
@InProceedings{arif_et_al:LIPIcs.ESEM.2026.10,
author = {Arif, Hera and Kuutila, Miikka and Ralph, Paul},
title = {{Assessing the Construct Validity of Object-Oriented, Class-Level Code Quality Metrics}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {10:1--10:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.10},
URN = {urn:nbn:de:0030-drops-279781},
doi = {10.4230/LIPIcs.ESEM.2026.10},
annote = {Keywords: Code quality metrics, Factor analysis, Exploratory factor analysis, Confirmatory factor analysis, Software quality, Size, Inheritance, Coupling, Cohesion}
}