Abstract:Current quality risk assessment for residential construction projects relies largely on practitioners’ experience. To improve assessment accuracy and strengthen its empirical basis, this study uses approximately 142,000 quality-risk events recorded during third-party inspections as the data source and conducts text cleaning, domain-specific word segmentation, word-frequency analysis, and association-rule mining. The results show that major-risk events account for 1.71% of the sample. The terms “roof,” “leakage,” and “external wall” occur 10,640, 10,429, and 10,047 times, respectively, identifying the most prominent high-frequency risk locations and problems. Defects such as leakage, cracking, exposed reinforcement, and hollowing are frequent and widely distributed. The Apriori algorithm identifies several strongly directional co-occurrence rules: the itemset containing “design” and “spacing” is associated with that containing “greater than” and “stirrup”; “waterproofing membrane” and “height” are associated with “upturn”; and “external wall” and “sealing” are associated with “form-tie hole.” Compared with qualitative methods such as expert scoring and comprehensive assessment, the proposed approach transforms unstructured descriptions of quality problems into quantified defect frequencies and priority locations, while revealing quantitative relationships among locations, defects, and construction processes. Differentiated inspection priorities are further proposed for structural-safety, serviceability, and installation-and-connection problems. The findings provide a quantitative basis for quality risk assessment, inspection prioritization, and targeted leakage-prevention inspections in residential construction projects, while establishing a data foundation for the quantitative characterization of quality-risk indicators.