ITcon Vol. 31, pg. 1067-1095, http://www.itcon.org/2026/45

Toward trustworthy construction safety hazard detection with visual-language models

DOI:10.36680/j.itcon.2026.045
submitted:May 2026
published:September 2026
editor(s):Bosché F
authors:Mina Sadat Orooje, PhD candidate
Department of Architecture, Construction Engineering and Built Environment, Politecnico di Milano, Italy
https://orcid.org/0009-0002-4131-982X
minasadat.orooje@polimi.it

Fulvio Re Cecconi, Associate Professor
Department of Architecture, Construction Engineering and Built Environment, Politecnico di Milano, Italy
https://orcid.org/0000-0001-7716-8854
fulvio.rececconi@polimi.it

Gaang Lee, Dr.-Ing., Assistant Professor
Department of Civil and Environmental Engineering, University of Alberta, Edmonton, Alberta, Canada
https://orcid.org/0000-0002-6341-2585
gaang@ualberta.ca

Qipei (Gavin) Mei, PhD, PEng, Assistant Professor
Department of Civil and Environmental Engineering, University of Alberta, Edmonton, Alberta, Canada
https://orcid.org/0000-0003-1409-3562
qipei@ualberta.ca

Muhammad Adil, M.Sc., Research Assistant
Department of Civil and Environmental Engineering, University of Alberta, Canada
https://orcid.org/0009-0009-1346-6479
madil2@ualberta.ca
summary:The construction industry continues to experience high rates of accidents and fatalities, underscoring the need for proactive and reliable safety management. Accurate hazard identification is essential for effective image-based monitoring and decision support on construction sites. Vision–language models (VLMs) have shown strong potential for interpreting complex visual environments; however, their deployment in safety-critical applications is limited by hallucination, where hazards are inferred without sufficient evidence. To address this limitation, this study proposes a retrieval-augmented hazard detection framework that grounds VLM outputs in visual evidence through image-to-image retrieval of visually similar hazard scenarios and structured domain knowledge within a retrieval-augmented generation (RAG) pipeline. The framework integrates image-to-image retrieval of real-world hazard scenarios, supported by a curated knowledge base of 1,040 annotated construction site images, structured hazard taxonomies, and regulation-aware verification to enforce condition-constrained hazard reasoning before compliance assessment. The proposed approach is evaluated on a separate set of 260 construction site images spanning ten construction-safety hazard categories aligned with OSHA safety domains and the construction-safety literature. At the end-to-end report level, the complete pipeline achieves an F1-score of 0.917 (precision 0.939, recall 0.896) and a hallucination rate of 6.1%, substantially lower than the VLM-only baseline. At the hazard-verification gate, adding taxonomy-constrained reasoning to RAG-1 yields a precision of 0.987 and an F1-score of 0.941, while RAG-1 alone achieves an F1-score of 0.923 with 89.8% retrieval coverage. These results indicate that visual grounding provides the largest reliability gain, with structured condition verification adding further precision and gated regulatory retrieval providing traceable compliance support after hazard confirmation.
keywords:Vision–Language Models (VLMs), Retrieval-Augmented Generation (RAG), Construction Safety, Hallucination Mitigation, Hazard Detection
full text: (PDF file, 1.241 MB)
citation:Orooje, M. S., Re Cecconi, F., Lee, G., Mei, Q., & Adil, M. (2026). Toward trustworthy construction safety hazard detection with visual-language models. Journal of Information Technology in Construction (ITcon), 31, 1067-1095. https://doi.org/10.36680/j.itcon.2026.045
statistics: