Leveraging large language models for synthetic data generation for AI applications for water infrastructure management

Wala, Jens and Rüppel, Uwe; Moreno-Rangel, Alejandro and Kumar, Bimal, eds. (2025) Leveraging large language models for synthetic data generation for AI applications for water infrastructure management. In: EG-ICE 2025. University of Strathclyde Publishing, GBR, pp. 635-644. ISBN 9781914241826 (https://doi.org/10.17868/strath.00093226)

[thumbnail of Wala-Ruppel-EG-ICE-2025-Leveraging-large-language-models-for-synthetic-data-generation]
Preview
Text. Filename: Wala-Ruppel-EG-ICE-2025-Leveraging-large-language-models-for-synthetic-data-generation.pdf
Final Published Version
License: Creative Commons Attribution 4.0 logo

Download (4MB)| Preview

Abstract

The effectiveness of AI models in water infrastructure management is often hindered by limited access to high-quality, domain-specific data. This paper explores the use of Large Language Models (LLMs), specifically the GPT-4 family of models, to generate synthetic hydrological datasets that augment real-world data and improve AI model performance. Leveraging prompt engineering tailored to hydrological patterns, we generate synthetic data that maintains key statistical properties of real observations. Our results demonstrate that AI models trained with synthetic data exhibit improved accuracy in forecasting tasks. This approach offers a scalable, privacy-preserving solution to data scarcity in critical infrastructure contexts.