Publications

paper
Enhancing firefighter safety and efficiency through UAV-assisted AI-based human motion recognition system
Wang Hong, Chenyang Zhao, Yuan Feng, Xu Huang, Chenyi Qu, Yaguang Zhu, Ming Xin, and Wenbin Guo.
Expert Systems with Applications, 2025
Fires pose grave threats to lives, properties, and natural resources, necessitating effective fire suppression and rescue efforts. This paper proposes a solution to enhance firefighter safety and efficiency by integrating Unmanned Aerial Vehicles (UAVs) with Artificial Intelligence (AI)-based human motion recognition. The proposed system employs supervised machine learning algorithms trained on firefighting video datasets to recognize six crucial firefighter actions: ‘advance’, ‘climb’, ‘connect’, ‘fetchwater’, ‘hoselaying’, and ‘watersupply’. The trained model demonstrates 74.8% accuracy in identifying these actions during fire suppression activities. By utilizing UAVs and AI, this system offers extended detection range, deployment flexibility, and live, high-resolution image capture without risking pilots’ lives. The integration of UAVs and AI holds great promise in mitigating the devastating impact of fires and safeguarding the lives of firefighters.
paper
Enhancing Human Detection in Post-Disaster Scenarios through Generative Adversarial Networks
Chengyi Qu, Xiaodong Hao, Feiyu Ma, Jizhuang Hui , Wenbin Guo, and Hong Wang.
InInternational Conference on Human-Computer Interaction, 2025
Human body detection is critical in post-disaster response, especially in UAV-assisted search and rescue missions, where accurate localization of stranded individuals is essential for efficient emergency operations. However, current disaster scene image datasets are often scarce, poorly annotated, and of inconsistent quality, limiting the performance of deep learning models in complex disaster environments (e.g., ruins, fires, floods). These challenges are exacerbated when human targets are intertwined with complex backgrounds, making existing techniques inadequate for practical needs. To address these limitations, this paper presents an integrated approach combining image generation and target detection techniques to enhance the adaptability and accuracy of deep learning models. By constructing diverse, high-quality disaster scene datasets, the proposed method alleviates data scarcity and improves generalization under complex scenarios. Models trained on the expanded dataset achieve significant improvement, with generated images yielding an FID score of 45 and mAP50 scores exceeding 98% across three detection models. Additionally, this study explores the integration of computer vision and human-computer interaction in UAV search and rescue missions. The proposed framework provides an intelligent, efficient solution for post-disaster search and rescue, demonstrating the transformative potential of data augmentation and target detection technologies in disaster management. This work establishes a robust technological foundation for future post-disaster response and intelligent search and rescue operations.
paper
Autonomous Video Transmission and Air-to-Ground Coordination in UAV-Swarm-Aided Disaster Response Platform
Chengyi Qu, Xiaodong Hao, Feiyu Ma, Jizhuang Hui, Wenbin Guo, and Hong Wang.
InInternational Conference on Human-Computer Interaction, 2025
Unmanned Aerial Vehicles (UAVs), or drones with cameras, are crucial for environmental awareness in applications like smart agriculture, border security, and disaster response. Building realistic UAV testbeds for novel network control algorithms is challenging due to time constraints and regulatory limitations. To develop an autonomous air-to-ground coordination platform, in this paper, we introduce a UAV-swarm-aided Disaster Response Platform (DRP) that simulates video transmission and coordination that integrates simulation for both drones and networks, allowing experimentation with protocols (HTTP/TCP, UDP/RTP, QUIC) and video properties (codec, resolution). Our design combined human-centered interaction principles, prioritizing user experience and indicating multi-modal interactions. In addition, our approach utilizes trace-based experiments to demonstrate its effectiveness in delivering video quality (e.g., PSNR) aligning with real-world measurements, while in the meantime validating model accuracy. Evaluation results demonstrate that our intelligent video transmission and air-to-ground coordination strategies present reasonable video quality under measurement of subjective and objective metrics, and in the meantime achieve \approx 85\% accuracy with dynamic decision-making with competitive time efficiency and energy savings (\approx 18\% gain) compared to on-boarding only. Implementing these strategies in UAV-swarm-aided DRPs significantly enhances overall response efficiency, ensuring the safety of lives and property.
paper
Emotion Recognition in Dance: A Novel Approach Using Laban Movement Analysis and Artificial Intelligence
Hong Wang, Chenyang Zhao, Xu Huang, Yaguang Zhu, Chengyi Qu, and Wenbin Guo.
InInternational Conference on Human-Computer Interaction, 2024
Dance, as a highly expressive form of art, conveys intense emotions through bodily movements and postures. In the field of human-computer interaction, the automated recognition of dance movements poses a significant challenge concerning artistic expression and emotional classification. Analyzing dance movements enables us to extract rich emotional information. This paper introduces a novel approach for dance emotion recognition—the Laban Movement Analysis (LMA)—which characterizes the human body based on three aspects: body distribution, body structure, and dynamic trends. Leveraging artificial intelligence-based computer vision technology, we conduct a comparative analysis and supervised learning on existing dance performance video datasets. Various machine learning algorithms are trained and compared. The results indicate that recognizing emotional information from the perspective of dance movements achieves a high level of accuracy.
paper
An AI-Based Action Detection UAV System to Improve Firefighter Safety
Hong Wang, Yuan Feng, Xu Huang, and Wenbin Guo.
InInternational Conference on Human-Computer Interaction, 2023
Human hazardous fires can inflict massive harm to life, property, and the environment. Close contact with fire sources threatens firefighters’ lives who are critical first responders to fire suppression and rescue. Unmanned aerial vehicles (UAVs) are recently introduced to improve firefighters’ performance by monitoring fire characteristics and inferring trajectory. Some UAVs studies used artificial intelligence (AI) for pattern recognition for wildfire prevention and other relevant tasks. However, how to coordinate firefighters is equally important and needs to be explored. Therefore, this research offers an analysis and comparison of AI-based computer vision methods that can annotate human movement automatically. This study proposes to use UAVs combined with AI-based human motion detection to recognize firefighters’ action patterns. Based on existing human motion datasets of firefighting videos, we have successfully trained a supervised machine learning algorithm recognizing firefighters’ forward movement and water splashing actions. The trained model was tested in the existing video and reached an 81.55% accuracy rate. Applying this model in the UAV system is able to improve firefighters’ functioning and safeguard their lives.
paper
An autonomous robot for shell and tube heat exchanger inspection
Zheng, Bujingda, Jheng‐Wun Su, Yunchao Xie, Jonathan Miles,Hong Wang, Wenxin Gao, Ming Xin, and Jian Lin.
Journal of Field Robotics, 2022
Shell and tube heat exchangers (STHEs) are critical to energy conversion efficiency of power plants. Eddy current examination is a way to evaluate working conditions of these tubes. However, the current testing apparatus requires human to manually insert an eddy current testing (ECT) probe into and extract it out of individual tubes, and meanwhile monitor measurement results for diagnosis. It is a time-consuming and labor-intensive procedure even for an experienced technician. To tackle this challenge, in this study, we developed a robot enabled ECT system for autonomous inspection of STHEs. The robotic platform employs Mecanum wheeled chassis for high mobility, machine vision to locate tube bundle and tube inlets, a rotational Cartesian mechanism to operate at planes with all possible inclinations, and a task-specific mechanism for ECT probe delivery. Machine vision locates tube bundle and tube inlets by an April tag detection algorithm and a Circle Hough Transform algorithm, respectively. Assisted by a guiding cone, the ECT probe is continuously fed into the tubes with a fill factor of 0.819. During this process, the eddy current data are automatically collected and real-time analyzed by convolutional neural networks, showing accuracy of nearly 100% for identifying defective and nondefective tubes and 85% for four types of defective tubes and nondefective tubes.
paper
Inverse design of two-dimensional graphene/h-BN hybrids by a regressional and conditional GAN
Dong Yuan, Dawei Li, Chi Zhang, Chuhan Wu,Hong Wang, Ming Xin, Jianlin Cheng, and Jian Lin.
Carbon, 2020
Design of materials with desired properties is currently laborious and heavily relies on intuition of researchers through a trial-and-error process. To tackle this challenge, we propose a novel regressional and conditional generative adversarial network (RCGAN) for inverse design of representative two-dimensional materials, the graphene and boron-nitride (BN) hybrids. RCGAN incorporates a supervised regressor network, thus overcoming the common technical barrier in the traditional unsupervised GANs, which cannot generate data when fed with continuous and quantitative labels. RCGAN can autonomously generate graphene/BN hybrids given any target bandgap values. These structures are distinguished from the ones used for training and exhibit high diversity for a given bandgap. Moreover, they exhibit high fidelity, yielding bandgaps within ∼10% MAEF of the desired bandgaps as validated by density functional theory (DFT) calculations. Analysis by the principle component analysis (PCA) and modified locally linear embedding (MLLE) reveals that the generator has successfully generated structures following the statistical distribution of the real structures. It implies the possibility of the RCGAN in recognizing physical rules hidden in the high-dimensional data. The novel strategy for designing regressional GAN architecture together with the successful application to inverse design of materials would inspire further exploration in research fields beyond materials.
paper
Rapid identification of X-ray diffraction patterns based on very limited data by interpretable convolutional neural networks
Hong Wang, Yunchao Xie, Dawei Li, Heng Deng, Yunxin Zhao, Ming Xin, and Jian Lin.
Journal of chemical information and modeling, 2020
Large volumes of data from material characterizations call for rapid and automatic data analysis to accelerate materials discovery. Herein, we report a convolutional neural network (CNN) that was trained based on theoretical data and very limited experimental data for fast identification of experimental X-ray diffraction (XRD) patterns of metal–organic frameworks (MOFs). To augment the data for training the model, noise was extracted from experimental data and shuffled; then it was merged with the main peaks that were extracted from theoretical spectra to synthesize new spectra. For the first time, one-to-one material identification was achieved. Theoretical MOFs patterns (1012) were augmented to a whole data set of 72 864 samples. It was then randomly shuffled and split into training (58 292 samples) and validation (14 572 samples) data sets at a ratio of 4:1. For the task of discriminating, the optimized model showed the highest identification accuracy of 96.7% for the top 5 ranking on a test data set of 30 hold-out samples. Neighborhood component analysis (NCA) on the experimental XRD samples shows that the samples from the same material are clustered in groups in the NCA map. Analysis on the class activation maps of the last CNN layer further discloses the mechanism by which the CNN model successfully identifies individual MOFs from the XRD patterns. This CNN model trained by the data augmentation technique would not only open numerous potential applications for identifying XRD patterns for different materials, but also pave avenues to autonomously analyze data by other characterization tools such as FTIR, Raman, and NMR spectroscopies.