Speakers
Description
Industrial robots inevitably suffer from performance degradation and mechanical faults during long-term operation, making accurate fault diagnosis essential for ensuring system reliability and reducing maintenance costs. Recently, multimodal fault diagnosis methods have attracted increasing attention because multiple sensor modalities are generally assumed to provide complementary fault information. However, the effectiveness of multimodal fusion strongly depends on the fault relevance of each modality, which remains insufficiently investigated in industrial applications. This paper presents an empirical study on pose-vibration multimodal fusion for industrial robot fault diagnosis, with particular emphasis on the phenomenon of multimodal fusion failure. A ResNet1D-based vibration encoder and a lightweight pose encoder are developed to extract modality-specific representations, while concatenation-based fusion and channel-attention fusion strategies are further investigated. Experimental results show that vibration signals alone achieve the best diagnostic performance with 99.23% test accuracy, whereas pose signals only achieve 67.88%, indicating limited fault-discriminative capability in the pose modality. More importantly, introducing pose information consistently degrades diagnostic performance: concatenation fusion reduces the accuracy to 94.42%, while attention fusion achieves 96.34%, both remaining inferior to the vibration-only baseline. Attention weight analysis further reveals that the learned fusion model consistently assigns higher importance to vibration features while suppressing pose-related representations. The results suggest that weakly fault-related modalities may introduce negative transfer and cross-modal interference rather than performance gains. This study provides practical insights into modality selection and reliability evaluation for multimodal industrial fault diagnosis systems.