상세 보기
초록
In this study, we used deep learning to align bone-conducted speech signals with air-conducted speech signals, aiming to replace traditional air conduction microphones in voice-based services capturing surrounding sounds. We fabricated headphones, placing bone conduction microphones on the rami (the branches of a bone in the jaw area), in line with traditional bone conduction headphone configurations. Using LSTM, CNN, and CRNN models, we created databases that aligned bone-conducted speech signals with their air-conducted counterparts and tested them with bone-conducted speech signals captured via our custom-made headphones. The CNN model demonstrated superior performance in accurately distinguishing three English words (“apple,” “hello,” and “pass”), including their voiceless pronunciations. In conclusion, our study shows that deep learning models can effectively use bone-conducted speech signals extracted from the rami for automatic speech recognition (ASR), paving the way for future ASR technology that precisely recognizes only the speaker’s voice.
키워드
- 제목
- 골전도 헤드폰 형태로 추출된 골전도 음성 신호의 딥러닝 활용
- 제목 (타언어)
- Application of Deep Learning Models for Bone-Conducted Speech Signals Extracted in the Form of Bone Conduction Headphones
- 저자
- 송희주; 유선아; 손세강; 장웅기; 황향희; 김현욱; 김병희; 이형석
- 발행일
- 2024-02
- 유형
- Y
- 저널명
- 한국생산제조학회지
- 권
- 33
- 호
- 1
- 페이지
- 27 ~ 34