{"id":2328,"date":"2025-06-25T16:35:33","date_gmt":"2025-06-25T07:35:33","guid":{"rendered":"https:\/\/aida.korea.ac.kr\/?page_id=2328"},"modified":"2025-06-25T16:37:46","modified_gmt":"2025-06-25T07:37:46","slug":"632-2","status":"publish","type":"page","link":"https:\/\/aida.korea.ac.kr\/?page_id=2328","title":{"rendered":""},"content":{"rendered":"\n<h1 class=\"wp-block-heading\">Deep Learning \u2013 Sound Event Localization and Detection<\/h1>\n\n\n\n\n<hr class=\"wp-block-separator has-css-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">AD-YOLO: You Look Only Once in Training Multiple Sound Event Localization and Detection<\/h2>\n\n\n\n<p><strong>Objective<\/strong><\/p>\n\n\n\n<p>Given a multi-channel audio input, sound event localization and detection (SELD) combines sound event detection (SED) along the temporal progression and the identification of the direction-of-arrival (DOA) of the corresponding sounds. Several prior works proposed methods to train deep neural networks by representing targets in event\/track-oriented approaches. However, the event-oriented track output formats intrinsically contain the limitation of presetting the number of tracks, constraining the generality and expandability of the method itself.<\/p>\n\n\n\n<p><strong>Data<\/strong><\/p>\n\n\n\n<p>A series of development sets of DCASE Task 3 from 2020 to 2022 [1, 2, 3] are used. In addition, the simulated acoustic scenes to train 2022 baseline [4] is also exploited.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[1] A. Politis, S. Adavanne, and T. Virtanen, \u201cA dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection,\u201d in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020), Tokyo, Japan, November 2020, pp. 165\u2013169. [Online]. Available: https: \/\/dcase.community\/workshop2020\/proceedings<\/p>\n\n\n\n<p class=\"has-small-font-size\">[2] A. Politis, S. Adavanne, D. Krause, A. Deleforge, P. Srivastava, and T. Virtanen, \u201cA dataset of dynamic reverberant sound scenes with directional interferers for sound event localization and detection,\u201d in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021), Barcelona, Spain, November 2021, pp. 125\u2013129. [Online]. Available: <a href=\"https:\/\/dcase.community\/workshop2021\/proceedings\">https:\/\/dcase.community\/workshop2021\/proceedings<\/a><\/p>\n\n\n\n<p class=\"has-small-font-size\">[3] A. Politis, K. Shimada, P. Sudarsanam, S. Adavanne, D. Krause, Y. Koyama, N. Takahashi, S. Takahashi, Y. Mitsufuji, and T. Virtanen, \u201cStarss22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events,\u201d 2022. [Online]. Available: <a href=\"https:\/\/arxiv.org\/abs\/2206.01948\">https:\/\/arxiv.org\/abs\/2206.01948<\/a><\/p>\n\n\n\n<p class=\"has-small-font-size\">[4] [DCASE2022 Task 3] Synthetic SELD mixtures for baseline training, [Online]. Available: 10.5281\/zenodo.6406873<\/p>\n\n\n\n<p><strong>Related Work<\/strong><\/p>\n\n\n\n[5] SELDnet have adopted a two-branch output format, considering SELD as the performing of two separate sub-tasks from each branch, SED and DOA (SED-DOA)<\/p>\n\n\n\n[6, 7] solve the task in a single-branch output through a Cartesian unit vector (proposed as ACCDOA [11]), combining SED and DOA representations, where the zero-vector represents none.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"325\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-26-1024x325.png\" alt=\"\" class=\"wp-image-1962\" srcset=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-26-1024x325.png 1024w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-26-300x95.png 300w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-26-768x244.png 768w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-26-1536x488.png 1536w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-26.png 1618w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p class=\"has-small-font-size\">[5] S. Adavanne, A. Politis, J. Nikunen, and T. Virtanen, \u201cSound event localization and detection of overlapping sources using convolutional recurrent neural networks,\u201d IEEE JSTSP, vol. 13, no. 1, pp. 34\u201348, 2018.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[6] K. Shimada, Y. Koyama, N. Takahashi, S. Takahashi, and Y. Mitsufuji, \u201cAccdoa: Activity-coupled cartesian direction of arrival representation for sound event localization and detection,\u201d in Proc. of IEEE ICASSP, 2021, pp. 915\u2013919.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[7] K. Shimada, Y. Koyama, S. Takahashi, N. Takahashi, E. Tsunoo, and Y. Mitsufuji, \u201cMulti-accdoa: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating permutation invariant training,\u201d in Proc. of IEEE ICASSP, 2022, pp. 316\u2013320.<\/p>\n\n\n\n<p><strong>Proposed Method<\/strong><\/p>\n\n\n\n<p>We proposed an angular-distance-based YOLO (AD-YOLO) approach to perform sound event localization and detection (SELD) on a spherical surface. AD-YOLO assigns multilayered responsibilities, which are based on the angular distance from the target events, to predictions according to each estimated direction of arrival. Avoiding the primal format of the event-oriented track output, AD-YOLO addresses the SELD problem in an unknown polyphony environment.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"454\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-27-1024x454.png\" alt=\"\" class=\"wp-image-1963\" srcset=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-27-1024x454.png 1024w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-27-300x133.png 300w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-27-768x341.png 768w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-27-1536x682.png 1536w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-27.png 1627w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">[8] J. S. Kim, H. J. Park, W. Shin and S. W. Han, &#8220;AD-YOLO: You Look Only Once in Training Multiple Sound Event Localization and Detection,&#8221; in Proc. of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023.<\/figcaption><\/figure>\n\n\n\n<p>We evaluate several existing formats handling the SELD problem using the same backbone network. The model trained in AD-YOLO format achieves the lowest SELD-error (\u03b5_SELD) from all setups. In particular, AD-YOLO outperformed the other approaches in terms of F_(20\u00b0) and LE_CD evaluation metrics.<\/p>\n\n\n\n<p>AD-YOLO proves robustness in class-homogeneous polyphony by the minimum performance degradation (\u0394\u03b5_SELD) compared to the overall evaluation.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"538\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-28-1024x538.png\" alt=\"\" class=\"wp-image-1964\" srcset=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-28-1024x538.png 1024w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-28-300x158.png 300w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-28-768x403.png 768w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-28.png 1200w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">[8] J. S. Kim, H. J. Park, W. Shin and S. W. Han, &#8220;AD-YOLO: You Look Only Once in Training Multiple Sound Event Localization and Detection,&#8221; in Proc. of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023.<\/figcaption><\/figure>\n\n\n\n\n<hr class=\"wp-block-separator has-css-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">A Robust Framework for Sound Event Localization and Detection on Real Recordings<\/h2>\n\n\n\n<p><strong>Objective<\/strong><\/p>\n\n\n\n<p>We address the sound event localization and detection (SELD) problem, of one held at the IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) in 2022. The task aims identification of both the sound-event occurrence (SED) and the direction of arrival from the sound source (DOA). However, by allowing large external datasets and giving small real recordings, the challenge also encompasses a key issue of exploiting synthetic acoustic scenes to perform SELD well in the real world.<\/p>\n\n\n\n<p><strong>Data<\/strong><\/p>\n\n\n\n<p>STARSS22 [1] is in public for the 2022 challenge, comprising real-world recording and label pairs that were man-annotated. To synthesize emulated sound scenarios from the external data, we use class-wise audio samples extracted from seven external datasets, which are AudioSet [2], FSD50K [3], DCASE2020 and 2021 SELD datasets [4, 5], ESC-50 [6], IRMAS [7], and Wearable SELD [8]. As the same way in former SELD task challenges, extracted audio samples are synthesized through SRIR and SNoise from TAU-SRIR DB [9] emulating the spatial sound environment.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[1] A. Politis, K. Shimada, P. Sudarsanam, S. Adavanne, D. Krause, Y. Koyama, N. Takahashi, S. Takahashi, Y. Mitsufuji, and T. Virtanen, \u201cStarss22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events,\u201d 2022. [Online]. Available: <a href=\"https:\/\/arxiv.org\/abs\/2206.01948\">https:\/\/arxiv.org\/abs\/2206.01948<\/a><\/p>\n\n\n\n<p class=\"has-small-font-size\">[2] J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, \u201cAudio set: An ontology and human-labeled dataset for audio events,\u201d in Proc. IEEE ICASSP 2017, New Orleans, LA, 2017.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[3] E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, \u201cFSD50K: an open dataset of human-labeled sound events,\u201d IEEE\/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 829\u2013852, 2022.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[4] A. Politis, S. Adavanne, and T. Virtanen, \u201cA dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection,\u201d in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020), Tokyo, Japan, November 2020, pp. 165\u2013169. [Online]. Available: https: \/\/dcase.community\/workshop2020\/proceedings<\/p>\n\n\n\n<p class=\"has-small-font-size\">[5] A. Politis, S. Adavanne, D. Krause, A. Deleforge, P. Srivastava, and T. Virtanen, \u201cA dataset of dynamic reverberant sound scenes with directional interferers for sound event localization and detection,\u201d in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021), Barcelona, Spain, November 2021, pp. 125\u2013129. [Online]. Available: <a href=\"https:\/\/dcase.community\/workshop2021\/proceedings\">https:\/\/dcase.community\/workshop2021\/proceedings<\/a><\/p>\n\n\n\n<p class=\"has-small-font-size\">[6] K. J. Piczak, \u201cESC: Dataset for Environmental Sound Classification,\u201d in Proceedings of the 23rd Annual ACM Conference on Multimedia. ACM Press, pp. 1015\u2013 1018. [Online]. Available: http:\/\/dl.acm.org\/citation.cfm? doid=2733373.2806390<\/p>\n\n\n\n<p class=\"has-small-font-size\">[7] J. J. Bosch, F. Fuhrmann, and P. Herrera, \u201cIRMAS: a dataset for instrument recognition in musical audio signals,\u201d Sept. 2014. [Online]. Available: https:\/\/doi.org\/10.5281\/zenodo. 1290750<\/p>\n\n\n\n<p class=\"has-small-font-size\">[8] K. Nagatomo, M. Yasuda, K. Yatabe, S. Saito, and Y. Oikawa, \u201cWearable seld dataset: Dataset for sound event localization and detection using wearable devices around head,\u201d in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 156\u2013160.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[9] A. Politis, S. Adavanne, and T. Virtanen, \u201cTAU Spatial Room Impulse Response Database (TAU- SRIR DB),\u201d Apr. 2022. [Online]. Available: https:\/\/doi.org\/10.5281\/zenodo.6408611<\/p>\n\n\n\n<p><strong>Related Work<\/strong><\/p>\n\n\n\n<p>SELDnet [10] established the basic neural network structure to perform SELD, which comprises the layers processing multichannel spectrogram input (2D CNN) followed by sequential processing layers (Bi-GRU) and lastly, fully-connected linear layers.<\/p>\n\n\n\n<p>The squeeze-and-excitation residual networks [11] (SE-ResNet) have recently been applied to audio classification [12, 13], as the SELD encoder.<\/p>\n\n\n\n[14] proposed the method of rotating the sound direction of arrival as the data-augmentation for SELD.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1017\" height=\"499\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-17.png\" alt=\"\" class=\"wp-image-1946\" srcset=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-17.png 1017w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-17-300x147.png 300w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-17-768x377.png 768w\" sizes=\"auto, (max-width: 1017px) 100vw, 1017px\" \/><\/figure><\/div>\n\n\n<p class=\"has-small-font-size\">[10] S. Adavanne, A. Politis, J. Nikunen, and T. Virtanen, \u201cSound event localization and detection of overlapping sources using convolutional recurrent neural networks,\u201d IEEE Journal of Selected Topics in Signal Processing, vol. 13, no. 1, pp. 34\u201348, 2018.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[11] J. Hu, L. Shen, and G. Sun, \u201cSqueeze-and-excitation networks,\u201d in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132\u20137141.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[12] J. H. Yang, N. K. Kim, and H. K. Kim, \u201cSe-resnet with ganbased data augmentation applied to acoustic scene classification,\u201d in DCASE 2018 workshop, 2018.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[13] H. Shim, J. Kim, J. Jung, and H.-j. Yu, \u201cAudio tagging and deep architectures for acoustic scene classification: Uos submission for the dcase 2020 challenge,\u201d Proceedings of the DCASE2020 Challenge, Virtually, pp. 2\u20134, 2020.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[14] L. Mazzon, Y. Koizumi, M. Yasuda, and N. Harada, \u201cFirst order ambisonics domain spatial augmentation for dnn-based direction of arrival estimation,\u201d arXiv preprint arXiv:1910.04388, 2019.<\/p>\n\n\n\n<p><strong>Proposed Method<\/strong><\/p>\n\n\n\n<p>To keep the model fit in real-world scenario contexts while taking advantage of various audio samples from external datasets, a dataset mixing technique (External Mix) is adopted to consist model training dataset. The technique balances the size of each dataset on the model training phase, between the small real recording set and the large emulated scenarios.<\/p>\n\n\n\n<p>A test time augmentation (TTA) is widely used in computer vision to increase the robustness and performance of models. On the other hand, the unknown number of events and the presence of coordinates information make it challenging to apply TTA on SELD. To utilize TTA on SELD, we propose a clustering-based aggregation method to obtain confident-predicted outputs and aggregate them. We take 16 pattern rotation augmentation for test time augmentation, making 16 predicted outputs, that is candidates. To obtain confident aggregated outputs, we use DBSCAN [16] for clustering candidates.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-18.png\" alt=\"\" class=\"wp-image-1947\" width=\"911\" height=\"461\" srcset=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-18.png 911w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-18-300x152.png 300w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-18-768x389.png 768w\" sizes=\"auto, (max-width: 911px) 100vw, 911px\" \/><\/figure><\/div>\n\n\n<p class=\"has-small-font-size\">[15] J. S. Kim*, H. J. Park*, W. Shin*, and S. W. Han**, \u201cA robust framework for sound event localization and detection on real recordings,\u201d Tech. Rep., 3<sup>rd<\/sup> prize for Sound Event Localization and Detection, IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE), 2022.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[16] M. Ester, H.-P. Kriegel, J. Sander, X. Xu, et al., \u201cA densitybased algorithm for discovering clusters in large spatial databases with noise.\u201d in kdd, vol. 96, no. 34, 1996, pp. 226\u2013 231.<\/p>\n\n\n\n<p>We validate the influence of the proposed framework in experiments. The second row (w\/o Ext. Data) of both the first and second blocks is the result of only using small real-world recording data as the training set. As reported in the first rows at the first and second blocks, We observed that using emulated data (Baseline synthesized data [17]), simulated from FSD50K audio samples, enhances the performance of the same models. Concurrently, however, the third row (w\/ Larger Ext. Data) of each shows that the addition of larger emulated soundscapes does not guarantee performance improvement.<\/p>\n\n\n\n<p>In the last block, we found that significant improvements were obtained from three components (Augmentation, External Mix, and TTA). Among them, the external mix method contributed more to the performance improvement than the other methods.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"798\" height=\"455\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-19.png\" alt=\"\" class=\"wp-image-1948\" srcset=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-19.png 798w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-19-300x171.png 300w, https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2023\/05\/image-19-768x438.png 768w\" sizes=\"auto, (max-width: 798px) 100vw, 798px\" \/><\/figure><\/div>\n\n\n<p class=\"has-small-font-size\">[15] J. S. Kim*, H. J. Park*, W. Shin*, and S. W. Han**, \u201cA robust framework for sound event localization and detection on real recordings,\u201d Tech. Rep., 3<sup>rd<\/sup> prize for Sound Event Localization and Detection, IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE), 2022.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[17] [DCASE2022 Task 3] Synthetic SELD mixtures for baseline training, [Online]. Available: 10.5281\/zenodo.6406873<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Deep Learning \u2013 Sound Event Localization and Detection AD-YOLO: You Look Only Once in Training Multiple Sound Event Localization and Detection Objective Given a multi-channel audio input, sound event localization and detection (SELD) combines sound event detection (SED) along the temporal progression and the identification of the direction-of-arrival (DOA) of the corresponding sounds. Several prior &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/aida.korea.ac.kr\/?page_id=2328\" class=\"more-link\">Read more<span class=\"screen-reader-text\"> &#8220;&#8221;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-2328","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/pages\/2328","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2328"}],"version-history":[{"count":2,"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/pages\/2328\/revisions"}],"predecessor-version":[{"id":2355,"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/pages\/2328\/revisions\/2355"}],"wp:attachment":[{"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2328"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}