{"title":"Individual HRTF Prediction Based on Anthropometric Data and Multi-Stage Model","authors":"Yinliang Qiu, Zhiyu Li, Jing Wang","doi":"10.1109/ICMEW59549.2023.00060","DOIUrl":null,"url":null,"abstract":"Getting individual head related transfer function (HRTF) is an important step in rendering binaural immersive audio. Individual HRTF can provide a more realistic experience than general HRTF. For more accurate prediction results, we propose a multi-stage model perform individual HRTF prediction based on anthropometric data. This model can combine global and local features through different stages. In the first stage, light gradient boosting machine(LightGBM) is chosen as decision tress model to predict HRTF according to anthropometric data and different angels. In the second stage, Transformer encoder is chosen to learn the global information between different frequency points. According to the experimental results, the effect of using a multi-stage model is better than that of a single model. The spectral distortion of the results predicted by our model is smaller, which can illustrate the effectiveness of our model.","PeriodicalId":111482,"journal":{"name":"2023 IEEE International Conference on Multimedia and Expo Workshops (ICMEW)","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2023-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2023 IEEE International Conference on Multimedia and Expo Workshops (ICMEW)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICMEW59549.2023.00060","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0
Abstract
Getting individual head related transfer function (HRTF) is an important step in rendering binaural immersive audio. Individual HRTF can provide a more realistic experience than general HRTF. For more accurate prediction results, we propose a multi-stage model perform individual HRTF prediction based on anthropometric data. This model can combine global and local features through different stages. In the first stage, light gradient boosting machine(LightGBM) is chosen as decision tress model to predict HRTF according to anthropometric data and different angels. In the second stage, Transformer encoder is chosen to learn the global information between different frequency points. According to the experimental results, the effect of using a multi-stage model is better than that of a single model. The spectral distortion of the results predicted by our model is smaller, which can illustrate the effectiveness of our model.