Identifying robust and dataset-independent acoustic biomarkers of depression through multi-model feature consensus analysis

IF 3 3区 计算机科学 Q2 COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE
Computer Speech and Language Pub Date : 2026-10-01 Epub Date: 2026-02-19 DOI:10.1016/j.csl.2026.101960
Musyysb Yousufi, Rytis Maskeliunas
{"title":"Identifying robust and dataset-independent acoustic biomarkers of depression through multi-model feature consensus analysis","authors":"Musyysb Yousufi,&nbsp;Rytis Maskeliunas","doi":"10.1016/j.csl.2026.101960","DOIUrl":null,"url":null,"abstract":"<div><div>Speech is one of the most abundant and natural sources of acoustic data containing prosodic and spectral information. Acoustic features help diagnose mental and emotional health issues. In recent years, several researchers have looked at speech features as a way to detect depression. However, most of the frameworks only work with the data on which they were trained and do not work with new speakers, recording devices, or languages. This research aims to identify reliable and interpretable acoustic features that serve as stable indicators of depression in various speech datasets.</div><div>This study used two publicly available datasets, E-DAIC and MODMA. A total of 107 handcrafted prosodic, spectral, and voice quality acoustic features were extracted from 4-second segments, with 1-second overlap for long audios and padding for short audio clips. Subject-aware pre-processing was used to prevent speaker level overlap. Five feature selection algorithms were used and their findings were integrated using a consensus-based rank aggregation framework to identify consistent depression related features in both datasets. The resulting set of characteristics was evaluated using four classifier architectures through a K-sweep analysis. The adaptation of the correlation alignment domain was used to reduce distribution mismatches by aligning second-order statistics between the source and target domains, allowing robust cross-dataset transfer evaluation. Bidirectional cross-dataset evaluation demonstrated effective generalization in both transfer directions. Models trained on E-DAIC achieved F1=0.49-0.52 in MODMA (92%–94% of within-dataset performance), while MODMA trained models achieved F1=0.34–0.35 in E-DAIC, exceeding the baseline within-dataset of E-DAIC. The negative domain loss observed in E-DAIC (domain loss = −0.22 to −0.24) reflects high intra-dataset heterogeneity from naturalistic recording conditions rather than poor generalizability. These findings demonstrate that robust acoustic depression biomarkers can be learned from diverse datasets, enabling the detection of cross-linguistic depression.</div></div>","PeriodicalId":50638,"journal":{"name":"Computer Speech and Language","volume":"100 ","pages":"Article 101960"},"PeriodicalIF":3.0000,"publicationDate":"2026-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Computer Speech and Language","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0885230826000239","RegionNum":3,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/2/19 0:00:00","PubModel":"Epub","JCR":"Q2","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0

Abstract

Speech is one of the most abundant and natural sources of acoustic data containing prosodic and spectral information. Acoustic features help diagnose mental and emotional health issues. In recent years, several researchers have looked at speech features as a way to detect depression. However, most of the frameworks only work with the data on which they were trained and do not work with new speakers, recording devices, or languages. This research aims to identify reliable and interpretable acoustic features that serve as stable indicators of depression in various speech datasets.
This study used two publicly available datasets, E-DAIC and MODMA. A total of 107 handcrafted prosodic, spectral, and voice quality acoustic features were extracted from 4-second segments, with 1-second overlap for long audios and padding for short audio clips. Subject-aware pre-processing was used to prevent speaker level overlap. Five feature selection algorithms were used and their findings were integrated using a consensus-based rank aggregation framework to identify consistent depression related features in both datasets. The resulting set of characteristics was evaluated using four classifier architectures through a K-sweep analysis. The adaptation of the correlation alignment domain was used to reduce distribution mismatches by aligning second-order statistics between the source and target domains, allowing robust cross-dataset transfer evaluation. Bidirectional cross-dataset evaluation demonstrated effective generalization in both transfer directions. Models trained on E-DAIC achieved F1=0.49-0.52 in MODMA (92%–94% of within-dataset performance), while MODMA trained models achieved F1=0.34–0.35 in E-DAIC, exceeding the baseline within-dataset of E-DAIC. The negative domain loss observed in E-DAIC (domain loss = −0.22 to −0.24) reflects high intra-dataset heterogeneity from naturalistic recording conditions rather than poor generalizability. These findings demonstrate that robust acoustic depression biomarkers can be learned from diverse datasets, enabling the detection of cross-linguistic depression.
通过多模型特征共识分析识别稳健且与数据集无关的抑郁症声学生物标志物
语音是包含韵律和频谱信息的最丰富和最自然的声学数据来源之一。声学特征有助于诊断精神和情感健康问题。近年来,一些研究人员将语言特征作为一种检测抑郁症的方法。然而,大多数框架只适用于它们所训练的数据,而不适用于新的说话者、录音设备或语言。本研究旨在识别可靠和可解释的声学特征,作为各种语音数据集中抑郁的稳定指标。这项研究使用了两个公开可用的数据集,e - aic和MODMA。从4秒的音频片段中提取了107个手工制作的韵律、频谱和语音质量声学特征,长音频片段有1秒的重叠,短音频片段有1秒的填充。受试者感知预处理用于防止说话人水平重叠。使用了五种特征选择算法,并使用基于共识的排名聚合框架将他们的发现整合起来,以识别两个数据集中一致的抑郁症相关特征。通过k -扫描分析,使用四种分类器架构对结果特征集进行评估。利用相关对齐域的自适应,通过对齐源域和目标域之间的二阶统计量来减少分布不匹配,从而实现鲁棒的跨数据集传输评估。双向跨数据集评估表明,在两个迁移方向上都有有效的泛化。经e - aic训练的模型在MODMA中达到F1=0.49-0.52(92%-94%的数据集内性能),而MODMA训练的模型在e - aic中达到F1= 0.34-0.35,超过e - aic的数据集内基线。在e - aic中观察到的负域损失(域损失= - 0.22至- 0.24)反映了自然记录条件下数据集内部的高异质性,而不是较差的泛化性。这些发现表明,稳健的声学抑郁生物标志物可以从不同的数据集中学习,从而实现跨语言抑郁症的检测。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Computer Speech and Language
Computer Speech and Language 工程技术-计算机:人工智能
CiteScore
11.30
自引率
4.70%
发文量
80
审稿时长
22.9 weeks
期刊介绍: Computer Speech & Language publishes reports of original research related to the recognition, understanding, production, coding and mining of speech and language. The speech and language sciences have a long history, but it is only relatively recently that large-scale implementation of and experimentation with complex models of speech and language processing has become feasible. Such research is often carried out somewhat separately by practitioners of artificial intelligence, computer science, electronic engineering, information retrieval, linguistics, phonetics, or psychology.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书