Adva Levi, Raz Danino, Caren M Rotello, Yonatan Goshen-Gottstein
{"title":"Reject common measures like d' and Corrected Recognition, embrace d<sub>a</sub>: Simulation explorations of single- and multi-point recognition measures of sensitivity.","authors":"Adva Levi, Raz Danino, Caren M Rotello, Yonatan Goshen-Gottstein","doi":"10.3758/s13428-026-03130-w","DOIUrl":"https://doi.org/10.3758/s13428-026-03130-w","url":null,"abstract":"<p><p>In this article, we address the high prevalence of false discoveries in recognition memory research. Using Monte Carlo simulations, our goal was to find a valid measure of performance that reliably separates the contribution of sensitivity (accuracy) from that of bias. Relevant to myriad tasks, most notably old-new recognition memory, the simulations revealed that common measures confound sensitivity with bias, a finding termed the \"measurement crisis.\" As a solution, we propose a version of d-sub-a (d<sub>a</sub>). We ran comprehensive simulations to evaluate the validity of sensitivity measures, including P<sub>r</sub> = HR - FAR, A', d', and AUC<sub>g</sub>, in addition to d<sub>a</sub>. Memory \"signals\" were randomly sampled from lure and target distributions. Sensitivity measures generated from iso-sensitive conditions that differed in bias were compared using t-tests, across thousands of simulations. For bias-independent measures, the rate of significant results should be 5%. We manipulated several parameters, including the form of the distributions (i.e., from three prominent models of recognition-memory: unequal variance signal detection [UVSD], double-high threshold [2HT], dual-process signal detection [DPSD]), the distance between their means, their relative variance, the placement of response criteria, the sample size, and the number of simulated trials. Results demonstrated that under most experimental scenarios, only d<sub>a</sub> was unaffected by changes in bias. In contrast, all common measures typically exhibited alarmingly high false discovery rates, exceeding 5%. The rates rose to 100% with larger sample sizes and a large number of trials. These findings indicate that d<sub>a</sub> warrants serious consideration as the default measure of sensitivity.</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148886369","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Caroline X Gao, Shengqi Wang, Ye Zhu, Myriam Ziou, Shu Mei Teo, Catherine L Smith, Derek Chiu, Aline Talhouk, Mengmeng Wang, Wenhua Yu, Sue M Cotton, Dominic Dwyer
{"title":"Ensemble clustering: A practical tutorial.","authors":"Caroline X Gao, Shengqi Wang, Ye Zhu, Myriam Ziou, Shu Mei Teo, Catherine L Smith, Derek Chiu, Aline Talhouk, Mengmeng Wang, Wenhua Yu, Sue M Cotton, Dominic Dwyer","doi":"10.3758/s13428-026-03158-y","DOIUrl":"https://doi.org/10.3758/s13428-026-03158-y","url":null,"abstract":"<p><p>Cluster analysis is an explorative analytical method, serving as a critical tool in psychology, psychiatry, and related fields to map heterogeneous data into meaningful subgroups. Despite their extensive historical use, traditional clustering techniques suffer from a lack of stability, robustness, and generalisability. These issues stem from the inherent difficulties of the clustering optimisation problem as well as the stochastic nature of algorithm optimisers. To address these challenges, we demonstrate the use of methods utilising ensemble learning techniques to combine clustering results from different algorithms, model specifications, and/or sampled sub-datasets to form a single, more reliable consensus of clustering solutions. We detail ensemble clustering principles and variations in base clustering generation models and ensemble methods. More importantly, detailed introductions in existing R libraries and practical examples using R code are provided to guide users in both implementing and optimising ensemble clustering models. As a practical tutorial, we then include simulation studies of real-world data to demonstrate the substantial benefit of ensemble clustering compared with single-run clustering models. The resources presented here will enable researchers to apply advanced clustering techniques to decompose heterogeneous and complex psychological data into stable subgroups.</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148886322","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Examining within-person variability of each individual: How should we deal with non-varying individuals?","authors":"Xiaohui Luo, Yueqin Hu, Hongyun Liu","doi":"10.3758/s13428-026-03171-1","DOIUrl":"https://doi.org/10.3758/s13428-026-03171-1","url":null,"abstract":"<p><p>Researchers have witnessed a rapid increase in attention to within-person dynamic processes. However, individuals with observed zero within-person variability (i.e., non-varying individuals) remain understudied. This study investigates how non-varying individuals affect the estimation of within-person dynamic processes and offers practical recommendations. A motivating example of daily stressors demonstrates that including non-varying individuals can change substantive conclusions for univariate autoregressive models. Two simulation studies further explored how different proportions of non-varying individuals affect parameter estimation. Study 1 (univariate) found that non-varying individuals induce systematic upward bias in autoregressive estimates. Study 2 (bivariate) showed that, although average (fixed-effect) cross-lagged estimates may appear accurate, person-specific cross-lagged effects are systematically distorted, with some overestimated and others underestimated. Within this widely used Gaussian autoregressive modeling framework, we therefore recommend estimating dynamic parameters using a subsample that excludes non-varying individuals. We also indicate when subsample estimates can be interpreted as approximations to the full population, versus when they should be interpreted as describing the subsample population (e.g., when the proportion of non-varying individuals reaches around 10% or more in bivariate processes). The study highlights the importance of examining each individual's within-person variability and offers valuable guidance for applied researchers on handling non-varying individuals.</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148873011","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Gunnar P Epping, Andrew Caplin, Erik Duhaime, William R Holmes, Daniel Martin, Jennifer S Trueblood
{"title":"Improving crowdsourcing for AI through cognitive-inspired data engineering.","authors":"Gunnar P Epping, Andrew Caplin, Erik Duhaime, William R Holmes, Daniel Martin, Jennifer S Trueblood","doi":"10.3758/s13428-026-03160-4","DOIUrl":"10.3758/s13428-026-03160-4","url":null,"abstract":"<p><p>Crowdsourcing offers a fast and cost-efficient approach to obtaining human-labeled datasets. However, crowdsourced datasets and the models trained on them can inherit the cognitive constraints and biases of their annotators. In a process we refer to as cognitive-inspired data engineering, we investigate whether ideas from cognitive science can be applied to mitigate the presence of cognitive constraints and cognitive biases in crowdsourced datasets and, as a result, improve the performance of models trained on these datasets. We evaluate our approach by crowdsourcing labels for medical image diagnostic tasks using two different crowdsourcing platforms across two experiments. In Experiment 1, we collect subjective probability judgments from novice annotators through Amazon Mechanical Turk and, in Experiment 2, we collect subjective probability judgments and binary classifications from skilled annotators through DiagnosUs, a crowdsourcing platform specializing in medical and scientific data annotation. In both experiments, we find that recalibrating subjective probability judgments reduces bias, yielding more accurate crowdsourced datasets and more accurate models trained on them. Our results suggest that cognitive-inspired data engineering offers a promising avenue to improve the quality of crowdsourced datasets, with consistent downstream benefits for machine learning models.</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13534169/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148872980","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Time after time, the best-worst scaling method is more reliable than ranking.","authors":"Garston Liang, Mackenzie Glover, Guy E Hawkins","doi":"10.3758/s13428-026-03167-x","DOIUrl":"10.3758/s13428-026-03167-x","url":null,"abstract":"<p><p>Two popular methods of preference elicitation are rankings and best-worst scaling (BWS). Rankings, while simple and widely adopted, can be burdensome with larger item sets and fail to capture indifference between options that are neither loved nor hated. Best-worst scaling is a survey method that sidesteps the set size problem by capitalizing on people's natural capacity to identify preferences at the extremes. Across three experiments, our primary finding is that elicited preferences for ranking and BWS methods align, and that BWS methods provide additional resolution to resolve the indifference between middling options where rankings can struggle as well as the relative importance of each option. Moreover, we show that BWS methods exhibit greater test-retest reliability compared to rankings, even over time frames as short as minutes. Taken together, our results privilege BWS as a reliable and readily accessible alternative to ranking methods for preference elicitation.</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13534160/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148873033","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"How accurate are Bayes factor-based null hypothesis tests? A simulation study.","authors":"Daniel J Schad, Martin Modrák","doi":"10.3758/s13428-026-03168-w","DOIUrl":"10.3758/s13428-026-03168-w","url":null,"abstract":"<p><p>Bayes factor null hypothesis tests provide a viable alternative to frequentist measures of evidence quantification. Bayes factors for realistic data sets in areas like psychology cannot be calculated exactly and require numerical approximations to complex integrals. Crucially, the accuracy of these approximations, i.e., whether an approximate Bayes factor corresponds to the exact Bayes factor, is unknown, and may depend on data, prior, and likelihood. We recently developed a novel statistical procedure, marginal simulation-based calibration (SBC) for Bayes factors, to test whether computed Bayes factors for a given analysis are accurate. Here, we use marginal SBC for Bayes factors and calibration plots to test whether Bayes factors are calculated accurately for some common cognitive designs. We use the bridgesampling/brms packages in R. We run analyses for three commonly used designs in psychology and psycholinguistics: (a) a design with random effects for subjects only, (b) a Latin square design with crossed random effects for subjects and items, but a single fixed factor, and (c) a Latin square 2x2 design with crossed random effects for subjects and items. We find that Bayes factor estimates turn out accurate in cases when the bridgesampling algorithm does not issue a warning message, but can be biased and variable when a warning message is shown. These results support the use of brms/bridgesampling for null hypothesis Bayes factor tests in commonly used factorial designs. They also suggest that when a warning message is issued, Bayes factor results should not be trusted. The results show that it is practical to check whether Bayes factors are computed correctly.</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13534162/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148872964","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Andrei A Grigoriev, Konstantin V Sugonyaev, Svetlana A Pashneva, Ekaterina A Valueva, Nikolai A Grigorev
{"title":"Subjective age of acquisition norms for 30,849 Russian words.","authors":"Andrei A Grigoriev, Konstantin V Sugonyaev, Svetlana A Pashneva, Ekaterina A Valueva, Nikolai A Grigorev","doi":"10.3758/s13428-026-03138-2","DOIUrl":"10.3758/s13428-026-03138-2","url":null,"abstract":"<p><p>Age of acquisition (AoA) is a key psycholinguistic variable known to influence lexical and conceptual processing across a broad range of tasks. However, existing Russian datasets are small and mainly limited to picturable nouns and verbs. In this study, we collected AoA ratings for 30,849 Russian words that represent a broad vocabulary range. Ratings were collected from 2,201 adult respondents in both in-person and online sessions. Reliability assessed via bootstrap split-half correlations was high (mean r = 0.895, 95% CI [.893, .897]). To validate the subjective AoA estimates, we administered vocabulary tests to school students from grades 2, 4, 6, 8, and 10. Two validation procedures were used: grade-level validation and within-grade validation. The former related the grade level at which an item was reliably known with the adult ratings, the latter related adult AoA to students' accuracy. Subjective AoA strongly predicted objective AoA measures (r = 0.805) and showed consistent negative relations with student accuracy across grades. Correlations with existing Russian AoA datasets (r = 0.68-0.91) further supported the reliability of the norms. The new ratings also showed expected correlations with word frequency, length, concreteness, imageability, and familiarity. The resulting norms provide the largest and most comprehensive AoA resource for the Russian language and can support future research in psycholinguistics, lexical processing, and language development. All norms are publicly available on OSF.</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13529830/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148863090","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Diego Fernández-Regueras, M Cristina Guerrero-Escagedo, Ana Calero-Elvira, Eduardo Estrada
{"title":"Quantifying the dyadic interaction: A methodological framework for conducting observational moment-by-moment research.","authors":"Diego Fernández-Regueras, M Cristina Guerrero-Escagedo, Ana Calero-Elvira, Eduardo Estrada","doi":"10.3758/s13428-026-03162-2","DOIUrl":"10.3758/s13428-026-03162-2","url":null,"abstract":"<p><p>Systematic observation is a valuable method for studying interactive phenomena in dyads, as it captures sequences of behavior, potentially enhancing validity. However, observational studies remain underutilized compared with other tools, such as self-report questionnaires, because they require time and resources. This work aims to facilitate observational research by providing a step-by-step guide for preprocessing and exploring data from systematic observations, enabling flexible data handling beyond traditional sequential analyses. To illustrate the guide's application, we used a dataset of 154 psychotherapy sessions collected longitudinally across different treatment phases from 32 patient-therapist dyads. We coded sessions using the Therapeutic Relationship Coding System (TRCS), which captures moment-by-moment verbal interactions relevant to the therapeutic alliance. Post-session questionnaire data were collected using the Working Alliance Inventory (WAI). An R script and a set of functions were developed to assist researchers with data preprocessing and to generate descriptive visualizations addressing the following questions: (1) What behaviors (and sequences of interacting behaviors) are more frequent?; (2) How do behaviors evolve over time; (3) Do behavioral changes differ across groups? After applying our proposed analytical tools, we interpret our results in light of the dominant theoretical framework on the therapeutic alliance.</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-08-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13522006/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148839080","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Measuring surprisal in sound sequences.","authors":"Andrey Anikin","doi":"10.3758/s13428-026-03153-3","DOIUrl":"10.3758/s13428-026-03153-3","url":null,"abstract":"<p><p>Sensory input that violates prior expectations attracts attention, making unpredictability an important perceptual property to measure. In the auditory modality, knowing what sounds will be perceived as surprising, and therefore salient, is relevant both for studying vocal communication and for applied purposes such as managing noise pollution. Focusing on sequences of animal vocalizations and environmental sounds as ecologically important acoustic stimuli, I describe and benchmark several algorithms for measuring their perceived unpredictability. Information-theoretical approaches include Shannon surprisal and Bayesian surprise, both implemented here to detect deviant stimuli based on distributional acoustic properties. The second group of algorithms is based on detecting spectro-temporal recurrence assessed with autocorrelation functions (ACF surprisal) and self-similarity matrices (SSM novelty). The third approach uses neural networks. Based on the ratings of the predictability of 300 synthetic acoustic sequences by 195 human listeners, Shannon surprisal and SSM novelty capture the perceived unpredictability that is due to spectral variability, whereas ACF surprisal taps into the perceptual impact of irregular rhythm. Most algorithms converge on the time scale of about 1 s as the most perceptually relevant for spectral variability, which is consistent with the hypothesis that the perception of unpredictability stems from a relatively limited amount of auditory input held in short-term memory. Together, the presented open-source algorithms offer powerful and flexible tools for measuring acoustic surprisal and studying auditory attention, while the corpus of predictability ratings offers a resource for future benchmarking. All code and data are freely available from the R package soundgen and supplementary materials at https://osf.io/bgzvc .</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-08-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13503466/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148811888","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Julius Grote, Julia Elina Stocker, Jens Sommer, Anna-Maria Hamm, Henrik Kessler, Andreas Jansen
{"title":"Tetrix: A novel Tetris-based paradigm for neuroimaging research and clinical applications.","authors":"Julius Grote, Julia Elina Stocker, Jens Sommer, Anna-Maria Hamm, Henrik Kessler, Andreas Jansen","doi":"10.3758/s13428-026-03150-6","DOIUrl":"10.3758/s13428-026-03150-6","url":null,"abstract":"<p><p>Tetris is a widely used computer game, not only for entertainment purposes but increasingly in cognitive and clinical neuroscience research. Despite its frequent application in experiments, the neural mechanisms underlying gameplay remain insufficiently understood. Prior work suggests that cognitive control during complex visuospatial tasks, like Tetris, engages the attention network as well as regions supporting visuospatial working memory, mental imagery, and motor planning. However, researchers currently lack a standardized paradigm to investigate these processes. To address this gap, we introduce Tetrix, a novel and flexible Tetris-based paradigm designed for use in behavioral and functional magnetic resonance imaging (fMRI) experiments. With a broad range of customizable features, Tetrix can be adapted to specific experimental requirements or used in its default configuration. To validate the paradigm's feasibility for neuroimaging applications, we conducted a proof-of-concept fMRI study with six participants. Results revealed robust bilateral activity in the frontal eye fields and posterior parietal cortex - key nodes of the dorsal attention network. Additional activation was found in the cerebellum and occipital cortex. These regions are thought to support rapid spatial orienting, the encoding and maintenance of visuospatial working memory, and motor planning. The observed activation patterns therefore align with theoretical expectations and confirm Tetrix's utility for probing relevant cognitive functions in fMRI. Moreover, the pilot dataset and corresponding effect size maps provide a valuable resource for estimating statistical power in future research. In summary, this article introduces Tetrix as a standardized and validated paradigm for functional neuroimaging and beyond. It is available at: https://github.com/JuliusGrote/Tetrix_Psychopy .</p>","PeriodicalId":8717,"journal":{"name":"Behavior Research Methods","volume":"58 10","pages":""},"PeriodicalIF":5.0,"publicationDate":"2026-08-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13503407/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148811909","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}