{"title":"Robert J. Mislevy and Computational Psychometrics","authors":"Alina A. von Davier, Jiangang Hao","doi":"10.1111/emip.70024","DOIUrl":"https://doi.org/10.1111/emip.70024","url":null,"abstract":"<p>As part of this EM:IP special issue honoring Robert (Bob) J. Mislevy, this paper highlights his critical role in shaping the emerging field of computational psychometrics, which integrates computational methods from data science and artificial intelligence (AI) with core psychometric principles to address the challenges posed by complex, high-volume multimodal data from interactive digital assessments. We discuss how Bob's vision, mentorship, and intellectual leadership helped to establish the conceptual foundations of this field, and how his foresight continues to guide the ongoing transformation of assessment in the era of AI.</p>","PeriodicalId":47345,"journal":{"name":"Educational Measurement-Issues and Practice","volume":"45 3","pages":""},"PeriodicalIF":1.9,"publicationDate":"2026-06-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148153603","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"教育学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Digital Module 40: Introduction to Machine Learning and Generative AI: From AutoGluon to Amazon Bedrock","authors":"Ye Ma, Vinita Talreja","doi":"10.1111/emip.70029","DOIUrl":"https://doi.org/10.1111/emip.70029","url":null,"abstract":"<p>Machine learning (ML) and generative artificial intelligence (AI) are rapidly transforming the field of educational measurement. This module focuses on illustrating the process of (1) automated machine learning (AutoML) using AutoGluon via an application of detecting aberrant test behavior and (2) AI-based item generation using Amazon Bedrock. To support these demonstrations, two tools are used: AutoGluon, an open-source automated machine learning (AutoML) system, and Amazon Bedrock, a fully managed AWS service for accessing foundation models from leading AI providers. By the end of this module, participants will (1) understand the key concepts and fundamentals underlying these two applications and (2) be able to programmatically train a classification ML model via AutoML using the provided data, as well as conduct AI-based item generation via the LLMs.</p>","PeriodicalId":47345,"journal":{"name":"Educational Measurement-Issues and Practice","volume":"45 2","pages":""},"PeriodicalIF":1.9,"publicationDate":"2026-05-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148091995","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"教育学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Grounding Meaning and Warranting Inference in Educational Measurement: Theory, Epistemology, and Practice in Mislevy's Sociocognitive Framework","authors":"David Torres Irribarra, Joshua A. McGrane","doi":"10.1111/emip.70023","DOIUrl":"https://doi.org/10.1111/emip.70023","url":null,"abstract":"<p>Educational measurement has long been characterized by a productive tension between technical sophistication and the conceptual frameworks used to interpret what measurement outcomes mean. While the field has produced major methodological advances, comparatively fewer contributions have reshaped how psychometric evidence is connected to theories of cognition, epistemology, and use. This article examines Robert J. Mislevy's <i>Sociocognitive Foundations of Educational Measurement</i> as a landmark effort to address this imbalance. We argue that Mislevy's framework advances the field along three dimensions: reframing constructs through sociocognitive theories of learning and practice; grounding assessment in an epistemology of evidentiary reasoning under uncertainty, aligned with Bayesian inference; and articulating a principled pluralism across methodological traditions. Central to this contribution is a disciplined stance toward inference that treats epistemic restraint as rigor. We discuss the framework's relevance for contemporary challenges, including fairness, socially responsible use, and artificial intelligence in assessment.</p>","PeriodicalId":47345,"journal":{"name":"Educational Measurement-Issues and Practice","volume":"45 2","pages":""},"PeriodicalIF":1.9,"publicationDate":"2026-05-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148090600","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"教育学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Brooke Nash, Meagan Karvonen, Claudia Flowers, Amy Clark, Russell Swinburne Romine, Sue Bechard
{"title":"A Comprehensive Alignment Framework for Building Coherence in an Assessment System","authors":"Brooke Nash, Meagan Karvonen, Claudia Flowers, Amy Clark, Russell Swinburne Romine, Sue Bechard","doi":"10.1111/emip.70027","DOIUrl":"https://doi.org/10.1111/emip.70027","url":null,"abstract":"<p>Traditional alignment studies that are narrowly focused on content standards and assessment items and are conducted postadministration often lack the timeliness and utility needed to improve assessment systems. This paper introduces a flexible alignment framework centered on assessment-system coherence. Our framework broadens alignment evaluation to include all interconnected elements, from domain definition to score use. We advocate for proactive alignment evaluation by defining expectations in advance and gathering both procedural and empirical evidence throughout the assessment life cycle. By addressing three guiding questions regarding system elements that need to be aligned, the purpose of the alignment evidence, and the types of evidence to use, developers can create customized evaluation plans to enhance alignment during development. The paper provides example alignment questions and an example application of the framework. We discuss the framework's potential to enhance the coherence of assessment systems in the service of meaningful policy and practice.</p>","PeriodicalId":47345,"journal":{"name":"Educational Measurement-Issues and Practice","volume":"45 2","pages":""},"PeriodicalIF":1.9,"publicationDate":"2026-05-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1111/emip.70027","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148090601","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"教育学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Digital Module 41: Process Data","authors":"Susu Zhang, Qiwei He, Sunbeom Kwon","doi":"10.1111/emip.70030","DOIUrl":"https://doi.org/10.1111/emip.70030","url":null,"abstract":"<p>Process data, such as log files from digital assessments, provide detailed records of how examinees interact with assessment tasks. These data offer opportunities to study test-taking behavior, strategy use, and human-machine interaction in ways that final item scores alone cannot capture. This module introduces the structure and characteristics of process data in large-scale digital assessments and presents several approaches for transforming raw action sequences into numerical features that can be used in statistical and psychometric analyses. The module covers both expert-derived and data-driven feature extraction methods, including pattern-based indicators, n-grams, multidimensional scaling, and sequence autoencoders. A hands-on section demonstrates data wrangling and feature extraction in R using the PISA 2012 Climate Control item and the ProcData package. The final section presents case studies showing how process-derived features can be used to study test accommodations, improve measurement precision, and reduce and interpret differential item functioning. By the end of the module, learners should have a practical introduction to process data and a foundation for incorporating process-derived information into educational measurement research.</p>","PeriodicalId":47345,"journal":{"name":"Educational Measurement-Issues and Practice","volume":"45 2","pages":""},"PeriodicalIF":1.9,"publicationDate":"2026-05-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1111/emip.70030","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148090614","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"教育学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Bob Mislevy's Contributions to Large-Scale Assessments","authors":"John R. Donoghue","doi":"10.1111/emip.70025","DOIUrl":"https://doi.org/10.1111/emip.70025","url":null,"abstract":"<p>Bob Mislevy made several important contributions to large-scale educational survey assessments (LSAs). The contributions include implementing marginal maximum likelihood estimation of item response theory model parameters, using latent variable regression to condition proficiency estimates on background characteristics, and using multiple imputation (plausible values) to report assessment results. All these pieces were needed to analyze LSA data, report the results, and produce the data products needed for secondary users. In addition, Mislevy provided analyses of potential bias, giving users confidence that this novel approach to producing results yielded reliable information. Although procedures have since evolved, the approach and combination of techniques introduced in the 1984 NAEP reading assessment remain the cornerstones of LSA analysis and reporting in NAEP and internationally, some 40 years later.</p>","PeriodicalId":47345,"journal":{"name":"Educational Measurement-Issues and Practice","volume":"45 2","pages":""},"PeriodicalIF":1.9,"publicationDate":"2026-05-19","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148086594","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"教育学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Knowing What Students Know—An Essay on the Contributions of Robert Mislevy","authors":"James W. Pellegrino, Mark Wilson, Kadriye Ercikan","doi":"10.1111/emip.70026","DOIUrl":"https://doi.org/10.1111/emip.70026","url":null,"abstract":"<p>In this essay, we recount Robert Mislevy's influence on the three of us and several others in the context of the National Research Council's <i>Committee on the Foundations of Assessment</i>, which produced the report <i>Knowing What Student's Know: The Science and Design of Education Assessment</i> (aka <i>KWSK</i>; NRC). We start by providing a context for the Committee's work, describe critical elements of the <i>KWSK</i> report, including their impact on the field of educational assessment over the last 25 years, and emphasize ways in which Bob influenced their articulation and impact.</p>","PeriodicalId":47345,"journal":{"name":"Educational Measurement-Issues and Practice","volume":"45 2","pages":""},"PeriodicalIF":1.9,"publicationDate":"2026-05-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148077737","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"教育学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"When Is Classroom Assessment Educational Measurement?","authors":"Susan M. Brookhart, Sarah M. Bonner","doi":"10.1111/emip.70021","DOIUrl":"https://doi.org/10.1111/emip.70021","url":null,"abstract":"<p>The relationship between classroom assessment and educational measurement has been under discussion for some time. This article uses the TISM framework (<i>Theory, Instrumentation, Scales and units</i>, and <i>Modeling</i>) to clarify which aspects of classroom assessment are educational measurement (e.g., a grade on a performance assessment keyed to a learning standard) and which are not (e.g., extended elaborated feedback on that same assessment). We conclude that classroom assessment which produces ordinal or interval-level quantitative scores—by whatever name they are called, including scores, grades, and performance levels—is educational measurement because it implicates theory, instrumentation, scales and units, and modeling of error. On this basis, we claim that work in classroom assessment and educational measurement can and should be mutually informative.</p>","PeriodicalId":47345,"journal":{"name":"Educational Measurement-Issues and Practice","volume":"45 2","pages":""},"PeriodicalIF":1.9,"publicationDate":"2026-04-06","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147715036","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"教育学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Alexander Kwako, Susan Lottridge, Christopher Ormerod
{"title":"Using Confidence Modeling to Optimize Overall Score Quality in Hybrid Scoring Systems","authors":"Alexander Kwako, Susan Lottridge, Christopher Ormerod","doi":"10.1111/emip.70019","DOIUrl":"https://doi.org/10.1111/emip.70019","url":null,"abstract":"<p>In large-scale assessments, constructed response items are often scored using hybrid scoring systems, which combine human and automated scores. In this study, we augment automated scoring with confidence modeling to strategically route difficult-to-score responses for human review. We utilize <i>hybrid performance curves</i> to visualize the impact of routing on performance. Additionally, we propose several <i>hybrid scoring policies</i> for selecting optimal routing thresholds given practical constraints. Our findings reveal that hybrid scoring systems can achieve an overall performance that exceeds that of human- and automated-only systems. Moreover, the superior performance of the hybrid system is less expensive than a human-only system. These findings highlight the complementarity of human raters and automated scoring engines. Although current standards focus on the performance of human raters and automated scoring engines in isolation, we recommend that practitioners also report on the performance of the hybrid scoring system as a whole.</p>","PeriodicalId":47345,"journal":{"name":"Educational Measurement-Issues and Practice","volume":"45 2","pages":""},"PeriodicalIF":1.9,"publicationDate":"2026-04-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1111/emip.70019","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147714992","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"教育学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}