Natural Language Processing and soft data for motor skill assessment: A case study in surgical training simulations

IF 4.9 2区医学 Q1 COMPUTER SCIENCE, INTERDISCIPLINARY APPLICATIONS

Computer methods and programs in biomedicine Pub Date : 2025-03-04 DOI:10.1016/j.cmpb.2025.108686

Arash Iranfar , Mohammad Soleymannejad , Behzad Moshiri , Hamid D. Taghirad

{"title":"Natural Language Processing and soft data for motor skill assessment: A case study in surgical training simulations","authors":"Arash Iranfar , Mohammad Soleymannejad , Behzad Moshiri , Hamid D. Taghirad","doi":"10.1016/j.cmpb.2025.108686","DOIUrl":null,"url":null,"abstract":"<div><h3>Background and Objective:</h3><div>Automated surgical skill assessment using kinematic and video data (hard data) sources has been widely adopted in the literature. However, experts’ opinions (soft data) in the form of free-text could be an invaluable source for evaluating one’s skill level since the availability and semantic richness of the soft data are both higher than the hard data. In this paper, the feasibility of using soft data as a single source of skill assessment is analyzed with various Natural Language Processing (NLP) algorithms of different levels of complexity.</div></div><div><h3>Methods:</h3><div>An experiment named “Vertex Pursuit” was designed to address the absence of a dataset with free-text soft data in synchronization with hard data. This experiment challenges participants’ hand-eye coordination, both-hand coordination, precision, and dexterity by tracking haptic device movements along a star pentagon shape. Top-performing participants receive additional training to provide expert feedback through free-text comments evaluating their peers’ trials. Traditional machine learning approaches are employed, including various word and sentence embedding techniques combined with a diverse set of classifiers, to assess skill levels based on this soft data. Additionally, encoder-only and decoder-only large language models (LLMs) are applied to the data, with the latter leveraging three prompt engineering techniques.</div></div><div><h3>Results:</h3><div>The task of skill assessment using soft data is demonstrated to be a complex NLP task, and as the complexity of the method increases, the results improve. The top performance was achieved with the decoder-only LLMs and the rule-based prompting strategy.</div></div><div><h3>Conclusion:</h3><div>This paper studied the feasibility of using soft data in a simulated surgical skill assessment scenario. While further research is needed, the proposed methods can reduce subjectivity, alleviate the burden on human experts, and enable more widespread, scalable skill evaluation in surgical training programs.</div></div>","PeriodicalId":10624,"journal":{"name":"Computer methods and programs in biomedicine","volume":"264 ","pages":"Article 108686"},"PeriodicalIF":4.9000,"publicationDate":"2025-03-04","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Computer methods and programs in biomedicine","FirstCategoryId":"5","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0169260725001038","RegionNum":2,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, INTERDISCIPLINARY APPLICATIONS","Score":null,"Total":0}

引用次数: 0

Abstract

Background and Objective:

Automated surgical skill assessment using kinematic and video data (hard data) sources has been widely adopted in the literature. However, experts’ opinions (soft data) in the form of free-text could be an invaluable source for evaluating one’s skill level since the availability and semantic richness of the soft data are both higher than the hard data. In this paper, the feasibility of using soft data as a single source of skill assessment is analyzed with various Natural Language Processing (NLP) algorithms of different levels of complexity.

Methods:

An experiment named “Vertex Pursuit” was designed to address the absence of a dataset with free-text soft data in synchronization with hard data. This experiment challenges participants’ hand-eye coordination, both-hand coordination, precision, and dexterity by tracking haptic device movements along a star pentagon shape. Top-performing participants receive additional training to provide expert feedback through free-text comments evaluating their peers’ trials. Traditional machine learning approaches are employed, including various word and sentence embedding techniques combined with a diverse set of classifiers, to assess skill levels based on this soft data. Additionally, encoder-only and decoder-only large language models (LLMs) are applied to the data, with the latter leveraging three prompt engineering techniques.

Results:

The task of skill assessment using soft data is demonstrated to be a complex NLP task, and as the complexity of the method increases, the results improve. The top performance was achieved with the decoder-only LLMs and the rule-based prompting strategy.

Conclusion:

This paper studied the feasibility of using soft data in a simulated surgical skill assessment scenario. While further research is needed, the proposed methods can reduce subjectivity, alleviate the burden on human experts, and enable more widespread, scalable skill evaluation in surgical training programs.

查看原文本刊更多论文

求助全文

约1分钟内获得全文求助全文

来源期刊

Computer methods and programs in biomedicine 工程技术-工程：生物医学

CiteScore

12.30

自引率

6.60%

发文量

601

审稿时长

135 days

期刊介绍： To encourage the development of formal computing methods, and their application in biomedical research and medical practice, by illustration of fundamental principles in biomedical informatics research; to stimulate basic research into application software design; to report the state of research of biomedical information processing projects; to report new computer methodologies applied in biomedical areas; the eventual distribution of demonstrable software to avoid duplication of effort; to provide a forum for discussion and improvement of existing software; to optimize contact between national organizations and regional user groups by promoting an international exchange of information on formal methods, standards and software in biomedicine. Computer Methods and Programs in Biomedicine covers computing methodology and software systems derived from computing science for implementation in all aspects of biomedical research and medical practice. It is designed to serve: biochemists; biologists; geneticists; immunologists; neuroscientists; pharmacologists; toxicologists; clinicians; epidemiologists; psychiatrists; psychologists; cardiologists; chemists; (radio)physicists; computer scientists; programmers and systems analysts; biomedical, clinical, electrical and other engineers; teachers of medical informatics and users of educational software.