<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
	<id>https://www.jmir.org/issue/feed</id>
	<title>Journal of Medical Internet Research</title>
			<updated>2025-01-01T11:30:03-05:00</updated>
	
		<author>
		<name>JMIR Publications</name>
				<email>editor@jmir.org</email>
			</author>
		<link rel="alternate" href="https://www.jmir.org" />
	<link rel="self" type="application/atom+xml" href="https://www.jmir.org/feed/atom" />

	<generator uri="http://pkp.sfu.ca/ojs/" version="2.2.0.0">Open Journal Systems</generator>

				    	<subtitle> The leading peer-reviewed journal for digital medicine and health and health care in the internet age.&amp;nbsp; </subtitle>



	<entry>
		<id> https://www.jmir.org/2026/1/e98184 </id>
		<title>The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review</title>
		<updated>2026-08-19T17:00:03-04:00</updated>

					<author>
				<name>Euijun Yang</name>
			</author>
					<author>
				<name>Siyeon Ko</name>
			</author>
					<author>
				<name>Hyekyung Woo</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e98184" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e98184">&lt;strong&gt;Background:&lt;/strong&gt; Integration of large language models (LLMs) into health care has accelerated rapidly, yet reliability concerns pose potential risks to patient safety. Although human evaluation has been widely used as an important approach for assessing LLM reliability, a systematic understanding of how such evaluations have been operationalized across studies remains limited. &lt;strong&gt;Objective:&lt;/strong&gt; This study aimed to characterize the current landscape of human evaluation frameworks for LLM reliability in health care and to identify similarities and differences between the clinical and public health domains. &lt;strong&gt;Methods:&lt;/strong&gt; In line with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines, PubMed, Web of Science, the Cochrane Library, CINAHL, and Google Scholar were searched for studies published from January 2016 to July 2025. Eligible studies were English-language original research conducted in health care settings that assessed the reliability of LLM-generated responses through human evaluation. Key exclusion criteria were studies without human evaluation and studies focused primarily on LLM model selection, performance optimization, or technical development. Extracted data were analyzed across 3 dimensions: what was evaluated, who evaluated, and how evaluation was conducted. Reported methodological limitations were also categorized and compared between the clinical and public health domains. &lt;strong&gt;Results:&lt;/strong&gt; Of the 4347 records identified, 71 studies were included in the final analysis (clinical, n=26; public health, n=45). Six reliability indicators were used: accuracy, relevance, completeness, clarity, safety, and consistency. The clinical domain more frequently assessed guideline concordance, internal consistency, and structural coherence, whereas the public health domain more frequently assessed understandability, harm potential, and repeat response consistency. Single-specialty clinicians were the most common evaluators in both domains, although mixed evaluator panels were observed only in the public health domain. Evaluator panels generally consisted of 5 or fewer members. Five-point Likert scales and researcher-defined rubrics were commonly used evaluation approaches in both domains. Key methodological limitations included evaluator subjectivity, nonstandardized indicators, and limited evaluation scope and settings. &lt;strong&gt;Conclusions:&lt;/strong&gt; To our knowledge, this is the first review to systematically examine how human evaluations of LLM reliability have been conducted across health care. The focus of reliability evaluation differed across domains, with clinical evaluations giving relatively greater attention to clinical validity and logical rigor, whereas public health evaluations gave relatively greater attention to understandability, practical use, and safe use. These differences suggest that the reliability of health care LLMs is difficult to evaluate adequately using a single universal standard. In addition, the methodological limitations identified in this review indicate that current human evaluation approaches are insufficiently standardized. Therefore, future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts and encompass indicator definitions, judgment criteria, evaluator guidance, and evaluation procedures. </summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/4cc16cc411fed1f70e8e22ebcf80a2a4" />
		
		<published>2026-08-19T17:00:03-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e93162 </id>
		<title>Data Maturity Assessment in Long-Term Care: Mixed Methods Study</title>
		<updated>2026-08-19T16:00:20-04:00</updated>

					<author>
				<name>Suleyman Bouchmal</name>
			</author>
					<author>
				<name>Katya Sion</name>
			</author>
					<author>
				<name>Pepijn Janssens</name>
			</author>
					<author>
				<name>Manouk ten Thij</name>
			</author>
					<author>
				<name>Jan Hamers</name>
			</author>
					<author>
				<name>Sil Aarts</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e93162" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e93162">Background: The use of data to support decision-making and primary processes is central to establishing data-informed care. Yet, data remain underutilized for quality improvement in long-term care (LTC). Data maturity reflects an organization’s capability to use data, for example, to guide strategic objectives. Objective: This study aimed to assess the data maturity of LTC organizations and evaluate the opinions and ideas of different stakeholders (eg, client representatives, care professionals, and data and IT specialists) within these organizations regarding their data maturity. Methods: An exploratory mixed methods study was conducted in 3 Dutch LTC organizations. Quantitative data were collected using a 7-domain data maturity assessment comprising 77 items with 3 response options (applicable, not applicable, or in progress), which was completed through consensus-based scoring. Each organization established an interprofessional community of practice (CoP), consisting of 5 to 6 purposively recruited stakeholders representing care professionals, managers, client representatives, IT and data specialists, and researchers. Qualitative data were collected through consensus meetings with the same CoP stakeholders to elaborate on and contextualize their data maturity assessments. Quantitative data were analyzed descriptively, while qualitative data were analyzed using hybrid thematic analysis. Results: Findings highlighted that LTC organizations currently operate at a low data maturity level. Domains regarding strategy, governance, and data quality were low yet consistent across organizations, whereas variability and substantial gaps were mainly present in the domain regarding leadership and culture. The consensus meetings with 15 stakeholders in the CoPs (mean age 43, SD 11 y; mean organizational work experience 10, SD 9 y) identified 4 overarching themes: (1) the absence of a clear vision for data-informed care reflects low data maturity, (2) a data culture is a prerequisite for continuous learning and improvement, (3) innovation and transformation are constrained by limited interdisciplinary engagement and ineffective communication, and (4) sustainable data use is hindered by infrastructural capacity and cultural readiness. Conclusions: This study contributes to the limited evidence on data maturity in LTC, revealing generally low levels of data maturity across organizations. Advancing data-informed care requires integrated strategies that align technical, cultural, and interprofessional conditions. Future research should adopt longitudinal designs and broader stakeholder involvement to strengthen data-informed care in LTC.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/62707d9117e713dbac17589d237f16bf" />
		
		<published>2026-08-19T16:00:20-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e96347 </id>
		<title>Multilingual Evidence-Based Question-Answering for Stroke Discharge Summaries: Study of Cross-Lingual Heterogeneity in Clinical Reports</title>
		<updated>2026-08-19T15:30:14-04:00</updated>

					<author>
				<name>Vojtěch Lanz</name>
			</author>
					<author>
				<name>Aleksis Datseris</name>
			</author>
					<author>
				<name>Svetla Boytcheva</name>
			</author>
					<author>
				<name>Jiří Mayer</name>
			</author>
					<author>
				<name>Šárka Zikánová</name>
			</author>
					<author>
				<name>Robert Mikulík</name>
			</author>
					<author>
				<name>Pavel Pecina</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e96347" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e96347">Background: The Registry of Stroke Care Quality (RES-Q) is a health care quality improvement platform used globally. RES-Q collects structured quality-of-care data for patients with stroke, requiring clinicians to manually extract information from electronic health records or documents such as discharge summaries. This process is essential but time-consuming, particularly given the variability, length, and semistructured nature of clinical reports. Objective: This study aimed to develop and evaluate a multilingual Evidence-Based Question-Answering framework that identifies supporting text spans in clinical reports of patients with stroke and proposes answer suggestions for structured clinical forms, with the goal of reducing clinician workload while preserving full human oversight. Methods: We conduct a multilingual study using more than 1500 pseudonymized stroke discharge summaries in 5 languages, annotated with question-evidence-answer triplets. Encoder-based language models are used to extract evidence spans from the reports, while generative language models are used to predict normalized form answers based on the extracted evidences. We compare multiple training strategies, such as (1) models trained on reports in a single target language, (2) models trained jointly on reports in different languages, and (3) models trained on original reports combined with cross-lingual data augmentations. We evaluate performance on Evidence Extraction, Answer Prediction, and end-to-end Evidence-Based Question Answering across the 5 languages. Results: The presented Evidence-Based Question-Answering system achieves 88% end-to-end accuracy in form filling across 5 languages (77% for patient-specific questions and 95% for default or unverifiable items). Evidence Extraction is the primary bottleneck, reaching 85% and 79% exact match, whereas Answer Prediction based on extracted evidences is more stable, achieving 95% accuracy. The performance varies by question type, and cross-lingual training generally reduces Evidence Extraction performance but has little effect on Answer Prediction. Model performance is influenced more by reporting practices and dataset characteristics than by language itself. Conclusions: Evidence-Based Question Answering over multilingual stroke discharge summaries enables human-in-the-loop validation and effective answer prediction with moderate computational resources. Evidence Extraction is the main bottleneck, while Answer Prediction is robust across languages and model sizes. The approach supports structured data collection, although generalization to new languages requires target-language training data.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/35860156ad4b0f26684b35ad4c046b9e" />
		
		<published>2026-08-19T15:30:14-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e98699 </id>
		<title>The Continuity Trap in Data Science Health Research</title>
		<updated>2026-08-19T08:30:03-04:00</updated>

					<author>
				<name>Clement Adebamowo</name>
			</author>
					<author>
				<name>Sally Nneoma Adebamowo</name>
			</author>
					<author>
				<name>Adeola Akintola</name>
			</author>
					<author>
				<name>Peter Ikhane</name>
			</author>
					<author>
				<name>Simisola Akintola</name>
			</author>
					<author>
				<name>Temidayo Ogundiran</name>
			</author>
					<author>
				<name>Ayodele Jegede</name>
			</author>
					<author>
				<name>Olusegun Adeyemo</name>
			</author>
					<author>
				<name>Shawneequa Callier</name>
			</author>
					<author>
				<name>Muhammad Imam-Tamim</name>
			</author>
					<author>
				<name>Ibrahim Uthman</name>
			</author>
					<author>
				<name>BridgELSI Project as part of the DS-I Africa Consortium</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e98699" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e98699">Secondary use is now the ordinary condition of data science health research rather than an exception to it. Electronic health records collected for clinical care become prediction tools and inputs for generative AI; imaging archives become foundation-model corpora; genomic datasets become resources for polygenic risk scores; and legacy biospecimens become renewable, indefinitely distributable cell lines. Governance has responded by emphasizing verifiable instruments such as provenance logs, repository approvals, broad-consent forms, data-use agreements, model cards, records of processing, and locality-preserving architectures. These instruments are necessary, and they answer real questions about lineage, privacy, institutional responsibility, and accountability, but they are not sufficient to establish that a present use remains ethically justified. We define ethical continuity as the persistence of normatively relevant relationships between the original conditions of data generation or material collection and subsequent downstream uses, such that current uses remain justifiable in light of the expectations, permissions, meanings, and relational obligations present at entrustment. We then define the Continuity Trap as a review-stage governance error in which a salient signal of continuity in one domain is treated as sufficient evidence of ethical continuity overall, causing inquiry into the remaining domains to close prematurely. The trap is not ordinary noncompliance, ethics creep, or a demand for universal rereview; it is a cross-domain inference error that can arise even in careful, good-faith review. We distinguish it from proxy closure, of which it is a continuity-specific subtype, and from Goodhart’s and Campbell’s laws, which describe how measures degrade once they become targets. We operationalize ethical continuity across 4 domains: provenance, semantics, authorization, and relational standing, developed in our Representational Veracity framework, and we show that these domains can diverge as data are linked, transformed, modeled, and redeployed. We identify the institutional mechanisms—provenance privilege, descriptor sedimentation, authorization fossilization, and community effacement—that cause auditable signals to be overread, and we examine how the US Health Insurance Portability and Accountability Act (HIPAA) of 1996, the General Data Protection Regulation, the European Health Data Space, US Food and Drug Administration guidance, the US National Institute of Standards and Technology (NIST) AI Risk Management Framework, and federated-learning governance can reduce risk while still inducing continuity traps. We apply the framework to consent and nonconsent settings, including public health, immunization, syndromic, and wastewater surveillance, polygenic risk scores, induced pluripotent stem cells, federated learning, and health-related large language models. The policy implication is trigger-based continuity review: rather than rereviewing every reuse, investigators and reviewers should identify the weakest continuity domain at the present data stage and impose a domain-matched safeguard, recorded in a short continuity statement. This reframing is intended for the committees, repositories, funders, and governance bodies that decide whether reuse may proceed, and it matters most in cross-border and low-resource settings. Provenance should begin ethical review; it should not end it.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/fbfe965a9496413a6b4cfa09311eb6b8" />
		
		<published>2026-08-19T08:30:03-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e92764 </id>
		<title>Digital Physical Exercise Interventions for Cognitive Functions in Older Adults: Systematic Review and Bayesian Network Meta-Analysis of Randomized Controlled Trials</title>
		<updated>2026-08-18T16:30:15-04:00</updated>

					<author>
				<name>Qing Hu</name>
			</author>
					<author>
				<name>Siyi Sun</name>
			</author>
					<author>
				<name>Jingru Zhu</name>
			</author>
					<author>
				<name>Xiaoke Zhong</name>
			</author>
					<author>
				<name>Shengyu Dai</name>
			</author>
					<author>
				<name>Changhao Jiang</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e92764" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e92764">Background: Cognitive decline in older adults imposes a major global burden, with physical inactivity a leading modifiable risk factor for dementia. Digital physical exercise interventions offer scalable alternatives to traditional programs, but comparative effectiveness across cognitive domains remains unclear. Objective: The aim of the study is to compare and rank 4 digital physical exercise interventions—immersive virtual reality exercise (IVR_E), nonimmersive exergame (NI_ExG), remote exercise (RE), and virtual reality exercise combined with cognitive training (VR_EC)—against routine intervention (RI) or nonintervention (NI) on global cognition, executive function, and memory function in older adults aged ≥60 years. Methods: Eligible studies were English-language randomized controlled trials evaluating digital physical exercise on any untrained cognitive outcome in older adults. In total, 6 databases (PubMed, Embase, Web of Science, CENTRAL, PsycINFO, and CINAHL) and 2 trial registries were searched from January 2010 to April 2026; reference lists were screened. Screening, data extraction, and risk-of-bias assessment were conducted independently in duplicate using the Cochrane tool. Bayesian network meta-analyses were performed in R, reporting standardized mean differences (SMDs) with 95% CIs and prediction intervals (PIs); surface under the cumulative ranking curve (SUCRA) ranked interventions, and CINeMA (Confidence in Network Meta-Analysis) assessed certainty of evidence, following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 and PRISMA-NMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Network Meta-Analyses). Results: In total, 51 randomized controlled trials (3673 participants) were included. NI_ExG significantly outperformed both NI (SMD 0.51, 95% CI 0.26-0.78; PI −0.03 to 1.07) and RI (SMD 0.32, 95% CI 0.12-0.53; PI −0.20 to 0.86) for global cognition (30 studies); IVR_E also outperformed NI (SMD 0.74, 95% CI 0.11-1.36; PI −0.04 to 1.53) and ranked highest by SUCRA. For executive function (33 studies), only NI_ExG versus NI was significant (SMD 0.39, 95% CI 0.04-0.76; PI −0.48 to 1.29). For memory function (20 studies), RE was significantly superior to both NI (SMD 1.30, 95% CI 0.15-2.44; PI −0.05 to 2.70) and RI (SMD 1.22, 95% CI 0.15-2.28; PI 0.07-2.56), with its PI also excluding the null; VR_EC was significantly inferior to RE (SMD −1.54, 95% CI −2.89 to −0.24; PI −3.09 to −0.07). Cumulative training ≥1000 minutes was associated with more stable memory benefit. Certainty was very low for most comparisons, downgraded for unclear allocation concealment, heterogeneity, and suspected reporting bias in executive function. Conclusions: This Bayesian network meta-analysis compares 4 digital physical exercise categories against active and passive controls across 3 cognitive domains. Comparative effectiveness was domain-specific: NI_ExG most consistently benefited global cognition and executive function; RE produced the only statistically robust memory improvement; and IVR_E achieved the highest rankings but requires confirmatory trials. The scalability of RE and NI_ExG makes them practical for older adults in rural and resource-limited settings, providing an evidence base to inform clinical guidelines and digital health investment. Trial Registration: PROSPERO CRD420251030142; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251030142</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/aa4a4db046562b0595de15969271b6e3" />
		
		<published>2026-08-18T16:30:15-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e83343 </id>
		<title>The Impact of Digitalization on European General Practice From the Perspective of General Practitioners: Systematic Review</title>
		<updated>2026-08-18T15:30:15-04:00</updated>

					<author>
				<name>Julia Magdalena Fuger</name>
			</author>
					<author>
				<name>Lisa Niehoff</name>
			</author>
					<author>
				<name>Erika Zelko</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e83343" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e83343">Background: Digital health technologies, including telemedicine, electronic medical records, digital health tools (DHTs), and AI, are transforming European primary care, but evidence on how general practitioners (GPs) experience these tools remains fragmented. Objective: We aimed to synthesize European GPs’ perspectives on the adoption and implementation of digital health technologies, focusing on perceived usefulness, ease of use, implementation barriers, workload, and doctor-patient relationships. Methods: We conducted a systematic review following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 and Synthesis Without Meta-Analysis guidance. Eight databases, PubMed, Scopus, SpringerLink, Wiley Online Library, Cochrane Library, Web of Science, Google Scholar, and Ovid, along with Elicit AI, were searched for primary studies on GPs’ perspectives in European primary care, published in English or German between January 2014 and December 2024; a search update extended coverage to October 2025. Eligible qualitative, quantitative, and mixed methods studies reported GP-perceived adoption, implementation, or impact of digital tools. Risk of bias was assessed using the Mixed Methods Appraisal Tool and JBI (Joanna Briggs Institute) critical appraisal tools. Data were synthesized using a narrative synthesis framework by Popay et al and a combined technology acceptance model and normalization process theory (NPT) effect-direction approach. Results: Of 1580 references, 65 studies including 24,994 GPs across 14 European countries (61/65, 93.8% from Western and Northern Europe) met the inclusion criteria. Studies covered AI or clinical decision support (n=7), DHTs (n=36), electronic medical records (n=4), and telemedicine (n=18). Perceived usefulness was high (56/65, 86.2% positive), whereas perceived ease of use was universally poor (0/65, 0% positive; 43/65, 66.2% negative), while actual use remained near universal (63/65, 96.9%). NPT constructs showed partial normalization: coherence 50.8% (33/65) positive; cognitive participation 4.6% (3/65) positive (22/65, 33.8%, negative); collective action 100% (65/65) partial; reflexive monitoring 66.2% (43/65) positive. Cross-cutting barriers included poor interoperability, fragmented platforms, limited training, misaligned incentives, and unreimbursed digital workload. GPs reported improved access for some patient groups but with concerns about depersonalized care and digital exclusion among older, low-literacy, rural, and socioeconomically disadvantaged patients. Conclusions: European general practice is extensively digitized but only partially normalized. GPs use digital tools because they must, not because they find them easy or workflow-enhancing. This review is, to our knowledge, the first to systematically synthesize European GPs’ experiences of digitalization across multiple technologies using a combined technology acceptance model-NPT effect-direction approach. It identifies a quantifiable “implementation paradox” in which high perceived usefulness and near-universal use coexist with poor usability and only partial workflow integration. These findings highlight the need for interoperable, user-centered systems that replace rather than add tasks, reimburse digital workload, provide structured digital-competency training, and set governance frameworks that safeguard equity and the quality of remote care in everyday general practice.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/0362a2479ef0799e484be2dcfe3f01d5" />
		
		<published>2026-08-18T15:30:15-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e96600 </id>
		<title>A Hybrid Care Intervention for High-Risk Patients With Chronic Obstructive Respiratory Disorders: Mixed Methods Co-Design Study</title>
		<updated>2026-08-18T14:30:16-04:00</updated>

					<author>
				<name>Alba Gómez-López</name>
			</author>
					<author>
				<name>Núria Sánchez-Ruano</name>
			</author>
					<author>
				<name>Marta Sorribes</name>
			</author>
					<author>
				<name>Rubèn González-Colom</name>
			</author>
					<author>
				<name>Marina Paredes</name>
			</author>
					<author>
				<name>Jin-Seok You</name>
			</author>
					<author>
				<name>María José Gordillo</name>
			</author>
					<author>
				<name>Emili Vela</name>
			</author>
					<author>
				<name>Gerard Carot-Sans</name>
			</author>
					<author>
				<name>Alba Jiménez</name>
			</author>
					<author>
				<name>Xabier Michelena</name>
			</author>
					<author>
				<name>Ramon Farré</name>
			</author>
					<author>
				<name>Josep Roca</name>
			</author>
					<author>
				<name>Isaac Cano</name>
			</author>
					<author>
				<name>Alicia Aguado</name>
			</author>
					<author>
				<name>Jose Fermoso</name>
			</author>
					<author>
				<name>Néstor Soler</name>
			</author>
					<author>
				<name>Ebymar Arismendi</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e96600" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e96600">Background: Community-based management of exacerbations in high-risk patients with chronic obstructive respiratory diagnoses remains a major challenge. Hybrid care interventions, combining digital support with in-person, patient-centered care, have shown efficacy to reduce unplanned hospitalizations in controlled trials. However, an efficacy-effectiveness gap remains, indicating the complexities of its deployment and sustainable adoption in real-world scenarios. Objective: This study aimed to co-design the core components of a hybrid care intervention for the preventive management of exacerbations in high-risk patients with chronic obstructive respiratory conditions, generating insights to guide sustainable adoption in routine clinical practice. Methods: Four plan-do-study-act (PDSA) co-design cycles were conducted, using a convergent mixed methods approach, during the 2-year follow-up (2024‐2025) of a cohort of 205 high-risk patients pertaining to two different clinical programs: (1) the community-based program included multimorbid patients with chronic obstructive pulmonary disease (COPD), asthma, or bronchiectasis from the Integrated Health District of Barcelona-Esquerra (AISBE, 520 k citizens) and (2) the Severe Asthma Program included patients with severe asthma. In all cases, patients were managed following the corresponding disease-specific consensus guidelines. The specific aims of each PDSA cycle were: PDSA-1, patients’ profiling and applicability of technological tools; PDSA-2, definition of clinical aspects of the hybrid care intervention and refinement of technology; PDSA-3, evaluation of a consolidated version of the hybrid care intervention; and PDSA-4, final refinements. Results: At the end of PDSA-3 (August 2025), the operationalization of the three core components of hybrid care—(1) nurse-led in-person care, (2) personalization of the intervention, and (3) advanced digital support—was achieved. The main study outcome was established through consensus among all key stakeholders on two aspects: (1) applicability of the hybrid care intervention for management of these patients in clinically stable conditions and during exacerbations in the real-world setting and (2) a well-defined strategy for its short-term deployment and sustainable site adoption. Conclusions: The intervention rollout requires emphasis on (1) alignment with local care pathways and information systems, (2) clear role definition and escalation procedures across care tiers, and (3) adaptation of the digital layer to patients’ capabilities, including pragmatic support for those with limited digital literacy. The co-design process enabled the operationalization of a hybrid care intervention integrating nurse-led management, personalization of care, and advanced digital support. Stakeholders reached consensus regarding its applicability and implementation strategy. Future real-world implementation studies are needed to evaluate its effects on clinical outcomes, health care use, patient experience, and health care value generation.</summary>
		
        
        
		<published>2026-08-18T14:30:16-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e93466 </id>
		<title>Value of AI in Critical Care Using Real-World Evidence on Intensive Care Unit Mortality Prediction: Cost-Utility Analysis</title>
		<updated>2026-08-18T14:30:16-04:00</updated>

					<author>
				<name>Gyeong Min Lee</name>
			</author>
					<author>
				<name>Joo-Yun Won</name>
			</author>
					<author>
				<name>Eun Young Cho</name>
			</author>
					<author>
				<name>Ji-Hyun Kim</name>
			</author>
					<author>
				<name>Kyung Soo Chung</name>
			</author>
					<author>
				<name>Kwang Joon Kim</name>
			</author>
					<author>
				<name>Seok Jin Yun</name>
			</author>
					<author>
				<name>Hyun Jun Lee</name>
			</author>
					<author>
				<name>Jae-Hyun Kim</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e93466" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e93466">Background: AI-based models for predicting mortality have shown potential for intensive care unit (ICU) patients, but evidence regarding their cost-effectiveness remains limited. Objective: This study aimed to evaluate the population-level cost-utility of an AI-based mortality prediction strategy activated during ICU admission in Korea. Methods: A lifetime Markov model followed a hypothetical cohort of Korean adults from age 19 years in the general-population state to capture ICU admissions, including recurrent ICU admissions, occurring over the lifetime horizon. AI-based mortality risk monitoring and its implementation cost were applied only when an individual entered the ICU state. Model inputs were derived from national claims data, published utility estimates, and cost sources. Outcomes were expressed as incremental cost-effectiveness ratios (ICERs) in Korean won (KRW) per quality-adjusted life year (QALY). Deterministic and probabilistic sensitivity analyses were conducted. Results: The AI-assisted strategy yielded an ICER of 8.37 million KRW (approximately US $6200) per QALY, below the societal willingness-to-pay (WTP) threshold of 40 million KRW (approximately US $29,600) per QALY in Korea. The incremental costs and QALYs were lifetime expected values per member of the general-population starting cohort rather than outcomes per directly monitored ICU admission. Probabilistic sensitivity analysis showed an 83.94% probability of cost-effectiveness at the threshold. One-way sensitivity analysis identified the true-positive rate, AI implementation cost, and postdischarge well-state utility value as the most influential parameters. Additional claims-based analyses indicated that ICU transition pathways and patient age were strongly associated with survival outcomes. Conclusions: Under modeled assumptions, AI-assisted ICU mortality prediction showed potential cost-effectiveness compared with usual care in the Korean critical care setting. These findings should be interpreted as decision-analytic estimates rather than direct evidence that the AI system reduces ICU mortality in real patients. Prospective real-world evaluations are needed to confirm clinical effectiveness and implementation value.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/be23dd0c52b31e2f9230fa782154edca" />
		
		<published>2026-08-18T14:30:16-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e95043 </id>
		<title>Multimodal Digital Therapeutics Enhanced by Task Design and AI for Attention-Deficit/Hyperactivity Disorder Core Symptoms and Executive Functions in Children and Adolescents: Systematic Review and Network Meta-Analysis of Randomized Controlled Trials</title>
		<updated>2026-08-18T14:00:03-04:00</updated>

					<author>
				<name>Zhixiang Hao</name>
			</author>
					<author>
				<name>Hongli Xu</name>
			</author>
					<author>
				<name>Chen Wang</name>
			</author>
					<author>
				<name>Bingjie Wang</name>
			</author>
					<author>
				<name>Zhengang Qiu</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e95043" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e95043">Background: Attention-deficit/hyperactivity disorder (ADHD) is a prevalent neurodevelopmental disorder in children and adolescents. Digital therapeutics (DTx) show promise as nonpharmacological interventions, but the comparative efficacy of different DTx modalities remains unclear. Objective: This network meta-analysis (NMA) compared 4 DTx modalities (single-task, cognitive-motor dual-task, AI-integrated single-task, and AI-integrated cognitive-motor dual-task DTx) on core ADHD symptoms and executive functions, identified the optimal modality, and explored treatment moderators. Methods: We included randomized controlled trials (RCTs) in children and adolescents aged 4 to 17 years with ADHD diagnosed per the DSM-5 (Diagnostic and Statistical Manual of Mental Disorders [Fifth Edition]) or ICD-10 (International Classification of Diseases, Tenth Revision). We searched PubMed/MEDLINE, PsycINFO, Web of Science, EMBASE, Scopus, ProQuest Dissertations and Theses, Cochrane Library, and ClinicalTrials.gov (gray literature) to identify trials published between January 2000 to May 2026 (last search May 22, 2026) without language restrictions, supplemented by snowballing. Risk of bias was assessed with the Cochrane Risk of Bias (RoB) 2 tool. Data were synthesized using Bayesian NMA with random-effects models. The surface under the cumulative ranking curve (SUCRA) was used to rank interventions. Heterogeneity was evaluated via 95% prediction intervals (95% PI) and explored through subgroup analyses and meta-regression. Small-study effects were assessed using Egger test, and sensitivity analyses were also performed. Results: Thirty-two RCTs (2819 patients) were included. The risk of bias assessment identified a low risk in 37.5% of the studies, some concerns in 21.9% of the studies, and a high risk in 40.6% of the studies, mainly due to inadequate reporting of randomization or blinding. AI-integrated cognitive-motor dual-task DTx ranked first for all outcomes in Bayesian network meta-analysis. For the Attention-Deficit/Hyperactivity Disorder-Rating Scale (ADHD-RS; 7 studies, n=1642), SUCRA was 57.5% (mean difference [MD] –3.03, 95% credible intervals [95% CrI] –5.59 to –0.47). For the Swanson, Nolan, and Pelham Rating Scale (Version IV; SNAP-IV) inattention subscale (SNAP-IV-PI; 8 studies, n=468), SUCRA was 82.5% (MD –5.58, 95% CrI –8.76 to –2.39); for the SNAP-IV hyperactivity-impulsivity subscale (SNAP-IV-PHI; 8 studies, n=468), SUCRA was 92.6% (MD –6.84, 95% CrI –10.37 to –3.31). For the Behavior Rating Inventory of Executive Function (BRIEF; 23 studies, n=1927), SUCRA was 84.4% (MD –7.75, 95% CrI –10.06 to –5.43). In pairwise meta-analyses, the 95% PI for ADHD-RS did not cross zero (−7.19 to −0.11), whereas those for the SNAP-IV (PI subscale: −5.62 to 1.87; PHI subscale: −6.66 to 2.82) and BRIEF (−6.91 to 1.94) did, indicating limited generalizability and substantial between-study heterogeneity. Subgroup analyses suggested intervention duration as a heterogeneity source for the SNAP-IV (both subscales) and BRIEF and mean age as a heterogeneity source for the SNAP-IV-PI and BRIEF. Conclusions: This NMA provides the first dual-dimension classification framework for ADHD DTx, combining SUCRA ranking, PI, and GRADE (Grading of Recommendations Assessment, Development and Evaluation). AI-integrated cognitive-motor dual-task DTx had the highest probability of improving core symptoms and executive functions, with duration and age as potential heterogeneity sources. These findings inform clinical decision-making and DTx development, although interpretation should account for evidence limitations. Trial Registration: PROSPERO CRD420261304236; https://www.crd.york.ac.uk/PROSPERO/view/CRD420261304236</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/769838a5822ced1c086c7ee970915e3f" />
		
		<published>2026-08-18T14:00:03-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e91756 </id>
		<title>NeuroSift for Task-Aware Quality Assurance of Multimedia Data in Remote Parkinson Disease Assessment: Machine Learning Model Development and Validation Study</title>
		<updated>2026-08-18T13:30:30-04:00</updated>

					<author>
				<name>Md Saiful Islam</name>
			</author>
					<author>
				<name>Sooyong Park</name>
			</author>
					<author>
				<name>Evelyn Xiaoxiao Ma</name>
			</author>
					<author>
				<name>Tariq Adnan</name>
			</author>
					<author>
				<name>Ehsan Hoque</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e91756" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e91756">&lt;strong&gt;Background:&lt;/strong&gt; Automated multimedia analysis of remotely recorded tasks offers a scalable approach to screening and remote monitoring of movement disorders such as Parkinson disease (PD). However, unsupervised recordings often suffer from quality issues that compromise model reliability. General multimedia quality checks may not detect task-specific failures, such as poor hand visibility during finger-tapping, inadequate facial framing during smile tasks, or background noise during speech tasks. &lt;strong&gt;Objective:&lt;/strong&gt; This study aimed to develop and evaluate NeuroSift, a task-aware, interpretable machine learning framework for assessing recording quality and task compliance in home-recorded multimedia data collected for remote PD assessment. &lt;strong&gt;Methods:&lt;/strong&gt; We analyzed 2516 home-recorded audio and video segments from 3 tasks: finger-tapping, facial expression (smile), and speech (pangram utterance). Three experts rated recordings as poor, borderline, or good quality. Task-specific annotation guidelines were developed through iterative review and discussion. We extracted interpretable features aligned with observable quality and compliance criteria and trained task-specific quality classification models. Interrater reliability was evaluated before and after guideline implementation using quadratic weighted Cohen κ (QWK), pairwise agreement, complete agreement, and intraclass correlation. Model performance was evaluated on a held-out test set (labeled using expert consensus) using accuracy, QWK, and ordinal classification accuracy (OCA, an accuracy metric that accounts for the ordered relationship among poor, borderline, and good labels). Feature importance was examined using Shapley additive explanations (SHAP), a method for estimating how individual features contribute to model predictions. &lt;strong&gt;Results:&lt;/strong&gt; The task-specific guidelines significantly improved interrater reliability across all tasks (&lt;i&gt;P&lt;/i&gt;&amp;lt;.001)—QWK increased from 0.46 to 0.89 for finger-tapping, from 0.61 to 0.84 for smile, and from 0.64 to 0.90 for speech. For 3-class quality classification, the best-performing models achieved QWK values of 0.71 for finger-tapping, 0.56 for smile, and 0.72 for speech. OCA was 82.1% for finger-tapping, 76.9% for smile, and 89.9% for speech. Most errors were adjacent-class errors, such as classifying poor recordings as borderline, whereas severe errors, such as classifying good recordings as poor, were rare. SHAP analyses identified task-specific sources of quality degradation that aligned with the annotation guidelines. &lt;strong&gt;Conclusions:&lt;/strong&gt; NeuroSift was evaluated as a task-aware quality-classification framework for identifying low-quality or noncompliant multimedia data prior to downstream PD assessment. By combining structured quality guidelines, interpretable features, and explainable machine learning models, NeuroSift supports more transparent and user-correctable remote data collection. Although this study did not test downstream clinical impact, the framework offers a practical approach for quality-aware workflows that rely on user-recorded audio or video. This approach may extend beyond PD assessment to other remote settings, including telehealth, rehabilitation, and digital recruitment. Future studies should evaluate NeuroSift on independent datasets and examine its effects on downstream model performance, fairness, user experience, and workflow integration. &lt;strong&gt;Trial Registration:&lt;/strong&gt; </summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/fcb6989ea0be4fb387dc14626c243f7e" />
		
		<published>2026-08-18T13:30:30-04:00</published>
	</entry>
</feed>