<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
	<id>https://www.jmir.org/issue/feed</id>
	<title>Journal of Medical Internet Research</title>
			<updated>2025-01-01T11:30:03-05:00</updated>
	
		<author>
		<name>JMIR Publications</name>
				<email>editor@jmir.org</email>
			</author>
		<link rel="alternate" href="https://www.jmir.org" />
	<link rel="self" type="application/atom+xml" href="https://www.jmir.org/feed/atom" />

	<generator uri="http://pkp.sfu.ca/ojs/" version="2.2.0.0">Open Journal Systems</generator>

				    	<subtitle> The leading peer-reviewed journal for digital medicine and health and health care in the internet age.&amp;nbsp; </subtitle>



	<entry>
		<id> https://www.jmir.org/2026/1/e91807 </id>
		<title>Digital Gaze and Vicarious Trauma Among Intensive Care Unit Nurses in Alarm-Monitoring Ecologies: Qualitative Interview Study</title>
		<updated>2026-08-04T16:15:12-04:00</updated>

					<author>
				<name>Yuanyuan Wang</name>
			</author>
					<author>
				<name>Hongyu Chen</name>
			</author>
					<author>
				<name>Kui Fang</name>
			</author>
					<author>
				<name>Xu Guo</name>
			</author>
					<author>
				<name>Yueqin Gu</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e91807" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e91807">Background: Intensive care units (ICUs) rely on continuous physiological monitoring and frequent alarms to detect patient deterioration. Although alarm fatigue has been widely discussed as a patient safety and workflow issue, less is known about how monitoring systems shape nurses’ attention, visibility, perceived accountability, emotional strain, and recovery after distressing events. Understanding these experiences is important for designing safer monitoring displays, alarm behavior, communication routines, and future AI-supported systems. Objective: This study aimed to explore how ICU nurses experience continuous monitoring and alarms as a digital work environment, with particular attention to digital gaze, vicarious trauma–related emotional strain, perceived accountability, and design-relevant system needs. Methods: We conducted a qualitative interview study with 15 ICU nurses from a tertiary hospital in Hangzhou, China. Semistructured interviews probed 6 topics: everyday monitoring routines and alarm exposure; responses to alarms, patient deterioration, and death; perceived pressure related to visible physiological data; after-shift experiences following distressing events; coping and recovery strategies; and suggestions for alarm governance and monitoring system design. Data were analyzed using reflexive thematic analysis. Results: Four themes were generated. First, continuous monitoring created a form of digital gaze in which nurses maintained constant watch, experienced alarm-driven interruptions, and felt that visible physiological data made bedside responses open to scrutiny. Second, patient deterioration and death were experienced partly through monitoring technologies, including weakening waveforms, escalating alarms, numerical decline, and eventual silence. Third, monitoring-related stress extended beyond the ICU through lingering alarm sounds, monitor images, personal resonance, and work-related messages after shifts. Fourth, participants described recovery strategies but also emphasized system-level needs, including clearer alarm prioritization, fewer nonactionable alerts, gentler auditory design, more useful trend displays, postresuscitation buffering, and better digital communication boundaries. Conclusions: Continuous monitoring and frequent alarms shaped ICU nurses’ work beyond workflow disruption and patient safety. Alarm-intensive care should be understood as a digital health and sociotechnical design issue that affects attention, perceived accountability, emotional strain, and recovery. Future monitoring systems should be co-designed and evaluated not only for alarm reduction and technical accuracy but also for how displays, alarm behavior, work-related communication, and AI-supported tools affect clinicians’ work, recovery, and perceived surveillance.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/8564e8526d7bb703a646682619dbc9d3" />
		
		<published>2026-08-04T16:15:12-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e81095 </id>
		<title>Temporal Dynamics of Internal vs External Entrapment in Suicidal Ideation in Psychotherapy Outpatients: Prospective Longitudinal Cohort Study Using Ecological Momentary Assessment</title>
		<updated>2026-08-04T15:30:16-04:00</updated>

					<author>
				<name>Emmy Wichelhaus</name>
			</author>
					<author>
				<name>Dajana Schreiber</name>
			</author>
					<author>
				<name>Inken Höller</name>
			</author>
					<author>
				<name>Thomas Forkmann</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e81095" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e81095">Background: The Integrated Motivational-Volitional Model posits entrapment as a central driver of suicidal ideation (SI), yet little is known about how internal (IE) vs external entrapment (EE) unfolds in real time. Prior ecological momentary assessment (EMA) studies have generally treated entrapment as unidimensional and used sampling intervals of several hours, potentially obscuring rapid risk processes. Objective: This study examined (1) cross-sectional and longitudinal temporal associations between IE, EE, and SI across lags of 1‐3 prompts; (2) differences in immediacy of IE vs EE effects; and (3) whether the exact time interval (in minutes) between assessments moderates these associations. Methods: A total of 91 (62.2% female; M=31.6, SD 10.6) adults initiating outpatient psychotherapy (convenience sample) completed a 7-day EMA protocol with 10 daily prompts (30‐120 min intervals). Participants provided informed consent, and ethical approval was obtained. A total of 5414 (85% compliance) observations were analyzed using multilevel models with maximum likelihood estimation (assuming missing at random). SI (4 items) and entrapment (IE, EE; 1 item each) were rated on 5-point scales. Lagged effects (up to ≈6 h) controlled for prior SI, and time intervals (in minutes) were tested as moderators. Results: SI varied primarily within persons (intraclass correlation=0.15). Cross-sectionally, both IE (est=0.16, 95% CI 0.08 to 0.24; &lt;.001) and EE (est=0.19, 95% CI 0.09 to 0.27; &lt;.001) were significantly associated with SI, explaining 25.3% of within-person variance (Quasi ²). Lagged analyses revealed distinct temporal patterns. At Lag 1 (30‐120 min), IE predicted increases in SI (est=0.13, 95% CI 0.09 to 0.18; &lt;.001), whereas EE did not (est=0.01, 95% CI –0.03 to 0.06; =.60). At lag 2 (60‐240 min), EE predicted SI (est=0.05, 95% CI 0.01 to 0.10; &lt;.05), whereas IE was not significant. At lag 3 (90‐360 min), IE negatively predicted SI (est=–0.06, 95% CI –0.11 to –0.01; &lt;.05), while EE positively predicted SI (est=0.05, 95% CI 0.00 to 0.10; &lt;.05). Prior SI consistently predicted subsequent SI (est range=0.38‐0.42; &lt;.001). Time intervals did not moderate effects ( values&gt;.12). Conclusions: Entrapment is a dynamic correlate of SI, with IE and EE showing distinct temporal patterns. IE has rapid, within-minute effects on SI, whereas EE shows delayed associations. This study is innovative in modeling minute-level dynamics using high-frequency EMA. Unlike prior unidimensional, lower-frequency designs, it captures ultra-short-term risk processes and clarifies temporally distinct pathways within the Integrated Motivational-Volitional Model and enhances theoretical precision regarding proximal vs distal drivers of SI. Clinically, results suggest that interventions may benefit from targeting IE in real time.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/4bad878134cd4b4efa72e99ba1b12f24" />
		
		<published>2026-08-04T15:30:16-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e92373 </id>
		<title>The Scale for AI Literacy in Health Care Workers: Development and Validation</title>
		<updated>2026-08-04T15:30:16-04:00</updated>

					<author>
				<name>Chin-Siang Ang</name>
			</author>
					<author>
				<name>Sakura Ito</name>
			</author>
					<author>
				<name>Saumya Bajaj</name>
			</author>
					<author>
				<name>Minyang Chow</name>
			</author>
					<author>
				<name>Jennifer Cleland</name>
			</author>
					<author>
				<name>Jonty Heaversedge</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e92373" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e92373">Background: AI is increasingly embedded in health care systems; yet, validated instruments for assessing AI literacy among health care workers remain limited. Existing measures are often designed for students or general populations and may not adequately reflect competencies required in health care practice. Objective: This study aimed to develop and validate the Scale for AI Literacy in Health Care Workers (SAIL-HCW), a new instrument designed to assess AI literacy across domains relevant to health care practice. Methods: A 3-phase instrument development study was conducted. In Phase 1, conceptual domains were identified through a literature review, and an initial item pool was generated. In Phase 2, content validity was assessed by 4 subject-matter experts, and face validity was evaluated with 26 health care workers. Feedback from both groups informed item refinement. In Phase 3, psychometric testing was conducted using survey data from health care workers in a single health care organization. A total of 425 participants completed the survey. The dataset was randomly split into 2 subsamples for exploratory factor analysis (n=212) and confirmatory factor analysis (n=213). Model fit was evaluated using unidimensional, correlated-factor, higher-order, and bifactor models. Reliability was assessed using Cronbach alpha and McDonald omega. Item performance was examined using corrected item-total correlations (CITC), item discrimination analysis, and inter-item correlations. Construct validity was assessed using prior AI training, frequency of AI use, and self-rated AI literacy. Results: Phase 2 feedback from experts and health care workers supported the proposed domain structure and informed item refinement, including revision of wording and removal of redundant items. The final SAIL-HCW consists of 14 items across 7 domains, including AI concept, data fluency, AI evaluation, AI in practice, ethics and regulation, AI in system, and continuous learning. In Phase 3, the bifactor model showed the best fit compared with alternative models (comparative fit index and Tucker-Lewis index&gt;0.93; root-mean-square error of approximation&lt;0.06; standardized root-mean-square residual&lt;0.05), indicating a general AI literacy factor alongside domain-specific factors. Internal consistency for the total scale was high (Cronbach α=0.937; =0.938). Domain-level reliability ranged from 0.635 to 0.797. All items significantly discriminated between high- and low-scoring groups (&lt;.001), with CITC values ranging from 0.570 to 0.785. Construct validity was supported, with higher SAIL-HCW scores observed among participants with prior AI training, higher frequency of AI use, and higher self-rated AI literacy (all &lt;.001). Conclusions: The SAIL-HCW provides initial evidence of validity and reliability for assessing AI literacy among health care workers. Findings suggest that AI literacy may be represented as a general construct with additional domain-level components. The scale may be useful for research and educational evaluations, although further validation in other settings is required.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/37b4dd0f9aed15c29c1d11febe9ed057" />
		
		<published>2026-08-04T15:30:16-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e93618 </id>
		<title>Cognitive Workload and Mental Burden in Health Care Professionals Interacting With AI: Systematic Review and Meta-Analysis</title>
		<updated>2026-08-04T10:00:24-04:00</updated>

					<author>
				<name>Eun Jeong Gong</name>
			</author>
					<author>
				<name>Chang Seok Bang</name>
			</author>
					<author>
				<name>Jae Jun Lee</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e93618" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e93618">Background: AI adoption in health care has accelerated rapidly, with ambient documentation tools, diagnostic imaging AI, and clinical decision support systems (CDSSs) entering routine practice. However, the cognitive demands placed on clinicians supervising these systems remain understudied. Specifically, the concept of verification burden requires closer examination. Consequently, institutional decision-makers lack a structured, certainty-graded evidence base regarding the true impact of AI on clinician workload and burnout. Objective: This study aimed to systematically review evidence on cognitive workload and burnout in health care professionals that use AI-powered clinical tools, quantify pooled effects under a conservative inferential framework, and assess certainty of evidence by AI category. Methods: The study was registered in PROSPERO (CRD420261284298) and reported per PRISMA 2020 and PRISMA-S guidelines. We searched MEDLINE, Embase, Web of Science, and Cochrane CENTRAL (January 2015-2026) for studies measuring cognitive workload or burnout using validated instruments (NASA Task Load Index [NASA-TLX] and Professional Fulfillment Index [PFI]) among health care professionals using clinical AI. Risk of bias was assessed using ROB 2.0 and ROBINS-I; certainty was rated using GRADE. Meta-analyses applied Hartung-Knapp-Sidik-Jonkman adjustment with restricted maximum likelihood estimation, incorporating prediction intervals (PIs). Results: We included 21 studies representing 2885 health care professionals across 7 countries. The synthesis demonstrated that the cognitive impact of clinical AI varies according to its specific application. Pooled analyses of ambient AI documentation showed statistically significant reductions in NASA-TLX temporal demand (SMD −1.46, 95% CI −2.81 to −0.11; k=2; =31.1%) and effort (SMD −1.29, 95% CI −2.16 to −0.42; k=2; =0%), PFI work exhaustion (MD −0.35, 95% CI −0.58 to −0.12; k=3; =0%; 95% PI −1.03 to 0.33), and burnout prevalence (OR 0.47, 95% CI 0.25-0.86; k=3; =0%; 95% PI 0.06-3.82). Two pools favored ambient AI but did not reach significance at k=2: NASA-TLX mental demand (SMD −1.29, 95% CI −3.64 to 1.07) and documentation time (SMD −0.24, 95% CI −1.10 to 0.61). Diagnostic imaging AI and CDSS showed mixed or paradoxically increased workload. GRADE certainty was moderate for cognitive workload reduction with ambient AI, low for burnout reduction with ambient AI, and very low for imaging AI and CDSS outcomes. Conclusions: This review combines validated workload instruments, meta-analysis, and PIs in health care AI, delivering a GRADE certainty assessment across 5 AI categories that prior accuracy- or efficiency-focused reviews have not provided. Ambient AI documentation was associated with reduced cognitive workload and burnout, but only in voluntary early-adopter cohorts and based on few studies; the conservative CIs were wide and, where estimable, PIs crossed the null. Findings inform institutional pilots with prospective workload measurement, regulatory human-factors evaluation of AI medical devices, and human-centered AI design. Net benefit on the health care workforce remains an open empirical question. Trial Registration: PROSPERO CRD420261284298;</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/d2de1de06cb866b263bde9526278a1d6" />
		
		<published>2026-08-04T10:00:24-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e90046 </id>
		<title>Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review</title>
		<updated>2026-08-03T16:00:27-04:00</updated>

					<author>
				<name>Yuhan Zhao</name>
			</author>
					<author>
				<name>Yiqun Miao</name>
			</author>
					<author>
				<name>Rongrong Guo</name>
			</author>
					<author>
				<name>Yuan Luo</name>
			</author>
					<author>
				<name>Huiying Wang</name>
			</author>
					<author>
				<name>Ying Wu</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e90046" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e90046">Background: Inference-time retrieval augmentation is increasingly used to improve the traceability and verifiability of large language model (LLM) applications in health care. Evaluation practices for text-based retrieval-augmented generation (RAG) and graph-structured RAG (GraphRAG) systems remain heterogeneous, which limits comparison across studies and complicates judgments about clinical readiness. Objective: This review mapped evaluation methods for inference-time retrieval-augmented and graph-structured retrieval-augmented LLM systems in health care and characterized how evaluation constructs are defined, operationalized, and reported across system layers and evaluation-setting categories. Methods: We conducted a scoping review in accordance with PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews), with search reporting informed by PRISMA-S (PRISMA literature search extension). Searches were conducted through May 14, 2026, in PubMed (MEDLINE), Web of Science Core Collection, IEEE Xplore, ACM Digital Library, arXiv, and medRxiv, with backward and forward citation tracking of included studies. Eligible records described health care–relevant LLM systems using inference-time RAG and reported at least 1 evaluation component. Data were charted on study characteristics, system design, retrieval-layer evaluation, evidence linkage, safety-related and GraphRAG-specific evaluation, and selected reporting and governance characteristics. We also constructed an evidence-and-gap map cross-classifying evaluation-setting categories with key evaluation domains. Results: A total of 157 studies met the inclusion criteria. Clinical question answering was the most frequently represented application (89/157, 56.7%), followed by clinical decision support (70/157, 44.6%). Most evaluations were conducted in offline-only settings (140/157, 89.2%), whereas 17/157 (10.8%) studies reported workflow-facing, prospective, or deployment-level evaluation. Independent retrieval-layer evaluation was reported in 47/157 (29.9%) studies. Grounding and faithfulness evaluation was reported in 41/157 (26.1%) studies, and fine-grained evidence verification was reported in 22/157 (14%) studies. Human evaluation was reported in 94/157 (59.9%) studies, but interrater reliability was reported in 26/94 (27.7%) studies. LLM-as-judge evaluation was reported in 41/157 (26.1%) studies, with bias-control measures reported in 15/41 (36.6%) studies. Formal safety-related evaluation was reported in 45/157 (28.7%) studies. Among 27 (17.2%) GraphRAG studies, intermediate-artifact evaluation was reported in 11/27 (40.7%) studies, and graph construction evaluation was reported in 6/27 (22.2%) studies. The evidence-and-gap map showed limited coverage of fine-grained verification, contradiction handling, safety evaluation, LLM-as-judge safeguards, GraphRAG construction evaluation, and GraphRAG intermediate-artifact evaluation in workflow-facing, prospective, or deployment-level settings. Conclusions: Evaluation of health care RAG and GraphRAG systems has expanded rapidly, yet reporting and operational definitions remain inconsistent across evaluation layers. Current evidence remains concentrated in offline evaluation, with limited workflow-facing, prospective, or deployment-level assessment of retrieval quality, fine-grained evidence linkage, safety, LLM-as-judge safeguards, GraphRAG construction quality, and GraphRAG intermediate artifacts. This review maps these gaps across evaluation-setting categories and translates them into synthesis-informed evaluation considerations. These findings suggest that future evaluation may need to move beyond end-to-end benchmark performance toward more transparent, layer-specific, safety-oriented, and clinically contextualized assessment before workflow-facing implementation. Trial Registration: OSF Registries mtf5x; https://osf.io/mtf5x/overview</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/76bf7d39a43672c77b88264589e66dab" />
		
		<published>2026-08-03T16:00:27-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e97772 </id>
		<title>Social Media Discourse on Breast Cancer Screening Barriers in Japanese and English: Cross-Sectional Infodemiology Study</title>
		<updated>2026-07-31T16:00:04-04:00</updated>

					<author>
				<name>Mitsuo Terada</name>
			</author>
					<author>
				<name>Rie Kawabori Tahara</name>
			</author>
					<author>
				<name>Yumi Wanifuchi-Endo</name>
			</author>
					<author>
				<name>Tomoko Asano</name>
			</author>
					<author>
				<name>Nanae Horisawa</name>
			</author>
					<author>
				<name>Kazuki Nozawa</name>
			</author>
					<author>
				<name>Ayaka Isogai</name>
			</author>
					<author>
				<name>Nari Kureyama</name>
			</author>
					<author>
				<name>Hikaru Kawahara</name>
			</author>
					<author>
				<name>Mika Kotani</name>
			</author>
					<author>
				<name>Yuya Tanaka</name>
			</author>
					<author>
				<name>Marie Mizumoto</name>
			</author>
					<author>
				<name>Atsushi Fushimi</name>
			</author>
					<author>
				<name>Nami Yamashita</name>
			</author>
					<author>
				<name>Madoka Iwase</name>
			</author>
					<author>
				<name>Asumi Iesato</name>
			</author>
					<author>
				<name>Tatsuya Toyama</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e97772" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e97772">Background: Despite clinical advances, breast cancer screening adherence remains stagnant in Japan (&lt;50%) compared with the United States (&gt;70%). Understanding distinct cross-cultural barriers is essential; however, traditional methodologies often fail to capture visceral, real-world individual experiences and hidden deterrents to screening. Objective: This study aims to characterize and compare cross-cultural informatics profiles of barriers to breast cancer screening across Japanese-language and English-language social media discourse. Methods: We developed an automated natural language processing pipeline on a cloud-based informatics platform to analyze 46,823 screening-related posts (30,027 in Japanese and 16,796 in English) from X (formerly Twitter) collected in 2025, derived from an initial 76,955 posts after noise exclusion. The methodology integrated large language model–assisted sentiment polarity scoring with strict negation-handling, co-occurrence network topology analysis, and advanced distributional visualizations, including raincloud and ridgeline plots. Subgroup comparisons (prescreening vs postscreening and ultrasound with vs without mammography) were evaluated. Results: Among the 46,823 screening-related posts, overall sentiment distributions showed no practically meaningful cross-cultural divergence (Cohen =0.049); however, domain-specific analyses revealed sharp disparities in barrier prevalence. Within the English-language cohort containing 16,796 posts, discourse exhibited a concentrated, moderate negative sentiment regarding systemic barriers, featuring “Cost” as the primary barrier (n=1239, 7.4%), while “Pain” ranked considerably lower (n=842, 5.0%). In contrast, the Japanese-language cohort containing 30,027 posts was heavily bottlenecked by psychosomatic barriers, governed by a tightly interconnected network of “Pain,” “Fear,” and “Appointment.” “Pain” emerged as the overwhelmingly dominant barrier (n=5788, 19.3%). In the English-language cohort, “Dense” breasts emerged as a prominent clinical topic due to elevated public awareness, distinct from the financial narrative. The Japanese subgroup analyses identified a temporal transition from anticipatory psychological anxiety (“Fear,” 739/3805, 19.4%) and logistical concerns (“Appointment,” 1556/3805, 40.9%) before screening to a strong persistence of the discomfort memory of “Pain” (1160/7419, 15.6%). Furthermore, sentiment scores for mammography were significantly more negative than those for ultrasound alone (&lt;.001, Cohen =0.264), and pain-related descriptors for ultrasound spiked from 5.0% (106/2110, ultrasound alone) to 18.2% (733/4017) when performed concurrently with mammography. Conclusions: Despite comparable overall emotional equilibrium, a fundamental dichotomy emerged. The English-language discourse predominantly reflects US-specific systemic financial burdens, whereas the Japanese experience is characterized by emotional volatility transitioning from anticipatory anxiety to a tightly interconnected “Pain-Fear-Appointment” network. Physical discomfort of mammography dominates the screening narrative, overshadowing concurrent painless modalities like ultrasound. Improving adherence in Japan requires individualized pain-mitigating compression protocols and optimized clinical workflows to decouple mammographic discomfort from supplemental screening, thereby preventing pain-associated defensive avoidance and reducing logistical hurdles to improve equitable access.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/8659a333fdf77e75bead1f8ec3f7cd06" />
		
		<published>2026-07-31T16:00:04-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e93378 </id>
		<title>Accuracy of Machine Learning Algorithms Based on Electroencephalogram in Sleep Apnea Detection: Systematic Review and Meta-Analysis</title>
		<updated>2026-07-31T16:00:04-04:00</updated>

					<author>
				<name>Xiangshuo Li</name>
			</author>
					<author>
				<name>Lulu Wang</name>
			</author>
					<author>
				<name>Ting Tang</name>
			</author>
					<author>
				<name>Yuanyuan Chen</name>
			</author>
					<author>
				<name>Lan Yang</name>
			</author>
					<author>
				<name>Hao Cai</name>
			</author>
					<author>
				<name>Chen Wang</name>
			</author>
					<author>
				<name>Shuxiao Zhang</name>
			</author>
					<author>
				<name>Ning Ding</name>
			</author>
					<author>
				<name>Kouying Liu</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e93378" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e93378">Background: Sleep apnea (SA) is a serious sleep disorder, and its diagnostic gold standard, polysomnography, is costly and time-consuming. Electroencephalogram (EEG) signals, due to their direct correlation with neural activity and ease of extraction, represent a promising tool. Despite increasing research on machine learning (ML) and deep learning for EEG-based SA detection, model performance has not been consistently evaluated. Objective: This systematic review evaluated the accuracy of ML in detecting SA from EEG data and provided an evidence base for further clinical application and future research. Methods: Following the PRISMA-DTA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy) and PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 expanded checklists, we systematically searched PubMed, Embase, Web of Science, Cochrane Library (CENTRAL), Scopus, IEEE Xplore, and ClinicalTrials.gov databases from inception to April 2026. Studies evaluating the value of ML algorithms for detecting SA based only on EEG data were included. The Quality Assessment of Diagnostic Accuracy Studies-2 and Prediction Model Risk of Bias Assessment Tool for Artificial Intelligence tools were used to assess the risk of bias in each study. Statistical analysis was performed using the and packages in R (version 4.6.0; R Foundation for Statistical Computing) and the Meta-DiSc (version 1.4; Hospital Ramón y Cajal) software. We used GRADE (Grading of Recommendations Assessment, Development and Evaluation) to evaluate the certainty of evidence. Results: A total of 27 retrospective studies were included. Segment-level analyses showed high diagnostic performance, with a pooled sensitivity of 0.90 (95% CI 0.85‐0.94; 95% prediction interval 0.43‐0.99) and specificity of 0.92 (95% CI 0.87‐0.95; 95% prediction interval 0.46‐0.99). The pooled area under the summary receiver operating characteristic curve was 0.95 (95% CI 0.92‐0.99). Meta-regression identified EEG channel configuration, region, and validation strategy as significant sources of heterogeneity (=.004, =.003, and =.046, respectively). Multichannel EEG, deep learning approaches, and hold-out validation strategies generally demonstrated better diagnostic performance. Only 2 studies evaluated patient-level diagnostic performance, which was summarized qualitatively. Conclusions: To our knowledge, this is the first systematic review and meta-analysis specifically focused on the diagnostic accuracy of EEG-based ML models in the detection of SA. This meta-analysis indicates that ML models based on EEG demonstrate good diagnostic accuracy in detecting SA at the segment level and show promise as tools for SA screening and clinical decision support. However, most current studies are retrospective segment-level analyses, which may overestimate the practical value of this technology in real-world clinical settings. To reliably integrate EEG-based ML models into clinical diagnostic workflows, further prospective studies incorporating full-night monitoring and patient-level validation are needed. Trial Registration: PROSPERO CRD420251244156; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251244156</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/df69e4d5e8c633a9a2b76eb594225b6b" />
		
		<published>2026-07-31T16:00:04-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e93875 </id>
		<title>Barriers and Facilitators to Implementing Digital Health Technologies for Remote Management of NCDs in Rural Areas: Mixed Methods Systematic Review</title>
		<updated>2026-07-31T15:45:10-04:00</updated>

					<author>
				<name>Serine Sahakyan</name>
			</author>
					<author>
				<name>Selai Akseer</name>
			</author>
					<author>
				<name>Lusine Abrahamyan</name>
			</author>
					<author>
				<name>Olivia Metcalf</name>
			</author>
					<author>
				<name>Sara Allin</name>
			</author>
					<author>
				<name>Richard Chenhall</name>
			</author>
					<author>
				<name>Emily Seto</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e93875" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e93875">Background: Digital health technologies (DHTs) have the potential to improve care delivery and outcomes for patients with noncommunicable diseases. Yet their implementation in rural settings remains uneven, and the factors influencing uptake are not well understood. Objective: This mixed methods systematic review aimed to identify barriers and facilitators influencing the implementation and use of DHTs for remote management of noncommunicable diseases in rural areas. Methods: We searched Medline, Embase, and CINAHL from inception to February 12, 2026, using terms related to digital health, noncommunicable diseases, and rural settings. Following the Joanna Briggs Institute methodology for mixed-method systematic review, we synthesized quantitative and qualitative studies. Barriers and facilitators were categorized using the Consolidated Framework for Implementation Research, and study quality was appraised using the Mixed Methods Appraisal Tool. Results: From the initial 1491 records, 14 studies met the inclusion criteria, with most conducted in high-income countries (n=11). Key barriers included technical challenges (software instability and hardware issues), poor internet connectivity, financial constraints, and workforce constraints, such as staff shortages and heavy workloads. Key facilitators included user-friendly technology design, strong leadership, effective teamwork, and ongoing communication. Evidence was predominantly qualitative, with only limited quantitative data available. Conclusions: DHTs show promise for improving access and continuity of care for cardiovascular disease, hypertension, and diabetes in rural settings; however, their impact is constrained by structural inequities, including limited broadband access, workforce shortages, and financial fragility. These findings highlight important implications for research, policy, and practice, including the need for rigorous mixed methods evaluations sensitive to rural contexts, long-term equity-oriented financing mechanisms, and strengthened organizational readiness to support effective DHT uptake.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/50404779937dcb6f1112ba04f05552ca" />
		
		<published>2026-07-31T15:45:10-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e91620 </id>
		<title>Advancing Human-Centered AI in Clinical Decision Support: Sociocognitive Human-in-the-Loop Study in HIV Care</title>
		<updated>2026-07-31T15:45:10-04:00</updated>

					<author>
				<name>Dezhi Wu</name>
			</author>
					<author>
				<name>Valerie Vera</name>
			</author>
					<author>
				<name>Sai Krishna Revanth Vuruma</name>
			</author>
					<author>
				<name>Lucas Aust</name>
			</author>
					<author>
				<name>Bharat Sowrya Yaddanapalli</name>
			</author>
					<author>
				<name>Jiaxuan Zhang</name>
			</author>
					<author>
				<name>Rithika Markanti</name>
			</author>
					<author>
				<name>Jiajia Zhang</name>
			</author>
					<author>
				<name>Xiaoming Li</name>
			</author>
					<author>
				<name>Sharon Weissman</name>
			</author>
					<author>
				<name>Bankole Olatosi</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e91620" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e91620">Background: AI-powered clinical decision support systems (CDSS) have shown promise in improving prediction, monitoring, and treatment optimization across clinical domains, including HIV care. However, translating AI outputs derived from electronic health records into clinically meaningful, trustworthy, and actionable decision support remains challenging, underscoring the need for more human-centered and socioecologically grounded CDSS design. Objective: This study aimed to explore how we can effectively translate the outputs of machine learning models based on HIV electronic health records into a real AI-powered CDSS for HIV care. Using the human-in-the-loop method, we engaged a set of stakeholders, including HIV physicians, nurse practitioners, infectious disease pharmacists, social workers, and case managers. Stakeholders interacted with an AI-powered CDSS prototype to identify barriers and challenges to adoption, as well as to inform a more holistic and context-aware AI-powered CDSS design. Methods: We conducted a field study at Prisma Health in South Carolina that included pre- and postsurveys, interactive usability testing sessions, think-alouds, and in-depth interviews with 16 clinicians providing HIV care between March and September 2025. We analyzed survey responses using descriptive statistics, and then transcribed and analyzed think-aloud and interview data using an etic and emic approach. Results: Clinicians identified multiple challenges and design considerations for AI-powered HIV CDSS, demonstrating that clinician-AI interaction is inherently sociotechnical and embedded across multiple socioecological levels. While clinicians relied on familiar clinical indicators as cognitive anchors for interpreting AI predictions, they emphasized that social determinants of health were central to their own risk assessment and clinical decision-making. Additionally, clinicians&#039; trust in AI is conditional and develops over time, with explainability and actionability emerging as critical factors for translating predictions into meaningful clinical interventions. Conclusions: Findings highlight the need to move beyond technically accurate predictions toward AI-powered CDSS designs that align with clinicians’ cognitive practices and socioecological realities of HIV care. By extending a sociocognitive framework through empirical grounding in HIV clinical practice, this study offers design insights for developing AI-powered CDSS that are trustworthy, context-aware, and capable of supporting actionable decision-making in HIV care settings and beyond.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/6b5983212380bc0df5679df2b4c9b9a2" />
		
		<published>2026-07-31T15:45:10-04:00</published>
	</entry>
	<entry>
		<id> https://www.jmir.org/2026/1/e77307 </id>
		<title>Impact of Large Language Model–Based AI Tools on Physician-Patient Communication: Systematic Review and Meta-Analysis</title>
		<updated>2026-07-31T15:30:12-04:00</updated>

					<author>
				<name>Sven Richter</name>
			</author>
					<author>
				<name>Clara Helene Buszello</name>
			</author>
					<author>
				<name>Markus Prem</name>
			</author>
					<author>
				<name>Sophia Willkommen</name>
			</author>
					<author>
				<name>Elida Hasani</name>
			</author>
					<author>
				<name>Ortrud Uckermann</name>
			</author>
					<author>
				<name>Tareq A Juratli</name>
			</author>
					<author>
				<name>Ilker Y Eyüpoglu</name>
			</author>
					<author>
				<name>Witold H Polanski</name>
			</author>
				<link rel="alternate" href="https://www.jmir.org/2026/1/e77307" />
					<summary type="html" xml:base="https://www.jmir.org/2026/1/e77307">Background: Recent advances in large language models (LLMs) such as GPT-3/4 have spurred the development of artificial intelligence (AI) chatbots and advisory tools in medicine. These systems are posited to assist or augment physician-patient communication, potentially improving empathy, clarity, and responsiveness. However, their actual impact on communication outcomes remains uncertain. Objective: This study aimed to systematically review and meta-analyze peer-reviewed studies (2020‐2025) evaluating how LLM-based interventions affect physician-patient communication, including empathy, clarity, trust, and patient understanding. Methods: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines, we searched PubMed/MEDLINE, Embase, Scopus, and Web of Science for studies published from 2020 to 2025 examining LLM or chatbot applications in clinical communication contexts. Eligible designs included randomized, observational, cross-sectional, and qualitative studies. Two reviewers (WHP and SR) independently screened titles or abstracts, assessed full texts, and extracted data on study design, population, LLM type, communication measures, and outcomes. We conducted a qualitative synthesis and random-effects meta-analysis, reporting pooled standardized mean differences or odds ratios with 95% CIs. Results: From 312 records, 10 studies were included, all quantitative and predominantly cross-sectional. Populations ranged from patients with chronic conditions to health care professionals and laypersons. Outcomes assessed included empathy (8 studies), clarity or information quality (6 studies), satisfaction or usefulness (4 studies), and trust perceptions (2 studies). In 6 direct comparisons of AI- versus physician-generated responses, LLMs were rated significantly higher in empathy in 5 studies. One large study found that chatbot replies were judged empathetic in 45.1% of cases versus 4.6% for physician replies (odds ratio approximately 9.8, &lt;.001). Similarly, ChatGPT-4 answers scored higher in empathy on a 5-point scale than human-written responses (mean 4.18 vs 2.70, &lt;.001). One neurology study showed higher empathy scores (Consultation and Relational Empathy Scale +1.38, &lt;.01) for ChatGPT answers. Only 1 study found no significant empathy difference. LLM content was also longer and more information-rich, improving patient-perceived clarity and understanding. On the other hand, GPT-4 simplified pathology reports, increasing patient comprehension scores (7.98 vs 5.23/10, &lt;.001) and reducing consultation time by 70%. However, AI replies were sometimes less concise or less readable for low-literacy patients. In pooled analyses (=4 studies; total evaluations N=2604), LLM assistance showed a large positive effect on empathy (standardized mean difference 1.02, 95% CI 0.44‐1.60; random-effects model). Patient satisfaction results were mixed. No study directly assessed long-term trust. Conclusions: Current evidence suggests that LLM-based chatbots can enhance physician-patient communication by producing more empathetic, detailed, and understandable responses. These improvements may positively influence patient experience and engagement. However, LLMs may also generate overly lengthy or occasionally inaccurate advice, emphasizing the need for physician oversight. While meta-analytic findings are promising, robust randomized controlled trials, real-world and longitudinal studies are needed to confirm benefits, assess trust outcomes, and define optimal clinical integration strategies.</summary>
		
        
                	<content type="image/png" src="https://jmir-production.s3.us-east-2.amazonaws.com/thumbs/de8ca5cc3944d1b5a1c636e2b1f43525" />
		
		<published>2026-07-31T15:30:12-04:00</published>
	</entry>
</feed>