❌

Normal view

Evidence for Digital Health Tools Designed to Support the Triage of Musculoskeletal Conditions in Primary, Urgent, and Emergency Care Settings: Scoping Review

Background: The digital health research field is growing rapidly, and a summary of the available digital tools for triaging musculoskeletal conditions is needed. Effective and safe digital triage tools for musculoskeletal conditions could support patients in making informed care decisions, aid clinicians and patients in navigating care, and may contribute to reducing ED overcrowding and healthcare costs. Objective: To identify and describe digital health tools for use by adults to triage musculoskeletal conditions across primary, urgent, or emergency care settings. Methods: Our scoping review was conducted following the Johanna Briggs Institute recommendations for scoping reviews and Arksey & O’Malley’s framework. Systematic searches in MEDLINE (OVID), CINAHL (EBSCO), PsycINFO (EBSCO), Embase (OVID), Cochrane Library, Web of Science, OpenGrey, GoogleScholar, arXiv.org, medRxiv.org, and an extensive grey literature search were conducted with a librarian scientist from inception to Sept 18, 2025. Studies had to recruit adults (18+ years) with musculoskeletal conditions that identified a digital health tool designed to triage or diagnose in primary, urgent, or emergency care settings and report primary data to be included. Two reviewer pairs independently screened abstracts and full-text articles Relevant data were extracted in duplicate, and results were summarized descriptively. Results: The search yielded 5695 records, and we screened 189 full-text articles. Thirty-four studies (n=37,509 patients) met the inclusion criteria. The most common musculoskeletal conditions reported were rheumatoid/inflammatory arthritis (n=13, 38%). Nineteen (59%) studies reported on symptom checkers, 13 (44%) studies on triage/diagnosis tools, and 2 (6%) were studies of diagnostic predictor tools. There were 16 unique digital health tools. Two tools were built for triaging musculoskeletal conditions, and were not publicly available outside the UK National Health Service. Most tools were generic tools designed to screen for general health problems, including musculoskeletal conditions. The most common approach to evaluating performance (eg, accuracy) of the tools was to compare the concordance of the tool to a clinician diagnosis or triage recommendation. Sensitivity and specificity ranged from 39%-91% and 23%-80%, respectively. Reported accuracy of included tools ranged from 33% to 98%. Conclusions: Musculoskeletal conditions remain a blind spot for people designing, implementing, and evaluating digital health for triage: few tools were specifically designed for musculoskeletal conditions, and most existing tools performed poorly when applied to musculoskeletal populations. The evidence base supporting accuracy of digital health for triaging and diagnosing musculoskeletal conditions is weak, and tool performance was inconsistent and lacking transparency. We recommend health systems and clinicians use a multi-modal approach, integrating both digital health tools and clinical decision-making to safely triage and diagnose until a more robust tool for musculoskeletal conditions is available. Future tool developers need to use transparent, standardized processes that prioritize tool safety, clinical value, and trustworthiness when designing for clinicians and patients.

From Agents to Governance: Essential AI Skills for Clinicians in the Large Language Model Era

Large language models are rapidly transitioning from pilot schemes to routine clinical practice. This creates an urgent need for clinicians to develop the necessary skills to strike the right balance between seizing opportunities and taking accountability. We propose a 3-tier competency framework to support clinicians’ evolution from cautious users to responsible stewards of artificial intelligence (AI). Tier 1 (foundational skills) defines the minimum competencies for safe use, including prompt engineering, human–AI agent interaction, security and privacy awareness, and the clinician-patient interface (transparency and consent). Tier 2 (intermediate skills) emphasizes evaluative expertise, including bias detection and mitigation, interpretation of explainability outputs, and the effective clinical integration of AI-generated workflows. Tier 3 (advanced skills) establishes leadership capabilities, mandating competencies in ethical governance (delineating accountability and liability boundaries), regulatory strategy, and model life cycle management—specifically, the ability to govern algorithmic adaptation and change protocols. Integrating this framework into continuing medical education programs and role-specific job descriptions could enhance clinicians’ ability to use AI safely and responsibly. This could standardize deployment and support safer clinical practice, with the potential to improve patient outcomes.

Africa’s Digital Health Revolution: The Digital Fit-Viability Model to Move From Innovation to Scaled Implementation

Digital innovations hold immense potential to transform health care delivery, particularly in sub-Saharan Africa, where financial, geographical, and infrastructural constraints continue to hinder progress toward universal health care delivery. Although a growing health tech sector offers creative solutions, few digital health interventions reach scaled implementation. In this paper, we present the digital fit/viability model—an adapted determinant framework to describe facilitators and barriers to moving from digital tools to integrated digital health implementation. We then use this model to describe the specific challenges and recommended solutions when developing digital health tools for health systems in sub-Saharan Africa.
  • ✇MIT Technology Review
  • The Download: next-gen nuclear, and the data center backlash Charlotte Jee
    This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. How next-generation nuclear reactors break out of the 20th-century blueprint   The popularity of commercial nuclear reactors has surged in recent years as worries about climate change and energy independence drowned out concerns about meltdowns and radioactive waste. The problem is, building nuclear power plants is expensive and slow.   A new generati
     

The Download: next-gen nuclear, and the data center backlash

14 January 2026 at 21:10

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

How next-generation nuclear reactors break out of the 20th-century blueprint  

The popularity of commercial nuclear reactors has surged in recent years as worries about climate change and energy independence drowned out concerns about meltdowns and radioactive waste. The problem is, building nuclear power plants is expensive and slow.  

A new generation of nuclear power technology could reinvent what a reactor looks like—and how it works. Advocates hope that new tech can refresh the industry and help replace fossil fuels without emitting greenhouse gases.  Here’s what that might look like.

—Casey Crownhart

Next-gen nuclear is one of our 10 Breakthrough Technologies this year. If you want to learn more about why it made the list, sign up to receive The Spark, our weekly newsletter all about energy and climate change, tomorrow. You can also check out the rest of the technologies on the list here.

Data centers are amazing. Everyone hates them.

The hyperscale datacenter is a marvel of our age. A masterstroke of engineering across multiple disciplines. They are nothing short of a technological wonder. People hate them.  

People hate them in Virginia, which leads the nation in their construction. They hate them in Nevada, where they slurp up the state’s precious water. They hate them in Michigan, and Arizona, and South Dakota. They hate them all around the world, it’s true. But they really hate them in Georgia. Read our story about why they’re provoking so much fury. 

—Mat Honan

This story first featured in The Debrief with Mat Honan, a weekly newsletter about the biggest stories in tech from our editor in chief. Sign up here to get the next one in your inbox on Friday.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 Iran is systematically crippling Starlink
The satellite internet service is meant to be impossible to jam—but the Iranian authorities are doing just that. (Rest of World)  
+ Messages getting around Iran’s internet block suggest that thousands of people have been killed. (NYT $)
+ On the ground in Ukraine’s largest Starlink repair shop. (MIT Technology Review)

2 Studies claiming microplastics harm us are being called into question
Some scientists say the discoveries are probably the result of contamination and false positives. (The Guardian) 

3 Trump is trying to temper the data center backlash 
He hopes cajoling tech companies to pay more and thus reduce people’s energy bills will do the trick. (WP $) 
+ Microsoft has just become the first tech company to promise it will do just that. (NYT $)
+ We know AI is power hungry. But just how big is the scale of the problem? (MIT Technology Review) 

4 US emissions jumped last year
Thanks to a combination of rising electricity demand, and more coal being burned to meet it. (NYT $)
+ But it’s not all bad news: coal power generation in India and China finally started to decline. (The Guardian)
+ Four bright spots in climate news in 2025. (MIT Technology Review)

5 Elon Musk needs to face consequences for his actions
If we tolerate him unleashing a flood of harassment of women and children, what will come next? (The Atlantic $) 
+ The US Senate has passed a bill that could give non-consensual deepfake victims a new way to fight back. (The Verge $)

6 Why the US is set to lose the race back to the moon 🚀🌔
Cuts to NASA aren’t helping, but they’re not the only problem. (Wired $)

7 Google’s Veo AI model can now turn portrait images into vertical videos
Really slick ones, too. (The Verge $)
+ AI-generated influencers are sharing fake images of them in bed with celebrities on Instagram. (404 Media $)

8 Former NYC mayor Eric Adams has been accused of a crypto ‘pump and dump’ 
He promoted a token that saw its market cap briefly soar to $580 million before plummeting. (Coindesk)

9 Are you a middle manager? Here’s some good news for you
Your skills are not being replaced by AI any time soon. (Quartz) 

10 Even miniscule lifestyle tweaks can extend your lifespan
A study of 60,000 adults found just a little bit more sleep and exercise makes a huge difference. (New Scientist $)
+ Aging hits us in our 40s and 60s. But well-being doesn’t have to fall off a cliff. (MIT Technology Review)

Quote of the day

“What I’m hopeful for in ’26 is for more people speaking up. Speaking truth to power is the point of freedom of speech, is the point of American society.”

—LinkedIn cofounder Reid Hoffman tells Wired he wants more people in Silicon Valley to start pushing back against the Trump administration this year. 

One more thing

two women collaborating on their laptops in a lecture hall
DEEP LEARNING INDABA 2024

What Africa needs to do to become a major AI player

Africa is still early in the process of adopting AI technologies. But researchers say the continent is uniquely hospitable to it for several reasons, including a relatively young and increasingly well-educated population, a rapidly growing ecosystem of AI startups, and lots of potential consumers.  

However, ambitious efforts to develop AI tools that answer the needs of Africans face numerous hurdles. Taken together, researchers worry, they could hold Africa’s AI sector back and hamper its efforts to pave its own pathway in the global AI race. Read the full story.

—Abdullahi Tsanni

We can still have nice things

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)

+ Still keen to do a bit of reflecting on the year behind and the one ahead? This free guide might help!
+ Turns out British comedian Rik Mayall had some pretty solid life advice.
+ I want to stay in this house in São Paolo.  
+ If you want to stop doomscrolling, it’s worth looking at your sleep habits. ($)

A nowhere-to-hide mechanism ensures complete piRNA-directed DNA methylation

Nature, Published online: 14 January 2026; doi:10.1038/s41586-025-09940-w

In mice, a SPOCD1–TPR-dependent ‘nowhere-to-hide’ mechanism is required for complete non-stochastic piRNA-directed LINE1 DNA methylation by preventing transposons from escaping surveillance within heterochromatin.

Enhancing telesurgical safety with predictive digital twin synchronization: a framework for latency compensation in robotic surgery

npj Digital Medicine, Published online: 13 January 2026; doi:10.1038/s41746-025-02283-w

Enhancing telesurgical safety with predictive digital twin synchronization: a framework for latency compensation in robotic surgery
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Internal Deployment Gaps in AI Regulation Joe Kwon · Stephen Casper
    arXiv:2601.08005v1 Announce Type: new Abstract: Frontier AI regulations primarily focus on systems deployed to external users, where deployment is more visible and subject to outside scrutiny. However, high-stakes applications can occur internally when companies deploy highly capable systems within their own organizations, such as for automating R\&D, accelerating critical business processes, and handling sensitive proprietary data. This paper examines how frontier AI regulations in the Uni
     

Internal Deployment Gaps in AI Regulation

arXiv:2601.08005v1 Announce Type: new Abstract: Frontier AI regulations primarily focus on systems deployed to external users, where deployment is more visible and subject to outside scrutiny. However, high-stakes applications can occur internally when companies deploy highly capable systems within their own organizations, such as for automating R\&D, accelerating critical business processes, and handling sensitive proprietary data. This paper examines how frontier AI regulations in the United States and European Union in 2025 handle internal deployment. We identify three gaps that could cause internally-deployed systems to evade intended oversight: (1) scope ambiguity that allows internal systems to evade regulatory obligations, (2) point-in-time compliance assessments that fail to capture the continuous evolution of internal systems, and (3) information asymmetries that subvert regulatory awareness and oversight. We then analyze why these gaps persist, examining tensions around measurability, incentives, and information access. Finally, we map potential approaches to address them and their associated tradeoffs. By understanding these patterns, we hope that policy choices around internally deployed AI systems can be made deliberately rather than incidentally.

Semantic Laundering in AI Agent Architectures: Why Tool Boundaries Do Not Confer Epistemic Warrant

arXiv:2601.08333v1 Announce Type: new Abstract: LLM-based agent architectures systematically conflate information transport mechanisms with epistemic justification mechanisms. We formalize this class of architectural failures as semantic laundering: a pattern where propositions with absent or weak warrant are accepted by the system as admissible by crossing architecturally trusted interfaces. We show that semantic laundering constitutes an architectural realization of the Gettier problem: propositions acquire high epistemic status without a connection between their justification and what makes them true. Unlike classical Gettier cases, this effect is not accidental; it is architecturally determined and systematically reproducible. The central result is the Theorem of Inevitable Self-Licensing: under standard architectural assumptions, circular epistemic justification cannot be eliminated. We introduce the Warrant Erosion Principle as the fundamental explanation for this effect and show that scaling, model improvement, and LLM-as-judge schemes are structurally incapable of eliminating a problem that exists at the type level.

WaterCopilot: An AI-Driven Virtual Assistant for Water Management

arXiv:2601.08559v1 Announce Type: new Abstract: Sustainable water resource management in transboundary river basins is challenged by fragmented data, limited real-time access, and the complexity of integrating diverse information sources. This paper presents WaterCopilot-an AI-driven virtual assistant developed through collaboration between the International Water Management Institute (IWMI) and Microsoft Research for the Limpopo River Basin (LRB) to bridge these gaps through a unified, interactive platform. Built on Retrieval-Augmented Generation (RAG) and tool-calling architectures, WaterCopilot integrates static policy documents and real-time hydrological data via two custom plugins: the iwmi-doc-plugin, which enables semantic search over indexed documents using Azure AI Search, and the iwmi-api-plugin, which queries live databases to deliver dynamic insights such as environmental-flow alerts, rainfall trends, reservoir levels, water accounting, and irrigation data. The system features guided multilingual interactions (English, Portuguese, French), transparent source referencing, automated calculations, and visualization capabilities. Evaluated using the RAGAS framework, WaterCopilot achieves an overall score of 0.8043, with high answer relevancy (0.8571) and context precision (0.8009). Key innovations include automated threshold-based alerts, integration with the LRB Digital Twin, and a scalable deployment pipeline hosted on AWS. While limitations in processing non-English technical documents and API latency remain, WaterCopilot establishes a replicable AI-augmented framework for enhancing water governance in data-scarce, transboundary contexts. The study demonstrates the potential of this AI assistant to support informed, timely decision-making and strengthen water security in complex river basins.

Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock

arXiv:2601.08673v1 Announce Type: new Abstract: Recent reports of large language models (LLMs) exhibiting behaviors such as deception, threats, or blackmail are often interpreted as evidence of alignment failure or emergent malign agency. We argue that this interpretation rests on a conceptual error. LLMs do not reason morally; they statistically internalize the record of human social interaction, including laws, contracts, negotiations, conflicts, and coercive arrangements. Behaviors commonly labeled as unethical or anomalous are therefore better understood as structural generalizations of interaction regimes that arise under extreme asymmetries of power, information, or constraint. Drawing on relational models theory, we show that practices such as blackmail are not categorical deviations from normal social behavior, but limiting cases within the same continuum that includes market pricing, authority relations, and ultimatum bargaining. The surprise elicited by such outputs reflects an anthropomorphic expectation that intelligence should reproduce only socially sanctioned behavior, rather than the full statistical landscape of behaviors humans themselves enact. Because human morality is plural, context-dependent, and historically contingent, the notion of a universally moral artificial intelligence is ill-defined. We therefore reframe concerns about artificial general intelligence (AGI). The primary risk is not adversarial intent, but AGI's role as an endogenous amplifier of human intelligence, power, and contradiction. By eliminating longstanding cognitive and institutional frictions, AGI compresses timescales and removes the historical margin of error that has allowed inconsistent values and governance regimes to persist without collapse. Alignment failure is thus structural, not accidental, and requires governance approaches that address amplification, complexity, and regime stability rather than model-level intent alone.

All Required, In Order: Phase-Level Evaluation for AI-Human Dialogue in Healthcare and Beyond

arXiv:2601.08690v1 Announce Type: new Abstract: Conversational AI is starting to support real clinical work, but most evaluation methods miss how compliance depends on the full course of a conversation. We introduce Obligatory-Information Phase Structured Compliance Evaluation (OIP-SCE), an evaluation method that checks whether every required clinical obligation is met, in the right order, with clear evidence for clinicians to review. This makes complex rules practical and auditable, helping close the gap between technical progress and what healthcare actually needs. We demonstrate the method in two case studies (respiratory history, benefits verification) and show how phase-level evidence turns policy into shared, actionable steps. By giving clinicians control over what to check and engineers a clear specification to implement, OIP-SCE provides a single, auditable evaluation surface that aligns AI capability with clinical workflow and supports routine, safe use.

Imaging-anchored Multiomics in Cardiovascular Disease: Integrating Cardiac Imaging, Bulk, Single-cell, and Spatial Transcriptomics

arXiv:2601.07871v1 Announce Type: cross Abstract: Cardiovascular disease arises from interactions between inherited risk, molecular programmes, and tissue-scale remodelling that are observed clinically through imaging. Health systems now routinely generate large volumes of cardiac MRI, CT and echocardiography together with bulk, single-cell and spatial transcriptomics, yet these data are still analysed in separate pipelines. This review examines joint representations that link cardiac imaging phenotypes to transcriptomic and spatially resolved molecular states. An imaging-anchored perspective is adopted in which echocardiography, cardiac MRI and CT define a spatial phenotype of the heart, and bulk, single-cell and spatial transcriptomics provide cell-type- and location-specific molecular context. The biological and technical characteristics of these modalities are first summarised, and representation-learning strategies for each are outlined. Multimodal fusion approaches are reviewed, with emphasis on handling missing data, limited sample size, and batch effects. Finally, integrative pipelines for radiogenomics, spatial molecular alignment, and image-based prediction of gene expression are discussed, together with common failure modes, practical considerations, and open challenges. Spatial multiomics of human myocardium and atherosclerotic plaque, single-cell and spatial foundation models, and multimodal medical foundation models are collectively bringing imaging-anchored multiomics closer to large-scale cardiovascular translation.

GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards

arXiv:2601.08183v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve MLLMs were evaluated across a five-stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro-F1, mean Intersection-over-Union (mIoU), and multi-dimensional Likert scale. Gemini-3-Pro achieved state-of-the-art performance. In diagnostic reasoning, top-tier models (Macro-F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p>0.05). However, a critical "spatial grounding bottleneck" persisted; human lesion localization (mIoU >0.506) significantly outperformed the best model (0.345; p

Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs

arXiv:2601.08634v1 Announce Type: cross Abstract: While recent research has systematically documented political orientation in large language models (LLMs), existing evaluations rely primarily on direct probing or demographic persona engineering to surface ideological biases. In social psychology, however, political ideology is also understood as a downstream consequence of fundamental moral intuitions. In this work, we investigate the causal relationship between moral values and political positioning by treating moral orientation as a controllable condition. Rather than simply assigning a demographic persona, we condition models to endorse or reject specific moral values and evaluate the resulting shifts on their political orientations, using the Political Compass Test. By treating moral values as lenses, we observe how moral conditioning actively steers model trajectories across economic and social dimensions. Our findings show that such conditioning induces pronounced, value-specific shifts in models' political coordinates. We further notice that these effects are systematically modulated by role framing and model scale, and are robust across alternative assessment instruments instantiating the same moral value. This highlights that effective alignment requires anchoring political assessments within the context of broader social values including morality, paving the way for more socially grounded alignment techniques.

ISLA: A U-Net for MRI-based acute ischemic stroke lesion segmentation with deep supervision, attention, domain adaptation, and ensemble learning

arXiv:2601.08732v1 Announce Type: cross Abstract: Accurate delineation of acute ischemic stroke lesions in MRI is a key component of stroke diagnosis and management. In recent years, deep learning models have been successfully applied to the automatic segmentation of such lesions. While most proposed architectures are based on the U-Net framework, they primarily differ in their choice of loss functions and in the use of deep supervision, residual connections, and attention mechanisms. Moreover, many implementations are not publicly available, and the optimal configuration for acute ischemic stroke (AIS) lesion segmentation remains unclear. In this work, we introduce ISLA (Ischemic Stroke Lesion Analyzer), a new deep learning model for AIS lesion segmentation from diffusion MRI, trained on three multicenter databases totaling more than 1500 AIS participants. Through systematic optimization of the loss function, convolutional architecture, deep supervision, and attention mechanisms, we developed a robust segmentation framework. We further investigated unsupervised domain adaptation to improve generalization to an external clinical dataset. ISLA outperformed two state-of-the-art approaches for AIS lesion segmentation on an external test set. Codes and trained models will be made publicly available to facilitate reuse and reproducibility.

DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction

arXiv:2509.14507v2 Announce Type: replace Abstract: Natural Language to SQL (NL2SQL) provides a new model-centric paradigm that simplifies database access for non-technical users by converting natural language queries into SQL commands. Recent advancements, particularly those integrating Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT) reasoning, have made significant strides in enhancing NL2SQL performance. However, challenges such as inaccurate task decomposition and keyword extraction by LLMs remain major bottlenecks, often leading to errors in SQL generation. While existing datasets aim to mitigate these issues by fine-tuning models, they struggle with over-fragmentation of tasks and lack of domain-specific keyword annotations, limiting their effectiveness. To address these limitations, we present DeKeyNLU, a novel dataset which contains 1,500 meticulously annotated QA pairs aimed at refining task decomposition and enhancing keyword extraction precision for the RAG pipeline. Fine-tuned with DeKeyNLU, we propose DeKeySQL, a RAG-based NL2SQL pipeline that employs three distinct modules for user question understanding, entity retrieval, and generation to improve SQL generation accuracy. We benchmarked multiple model configurations within DeKeySQL RAG pipeline. Experimental results demonstrate that fine-tuning with DeKeyNLU significantly improves SQL generation accuracy on both BIRD (62.31% to 69.10%) and Spider (84.2% to 88.7%) dev datasets.

Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems

arXiv:2512.20387v4 Announce Type: replace Abstract: We propose a Vision-Language Simulation Model (VLSM) that unifies visual and textual understanding to synthesize executable FlexScript from layout sketches and natural-language prompts, enabling cross-modal reasoning for industrial simulation systems. To support this new paradigm, the study constructs the first large-scale dataset for generative digital twins, comprising over 120,000 prompt-sketch-code triplets that enable multimodal learning between textual descriptions, spatial structures, and simulation logic. In parallel, three novel evaluation metrics, Structural Validity Rate (SVR), Parameter Match Rate (PMR), and Execution Success Rate (ESR), are proposed specifically for this task to comprehensively evaluate structural integrity, parameter fidelity, and simulator executability. Through systematic ablation across vision encoders, connectors, and code-pretrained language backbones, the proposed models achieve near-perfect structural accuracy and high execution robustness. This work establishes a foundation for generative digital twins that integrate visual reasoning and language understanding into executable industrial simulation systems. Project page: https://danielhsu2014.github.io/GDT-VLSM-project/
❌