❌

Reading view

The Download: a peek at AI’s future

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

The State of AI: A vision of the world in 2030  

There are huge gulfs of opinion when it comes to predicting the near-future impacts of generative AI. In one camp there are those who predict that over the next decade the impact of AI will exceed that of the Industrial Revolution—a 150-year period of economic and social upheaval so great that we still live in the world it wrought. 

At the other end of the scale we have team ‘Normal Technology’: experts who push back not only on these sorts of predictions but on their foundational worldview. That’s not how technology works, they argue.

Advances at the cutting edge may come thick and fast, but change across the wider economy, and society as a whole, moves at human speed. Widespread adoption of new technologies can be slow; acceptance slower. AI will be no different. What should we make of these extremes? 

Read the full conversation between MIT Technology Review’s senior AI editor Will Douglas Heaven and Tim Bradshaw, FT global tech correspondent, about where AI will go next, and what our world will look like in the next five years.

This is the final edition of The State of AI, a collaboration between the Financial Times and MIT Technology Review. Read the rest of the series, and if you want to keep up-to-date with what’s going on in the world of AI, sign up to receive our free Algorithm newsletter every Monday.

How AI is changing the economy

There’s a lot at stake when it comes to understanding how AI is changing the economy at large. What’s the right outlook to have? Join Mat Honan, editor in chief, David Rotman, editor at large, and Richard Waters, FT columnist, at 1pm ET today to hear them discuss what’s happening across industries and the market. Sign up now to be part of this exclusive subscriber-only event.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 Trump says he’ll sign an order blocking states from regulating AI
But he’s facing a lot of pushback, including from members of his own party. (CNN)
+ The whole debacle can be traced back to congressional inaction. (Semafor)

2 Google’s new smart glasses are getting rave reviews 👓
You’ll be able to get your hands on a pair in 2026. Watch out, Apple and Meta. (Tech Radar)

3 Trump gave the go-ahead for Nvidia to sell powerful AI chips to China
The US gets a 25% cut of the sales—but what does it lose longer-term? (WP $)
+ And how much could China stand to gain? (NYT $)
+ How a top Chinese AI model overcame US sanctions. (MIT Technology Review)

4 America’s data center backlash is here
Republican and Democrat alike, local residents are sick of rapidly rising power bills. (Vox $)
+ More than 200 environmental groups are demanding a US-wide moratorium on new data centers. (The Guardian)
+ The data center boom in the desert. (MIT Technology Review)

5 A quarter of teens are turning to AI chatbots for mental health support
Given the lack of real-world help, can you really blame them? (The Guardian)
+ Therapists are secretly using ChatGPT. Clients are triggered. (MIT Technology Review)

6 ICEBlock is suing the US government over its App Store removal 
Its creator is arguing that the Department of Justice’s demands to Apple violated his First Amendment rights. (404 Media)
+ It’s one of a number of ICE-tracking initiatives to be pulled by tech platforms this year. (MIT Technology Review)

7 This band quit Spotify, but it’s been replaced by AI knockoffs
The platform seems to be struggling against the tide of slop. (Futurism) 
+ AI is coming for music, too. (MIT Technology Review)

8 Think you’re immune to online ads? Think again
If you’re scrolling on social media, you’re being sold to. Relentlessly. (The Verge $)

9 People really do not like Microsoft Copilot
It’s like Clippy all over again, except it’s even less avoidable. (Quartz $)

10 The longest solar eclipse for 100 years is coming
And we’ll only have to wait until 2027 to see it! (Wired $)

Quote of the day

“Governments and MPs are shooting themselves in the foot by pandering to tech giants, because that just tells young people that they don’t care about our future.”

—Adele Zeynep Walton, founding member of online safety campaign group Ctrl+Alt+Reclaim, tells The Guardian why young activists are taking matters into their own hands. 

One more thing

fleet of ships at sea
COURTESY OF OCEANBIRD

Inside the long quest to advance Chinese writing technology

Every second of every day, someone is typing in Chinese. Though the mechanics look a little different from typing in English—people usually type the pronunciation of a character and then pick it out of a selection that pops up, autocomplete-style—it’s hard to think of anything more quotidian. The software that allows this exists beneath the awareness of pretty much everyone who uses it. It’s just there.

What’s largely been forgotten is that a large cast of eccentrics and linguists, engineers and polymaths, spent much of the 20th century torturing themselves over how Chinese was ever going to move away from the ink brush to any other medium. Read the full story.

—Veronique Greenwood

We can still have nice things

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)

+ Pantone chose a ‘calming’ shade of white for its Color of 2026… and people are fuming. 
+ Ozempic needles on the Christmas tree, anyone? Here’s why we’re going crazy for weird baubles. 
+ Can relate to this baby seal for instinctively heading to the nearest pub.
+ Thrilled to see One Battle After Another get so many Golden Globes nominations.

  •  

AI Application in Anti-Money Laundering for Sustainable and Transparent Financial Systems

arXiv:2512.06240v1 Announce Type: new Abstract: Money laundering and financial fraud remain major threats to global financial stability, costing trillions annually and challenging regulatory oversight. This paper reviews how artificial intelligence (AI) applications can modernize Anti-Money Laundering (AML) workflows by improving detection accuracy, lowering false-positive rates, and reducing the operational burden of manual investigations, thereby supporting more sustainable development. It further highlights future research directions including federated learning for privacy-preserving collaboration, fairness-aware and interpretable AI, reinforcement learning for adaptive defenses, and human-in-the-loop visualization systems to ensure that next-generation AML architectures remain transparent, accountable, and robust. In the final part, the paper proposes an AI-driven KYC application that integrates graph-based retrieval-augmented generation (RAG Graph) with generative models to enhance efficiency, transparency, and decision support in KYC processes related to money-laundering detection. Experimental results show that the RAG-Graph architecture delivers high faithfulness and strong answer relevancy across diverse evaluation settings, thereby enhancing the efficiency and transparency of KYC CDD/EDD workflows and contributing to more sustainable, resource-optimized compliance practices.
  •  

KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening

arXiv:2512.05994v1 Announce Type: cross Abstract: With the rapid advancement of conversational and diffusion-based AI, there is a growing adoption of AI in educational services, ranging from grading and assessment tools to personalized learning systems that provide targeted support for students. However, this adaptability has yet to fully extend to the domain of children's speech, where existing models often fail due to their reliance on datasets designed for clear, articulate adult speech. Children, particularly those in early developmental stages or with speech and language pathologies, present unique challenges that current AI models and datasets are ill-equipped to handle. To address this, we introduce KidSpeak, a multi-task speech-enhanced Foundation Model capable of both generative and discriminative tasks specifically tailored to children's speech patterns. Our framework employs a two-stage training process that incorporates phonetic knowledge into the speech encoder, achieving an average accuracy of 87% across four separate tasks. Furthermore, recognizing the limitations of scalable human annotation and existing speech alignment tools, we propose the Flexible and Automatic Speech Aligner (FASA) and leverage the method to construct high quality datasets for training and evaluation. This novel alignment tool significantly improves the quality of aligned children's speech from noisy data, enhancing data quality by 13.6x compared to human annotations, as demonstrated on the CHILDES dataset. To the best of our knowledge, KidSpeak and FASA represent the first comprehensive solution designed for speech and language therapy in children, offering both a multi-purpose speech LLM and a robust alignment tool.
  •  

Data Visualization Support for Interdisciplinary Team Treatment Planning in Clinical Oncology: Scoping Review

Background: Complex and expanding datasets in clinical oncology applications require flexible and interactive visualization of patient data to provide physicians and other medical professionals with maximum amount of information. In particular, interdisciplinary tumor conferences profit from customized tools to integrate, link, and visualize relevant data from all professions involved. Objective: Our objective was to identify and present currently available data visualization tools for tumor boards and related areas. We wanted to provide an overview of not only the digital tools currently used in tumor board settings but also of the data they include, their respective visualization solutions, and their integration into hospital processes. Methods: This scoping review was based on the scoping study framework by Arksey and O’Malley and attempted to answer the following research question: “What are the key features of data visualization solutions used in molecular and organ tumor boards, and how are these elements integrated and used within the clinical setting?” The following electronic databases were searched for articles: PubMed, Web of Science, and Scopus. Articles were deemed eligible if published in English in the last 10 years. Eligible articles were first deduplicated, followed by screening of titles and abstracts. Full-text screening was then conducted to decide on article selection. All included articles were analyzed using a data extraction template. The template included a variety of meta-information, as well as specific fields aiming to answer the research question. Results: The review process started with 2049 articles, of which 1014 (49.49%) were included in the title and abstract screening. A total of 5.47% (112/2049) of the publications were eligible for full-text screening, leading to 2.93% (60/2049) of the publications being eligible for final inclusion. They covered 49 distinct visualization tools and applications. We discovered a variety of innovative visualization solutions, most often driven by the complexity of omics data, represented in 96% (47/49) of the tools. Tables remained the most used tool for the visualization of data types described in the articles. Approximately one-third of the identified tools (16/49, 33%) were systematically evaluated in some form. For most discovered tools (37/49, 76%), there was no documentation of implementation into the clinical routine. A significant number of applications (21/49, 43%) were available through open-source access. Conclusions: There is a wide range of projects providing visualization solutions for tumor boards and clinical oncology applications. Among the few tools that have made their way into clinical routine settings, there are both commercial and academic solutions. While tables for a variety of data types remain the dominant visualization strategy, the complexity of omics data appears to be the driving force behind many visualization innovations in the domain of tumor boards. Trial Registration:
  •  

Digital Health Technologies: Learnings and Perspectives From a Patient Engagement Stakeholder Expectations Matrix Study

As digital health technologies become increasingly integrated into health care systems worldwide, there is growing recognition that their full potential can be realized only when development is rooted in patient engagement (PE). Despite its proven value in clinical research and health care delivery, PE remains insufficiently embedded in digital health design and implementation. This perspective paper explores the current state of PE in digital health through findings from the Stakeholder Expectations Matrix program developed by Patient Focused Medicines Development. Drawing from 37 in-depth interviews across 6 key stakeholder groups, complemented by insights gathered during a multisession cocreation track at the Patient Engagement Open Forum, this paper highlights differing perspectives on digital health, the barriers to meaningful engagement, and the fragmented nature of data governance and technology adoption. Findings point not only to significant gaps in shared understanding, infrastructure, and policy but also to clear opportunities for collaboration, including early recommendations for building a more inclusive and patient-centered digital health ecosystem, one that supports sustainable innovation, trust, and systemwide impact.
  •  

WisPaper: Your AI Scholar Search Engine

arXiv:2512.06879v1 Announce Type: cross Abstract: Researchers struggle to efficiently locate and manage relevant literature within the exponentially growing body of scientific publications. We present \textsc{WisPaper}, an intelligent academic retrieval and literature management platform that addresses this challenge through three integrated capabilities: (1) \textit{Scholar Search}, featuring both quick keyword-based and deep agentic search modes for efficient paper discovery; (2) \textit{Library}, a customizable knowledge base for systematic literature organization; and (3) \textit{AI Feeds}, an intelligent recommendation system that automatically delivers relevant new publications based on user interests. Unlike existing academic tools, \textsc{WisPaper} provides a closed-loop workflow that seamlessly connects literature discovery, management, and continuous tracking of research frontiers. Our multilingual and multidisciplinary system significantly reduces the time researchers from diverse backgrounds spend on paper screening and management, enabling them to focus on their core research activities. The platform is publicly accessible and serves researchers across academia and industry.
  •  

A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning

arXiv:2512.07136v1 Announce Type: cross Abstract: Multimodal human action recognition (HAR) leverages complementary sensors for activity classification. Beyond recognition, recent advances in large language models (LLMs) enable detailed descriptions and causal reasoning, motivating new tasks: human action understanding (HAU) and human action reasoning (HARn). However, most LLMs, especially large vision language models (LVLMs), struggle with non-RGB modalities such as depth, IMU, and mmWave due to the lack of large-scale data-caption resources. Existing HAR datasets mainly provide coarse data-label annotations, which are insufficient to capture fine-grained action dynamics needed for HAU and HARn. We consider two ground-truth pair types: (1) data label (discrete category) and (2) data caption (textual description). Naively generating captions from labels often lacks logical and spatiotemporal consistency. We introduce CUHK-X, a large-scale multimodal dataset and benchmark suite for HAR, HAU, and HARn. CUHK-X contains 58,445 samples covering 40 actions performed by 30 participants across two indoor environments. To improve caption consistency, we propose a prompt-based scene creation method that leverages LLMs to generate logically connected activity sequences, followed by human validation. CUHK-X includes three benchmarks with six evaluation tasks. Experiments report average accuracies of 76.52% (HAR), 40.76% (HAU), and 70.25% (HARn). CUHK-X aims to enable the community to apply and develop data-intensive learning methods for robust, multimodal human activity analysis. Project page and code: https://openaiotlab.github.io/CUHK-X/ and https://github.com/openaiotlab/CUHK-X.
  •  

A Field Guide to Deploying AI Agents in Clinical Practice

arXiv:2509.26153v3 Announce Type: replace Abstract: Large language models (LLMs) integrated into agent-driven workflows hold immense promise for healthcare, yet a significant gap exists between their potential and practical implementation within clinical settings. To address this, we present a practitioner-oriented field manual for deploying generative agents that use electronic health record (EHR) data. This guide is informed by our experience deploying the "irAE-Agent", an automated system to detect immune-related adverse events from clinical notes at Mass General Brigham, and by structured interviews with 21 clinicians, engineers, and informatics leaders involved in the project. Our analysis reveals a critical misalignment in clinical AI development: less than 20% of our effort was dedicated to prompt engineering and model development, while over 80% was consumed by the sociotechnical work of implementation. We distill this effort into five "heavy lifts": data integration, model validation, ensuring economic value, managing system drift, and governance. By providing actionable solutions for each of these challenges, this field manual shifts the focus from algorithmic development to the essential infrastructure and implementation work required to bridge the "valley of death" and successfully translate generative AI from pilot projects into routine clinical care.
  •  

BEDI: A Comprehensive Benchmark for Evaluating Embodied Agents on UAVs

arXiv:2505.18229v2 Announce Type: replace-cross Abstract: With the rapid advancement of low-altitude remote sensing and Vision-Language Models (VLMs), Embodied Agents based on Unmanned Aerial Vehicles (UAVs) have shown significant potential in autonomous tasks. However, current evaluation methods for UAV-Embodied Agents (UAV-EAs) remain constrained by the lack of standardized benchmarks, diverse testing scenarios and open system interfaces. To address these challenges, we propose BEDI (Benchmark for Embodied Drone Intelligence), a systematic and standardized benchmark designed for evaluating UAV-EAs. Specifically, we introduce a novel Dynamic Chain-of-Embodied-Task paradigm based on the perception-decision-action loop, which decomposes complex UAV tasks into standardized, measurable subtasks. Building on this paradigm, we design a unified evaluation framework encompassing six core sub-skills: semantic perception, spatial perception, motion control, tool utilization, task planning and action generation. Furthermore, we develop a hybrid testing platform that incorporates a wide range of both virtual and real-world scenarios, enabling a comprehensive evaluation of UAV-EAs across diverse contexts. The platform also offers open and standardized interfaces, allowing researchers to customize tasks and extend scenarios, thereby enhancing flexibility and scalability in the evaluation process. Finally, through empirical evaluations of several state-of-the-art (SOTA) VLMs, we reveal their limitations in embodied UAV tasks, underscoring the critical role of the BEDI benchmark in advancing embodied intelligence research and model optimization. By filling the gap in systematic and standardized evaluation within this field, BEDI facilitates objective model comparison and lays a robust foundation for future development in this field. Our benchmark is now publicly available at https://github.com/lostwolves/BEDI.
  •  

Learning to Align Human Code Preferences

arXiv:2507.20109v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable potential in automating software development tasks. While recent advances leverage Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) to align models with human preferences, the optimal training strategy remains unclear across diverse code preference scenarios. This paper systematically investigates the roles of SFT and DPO in aligning LLMs with different code preferences. Through both theoretical analysis and empirical observation, we hypothesize that SFT excels in scenarios with objectively verifiable optimal solutions, while applying SFT followed by DPO (S&D) enables models to explore superior solutions in scenarios without objectively verifiable optimal solutions. Based on the analysis and experimental evidence, we propose Adaptive Preference Optimization (APO), a dynamic integration approach that adaptively amplifies preferred responses, suppresses dispreferred ones, and encourages exploration of potentially superior solutions during training. Extensive experiments across six representative code preference tasks validate our theoretical hypotheses and demonstrate that APO consistently matches or surpasses the performance of existing SFT and S&D strategies. Our work provides both theoretical foundations and practical guidance for selecting appropriate training strategies in different code preference alignment scenarios.
  •  

Algorithms Trained on Normal Chest X-rays Can Predict Health Insurance Types

arXiv:2511.11030v4 Announce Type: replace-cross Abstract: Artificial intelligence is revealing what medicine never intended to encode. Deep vision models, trained on chest X-rays, can now detect not only disease but also invisible traces of social inequality. In this study, we show that state-of-the-art architectures (DenseNet121, SwinV2-B, MedMamba) can predict a patient's health insurance type, a strong proxy for socioeconomic status, from normal chest X-rays with significant accuracy (AUC around 0.70 on MIMIC-CXR-JPG, 0.68 on CheXpert). The signal was unlikely contributed by demographic features by our machine learning study combining age, race, and sex labels to predict health insurance types; it also remains detectable when the model is trained exclusively on a single racial group. Patch-based occlusion reveals that the signal is diffuse rather than localized, embedded in the upper and mid-thoracic regions. This suggests that deep networks may be internalizing subtle traces of clinical environments, equipment differences, or care pathways; learning socioeconomic segregation itself. These findings challenge the assumption that medical images are neutral biological data. By uncovering how models perceive and exploit these hidden social signatures, this work reframes fairness in medical AI: the goal is no longer only to balance datasets or adjust thresholds, but to interrogate and disentangle the social fingerprints embedded in clinical data itself.
  •  

Simulating Life Paths with Digital Twins: AI-Generated Future Selves Influence Decision-Making and Expand Human Choice

arXiv:2512.05397v2 Announce Type: replace-cross Abstract: Major life transitions demand high-stakes decisions, yet people often struggle to imagine how their future selves will live with the consequences. To support this limited capacity for mental time travel, we introduce AI-enabled digital twins that have ``lived through'' simulated life scenarios. Rather than predicting optimal outcomes, these simulations extend prospective cognition by making alternative futures vivid enough to support deliberation without assuming which path is best. We evaluate this idea in a randomized controlled study (N=192) using multimodal synthesis - facial age progression, voice cloning, and large language model dialogue - to create personalized avatars representing participants 30 years forward. Young adults 18 to 28 years old described pending binary decisions and were assigned to guided imagination or one of four avatar conditions: single-option, balanced dual-option, or expanded three-option with a system-generated novel alternative. Results showed asymmetric effects: single-sided avatars increased shifts toward the presented option, while balanced presentation produced movement toward both. Introducing a system-generated third option increased adoption of this new alternative compared to control, suggesting that AI-generated future selves can expand choice by surfacing paths that might otherwise go unnoticed. Participants rated evaluative reasoning and eudaimonic meaning-making as more important than emotional or visual vividness. Perceived persuasiveness and baseline agency predicted decision change. These findings advance understanding of AI-mediated episodic prospection and raise questions about autonomy in AI-augmented decisions.
  •  

The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics

arXiv:2512.05765v1 Announce Type: new Abstract: Influential critiques argue that Large Language Models (LLMs) are a dead end for AGI: "mere pattern matchers" structurally incapable of reasoning or planning. We argue this conclusion misidentifies the bottleneck: it confuses the ocean with the net. Pattern repositories are the necessary System-1 substrate; the missing component is a System-2 coordination layer that selects, constrains, and binds these patterns. We formalize this layer via UCCT, a theory of semantic anchoring that models reasoning as a phase transition governed by effective support (rho_d), representational mismatch (d_r), and an adaptive anchoring budget (gamma log k). Under this lens, ungrounded generation is simply an unbaited retrieval of the substrate's maximum likelihood prior, while "reasoning" emerges when anchors shift the posterior toward goal-directed constraints. We translate UCCT into architecture with MACI, a coordination stack that implements baiting (behavior-modulated debate), filtering (Socratic judging), and persistence (transactional memory). By reframing common objections as testable coordination failures, we argue that the path to AGI runs through LLMs, not around them.
  •  

XR-DT: Extended Reality-Enhanced Digital Twin for Agentic Mobile Robots

arXiv:2512.05270v1 Announce Type: cross Abstract: As mobile robots increasingly operate alongside humans in shared workspaces, ensuring safe, efficient, and interpretable Human-Robot Interaction (HRI) has become a pressing challenge. While substantial progress has been devoted to human behavior prediction, limited attention has been paid to how humans perceive, interpret, and trust robots' inferences, impeding deployment in safety-critical and socially embedded environments. This paper presents XR-DT, an eXtended Reality-enhanced Digital Twin framework for agentic mobile robots, that bridges physical and virtual spaces to enable bi-directional understanding between humans and robots. Our hierarchical XR-DT architecture integrates virtual-, augmented-, and mixed-reality layers, fusing real-time sensor data, simulated environments in the Unity game engine, and human feedback captured through wearable AR devices. Within this framework, we design an agentic mobile robot system with a unified diffusion policy for context-aware task adaptation. We further propose a chain-of-thought prompting mechanism that allows multimodal large language models to reason over human instructions and environmental context, while leveraging an AutoGen-based multi-agent coordination layer to enhance robustness and collaboration in dynamic tasks. Initial experimental results demonstrate accurate human and robot trajectory prediction, validating the XR-DT framework's effectiveness in HRI tasks. By embedding human intention, environmental dynamics, and robot cognition into the XR-DT framework, our system enables interpretable, trustworthy, and adaptive HRI.
  •  

Simulating Life Paths with Digital Twins: AI-Generated Future Selves Influence Decision-Making and Expand Human Choice

arXiv:2512.05397v1 Announce Type: cross Abstract: Major life transitions demand high-stakes decisions, yet people often struggle to imagine how their future selves will live with the consequences. To support this limited capacity for mental time travel, we introduce AI-enabled digital twins that have ``lived through'' simulated life scenarios. Rather than predicting optimal outcomes, these simulations extend prospective cognition by making alternative futures vivid enough to support deliberation without assuming which path is best. We evaluate this idea in a randomized controlled study (N=192) using multimodal synthesis - facial age progression, voice cloning, and large language model dialogue - to create personalized avatars representing participants 30 years forward. Young adults 18 to 28 years old described pending binary decisions and were assigned to guided imagination or one of four avatar conditions: single-option, balanced dual-option, or expanded three-option with a system-generated novel alternative. Results showed asymmetric effects: single-sided avatars increased shifts toward the presented option, while balanced presentation produced movement toward both. Introducing a system-generated third option increased adoption of this new alternative compared to control, suggesting that AI-generated future selves can expand choice by surfacing paths that might otherwise go unnoticed. Participants rated evaluative reasoning and eudaimonic meaning-making as more important than emotional or visual vividness. Perceived persuasiveness and baseline agency predicted decision change. These findings advance understanding of AI-mediated episodic prospection and raise questions about autonomy in AI-augmented decisions.
  •  

ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset

arXiv:2511.15718v2 Announce Type: replace Abstract: Large Language Model (LLM) agents have developed rapidly in recent years to solve complex real-world problems using external tools. However, the scarcity of high-quality trajectories still hinders the development of stronger LLM agents. Most existing works on multi-turn dialogue synthesis validate correctness only at the trajectory level, which may overlook turn-level errors that can propagate during training and degrade model performance. To address these limitations, we introduce ToolMind, a large-scale, high-quality tool-agentic dataset with 160k synthetic data instances generated using over 20k tools and 200k augmented open-source data instances. Our data synthesis pipeline first constructs a function graph based on parameter correlations and then uses a multi-agent framework to simulate realistic user-assistant-tool interactions. Beyond trajectory-level validation, we employ fine-grained turn-level filtering to remove erroneous or suboptimal steps, ensuring that only high-quality reasoning traces are retained. This approach mitigates error amplification during training while preserving self-corrective reasoning signals essential for robust tool-use learning. Models fine-tuned on ToolMind show significant improvements over baselines on several benchmarks.
  •  

The AI Productivity Index (APEX)

arXiv:2509.25721v4 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.
  •  

Exploring a Digital Health Solution to Collect and Manage Health-Related Needs for Patients Who Undergo Complex Surgery: Mixed Methods Study

Background: Patients who undergo complex surgery (e.g., esophagectomy, liver resection) often experience substantial burden of health-related needs (medical, social, and behavioral health). A closed loop digital solution could facilitate the collection and resolution of health-related needs by care team members for patients who undergo complex surgery. A digital solution may facilitate adherence to a clear treatment plan and concomitantly reduce surgical complications and readmissions associated with unmet health-related needs, which remain persistent challenges across health care settings. Objective: To establish problems and gaps in the collection, integration, and management of health-related needs and identify a set of user specifications for a digital solution to collect and manage health-related needs, specifically medical, social, and behavioral needs for patients who undergo complex surgery. Methods: We applied the Double Diamond Framework and organized the study into two sequential phases: (1) qualitative methods to discover patients’ and care team members’ perspectives on health-related needs; (2) participatory design sessions to gain feedback and sentiment about ideal features of a digital solution. Both phases were conducted between December 2023 and March 2025. We supplemented both phases with analysis of electronic health record (EHR) data for patients who underwent complex surgery at our academic medical center (AMC). Results: Extensive themes emerged from interviews with patients (n=20) and care team members (n=24), capturing their health-related and surgical experiences as well as desired features for a proposed digital solution. Our swim lane diagram demonstrated four critical gaps in workflow: (1) heterogeneity in the approach to screening, monitoring, and managing health-related needs; (2) patients felt uncomfortable reporting health-related needs, particularly behavioral and social needs, to their care team; (3) lack of access to referral resources to resolve needs; and (4) the need for a closed loop intervention for patients and care team members. A subset of participants from Phase 1 (n=5 patients and n=9 care team members) provided feedback on preferred features, drawing from digital tools currently available in the EHR at our AMC. Among four existing EHR tools tested, there were notable variations in how patients and care team members felt about their potential use. Participants also provided extensive feedback for preferred components (e.g., goals and active plans) that should be available in an existing or custom digital solution to manage health-related needs. Findings from the qualitative interviews and design sessions were corroborated with EHR documentation. Conclusions: Digital solutions could provide a streamlined approach for collection and management of health-related needs in surgery, with the goal of addressing unmet needs and improving patient activation. This approach is critical to ensure patients, especially patients who undergo complex surgery, have positive health outcomes. We identified preferences for specific features in a proposed digital solution based on our systematic assessment that will inform future work.
  •  

EIF3M as a pan-cancer biomarker: prognostic significance and immune infiltration association

Front Mol Biosci. 2025 Nov 18;12:1697083. doi: 10.3389/fmolb.2025.1697083. eCollection 2025.

ABSTRACT

BACKGROUND: EIF3M, a core subunit of eukaryotic translation initiation factor 3, plays a pivotal role in protein synthesis by regulating the assembly of the 43S initiation complex. However, its biological functions in cancer remain poorly understood. To further investigate the clinical translational value and underlying mechanisms of EIF3M in tumors, this study conducted comprehensive bioinformatic analysis of EIF3M across various tumor types.

METHODS: We utilized publicly available databases to perform a comprehensive bioinformatics analysis of EIF3M's biological roles in oncogenesis, aiming to elucidate its pan-cancer expression patterns and prognostic significance. Furthermore, we conducted an integrative multi-omics analysis incorporating methylation profiling, co-expressed gene networks, targeted miRNA interactions, and tumor immune microenvironment infiltration to decipher the complex regulatory architecture and biological pathways mediated by EIF3M across cancer types. Finally, we used HCC cell lines for in vitro functional validation, determining how EIF3M expression modulates malignant phenotypic behaviors in hepatocellular carcinoma.

RESULTS: EIF3M was overexpressed in multiple cancers and correlated with advanced tumor stage and poor survival. Its dysregulation was primarily driven by gene amplification and regulated by promoter methylation and miRNAs. EIF3M functioned as a hub in cell cycle and transcriptional networks and was linked to an immunosuppressive microenvironment. In hepatocellular carcinoma models, EIF3M modulated tumor proliferation, migration, and activated oncogenic pathways like Wnt/β-catenin.

CONCLUSION: This study reveals that EIF3M expression correlates with immune infiltration and poor prognosis in multiple cancers. In vitro experiments in hepatocellular carcinoma models demonstrated that EIF3M critically regulates malignant cell behaviors. Collectively, our findings highlight EIF3M's value as a promising pan-cancer biomarker worthy of further investigation for its utility in prognosis prediction and as an indicator of immunotherapeutic response.

PMID:41341921 | PMC:PMC12669982 | DOI:10.3389/fmolb.2025.1697083

  •  

Detecting Sociodemographic Biases in the Content and Quality of Large Language Model–Generated Nursing Care: Cross-Sectional Simulation Study

Background: Large language models (LLMs) are increasingly applied in healthcare. However, concerns remain that their nursing care recommendations may reflect patients’ sociodemographic attributes rather than clinical needs. Objective: To investigate potential biases in nursing care plans generated by LLMs, we focused on whether outputs differ systematically based on patients’ sociodemographic characteristics and assessed the implications for equitable nursing care. Methods: We utilized a standardized clinical scenario with GPT to generate care plans for 96 sociodemographic identity combinations, drawing on 9,600 tests. We conducted statistical analyses (t-tests and ANOVA) to analyze how text length and the frequency of physiological and psychological nursing terms varied across sociodemographic factors. Additionally, we utilized Python for data processing and visualization to ensure methodological rigor throughout the study. Results: The analysis revealed significant sociodemographic biases in LLMs-generated nursing care plans. Female patients received shorter care plans (t = 4.864, P
  •  
❌