Normal view
-
TechCrunch
-
Humans& thinks coordination is the next frontier for AI, and they’re building a model to prove it
Humans&, a new startup founded by alumni of Anthropic, Meta, OpenAI, xAI, and Google DeepMind, is building the next generation of foundation models for collaboration, not chat.
-
TechCrunch
-
Former CEO of celeb fav gym Dogpound launches $5M fund to back wellness companies
Jenny Liu is a solo first-time GP looking to back underrepresented wellness founders.
Former CEO of celeb fav gym Dogpound launches $5M fund to back wellness companies
-
MIT Technology Review
-
“Dr. Google” had its issues. Can ChatGPT Health do better?
For the past two decades, there’s been a clear first step for anyone who starts experiencing new medical symptoms: Look them up online. The practice was so common that it gained the pejorative moniker “Dr. Google.” But times are changing, and many medical-information seekers are now using LLMs. According to OpenAI, 230 million people ask ChatGPT health-related queries each week. That’s the context around the launch of OpenAI’s new ChatGPT Health product, which debuted earlier this month.
“Dr. Google” had its issues. Can ChatGPT Health do better?
For the past two decades, there’s been a clear first step for anyone who starts experiencing new medical symptoms: Look them up online. The practice was so common that it gained the pejorative moniker “Dr. Google.” But times are changing, and many medical-information seekers are now using LLMs. According to OpenAI, 230 million people ask ChatGPT health-related queries each week.
That’s the context around the launch of OpenAI’s new ChatGPT Health product, which debuted earlier this month. It landed at an inauspicious time: Two days earlier, the news website SFGate had broken the story of Sam Nelson, a teenager who died of an overdose last year after extensive conversations with ChatGPT about how best to combine various drugs. In the wake of both pieces of news, multiple journalists questioned the wisdom of relying for medical advice on a tool that could cause such extreme harm.
Though ChatGPT Health lives in a separate sidebar tab from the rest of ChatGPT, it isn’t a new model. It’s more like a wrapper that provides one of OpenAI’s preexisting models with guidance and tools it can use to provide health advice—including some that allow it to access a user’s electronic medical records and fitness app data, if granted permission. There’s no doubt that ChatGPT and other large language models can make medical mistakes, and OpenAI emphasizes that ChatGPT Health is intended as an additional support, rather than a replacement for one’s doctor. But when doctors are unavailable or unable to help, people will turn to alternatives.
Some doctors see LLMs as a boon for medical literacy. The average patient might struggle to navigate the vast landscape of online medical information—and, in particular, to distinguish high-quality sources from polished but factually dubious websites—but LLMs can do that job for them, at least in theory. Treating patients who had searched for their symptoms on Google required “a lot of attacking patient anxiety [and] reducing misinformation,” says Marc Succi, an associate professor at Harvard Medical School and a practicing radiologist. But now, he says, “you see patients with a college education, a high school education, asking questions at the level of something an early med student might ask.”
The release of ChatGPT Health, and Anthropic’s subsequent announcement of new health integrations for Claude, indicate that the AI giants are increasingly willing to acknowledge and encourage health-related uses of their models. Such uses certainly come with risks, given LLMs’ well-documented tendencies to agree with users and make up information rather than admit ignorance.
But those risks also have to be weighed against potential benefits. There’s an analogy here to autonomous vehicles: When policymakers consider whether to allow Waymo in their city, the key metric is not whether its cars are ever involved in accidents but whether they cause less harm than the status quo of relying on human drivers. If Dr. ChatGPT is an improvement over Dr. Google—and early evidence suggests it may be—it could potentially lessen the enormous burden of medical misinformation and unnecessary health anxiety that the internet has created.
Pinning down the effectiveness of a chatbot such as ChatGPT or Claude for consumer health, however, is tricky. “It’s exceedingly difficult to evaluate an open-ended chatbot,” says Danielle Bitterman, the clinical lead for data science and AI at the Mass General Brigham health-care system. Large language models score well on medical licensing examinations, but those exams use multiple-choice questions that don’t reflect how people use chatbots to look up medical information.
Sirisha Rambhatla, an assistant professor of management science and engineering at the University of Waterloo, attempted to close that gap by evaluating how GPT-4 responded to licensing exam questions when it did not have access to a list of possible answers. Medical experts who evaluated the responses scored only about half of them as entirely correct. But multiple-choice exam questions are designed to be tricky enough that the answer options don’t give them entirely away, and they’re still a pretty distant approximation for the sort of thing that a user would type into ChatGPT.
A different study, which tested GPT-4o on more realistic prompts submitted by human volunteers, found that it answered medical questions correctly about 85% of the time. When I spoke with Amulya Yadav, an associate professor at Pennsylvania State University who runs the Responsible AI for Social Emancipation Lab and led the study, he made it clear that he wasn’t personally a fan of patient-facing medical LLMs. But he freely admits that, technically speaking, they seem up to the task—after all, he says, human doctors misdiagnose patients 10% to 15% of the time. “If I look at it dispassionately, it seems that the world is gonna change, whether I like it or not,” he says.
For people seeking medical information online, Yadav says, LLMs do seem to be a better choice than Google. Succi, the radiologist, also concluded that LLMs can be a better alternative to web search when he compared GPT-4’s responses to questions about common chronic medical conditions with the information presented in Google’s knowledge panel, the information box that sometimes appears on the right side of the search results.
Since Yadav’s and Succi’s studies appeared online, in the first half of 2025, OpenAI has released multiple new versions of GPT, and it’s reasonable to expect that GPT-5.2 would perform even better than its predecessors. But the studies do have important limitations: They focus on straightforward, factual questions, and they examine only brief interactions between users and chatbots or web search tools. Some of the weaknesses of LLMs—most notably their sycophancy and tendency to hallucinate—might be more likely to rear their heads in more extensive conversations and with people who are dealing with more complex problems. Reeva Lederman, a professor at the University of Melbourne who studies technology and health, notes that patients who don’t like the diagnosis or treatment recommendations that they receive from a doctor might seek out another opinion from an LLM—and the LLM, if it’s sycophantic, might encourage them to reject their doctor’s advice.
Some studies have found that LLMs will hallucinate and exhibit sycophancy in response to health-related prompts. For example, one study showed that GPT-4 and GPT-4o will happily accept and run with incorrect drug information included in a user’s question. In another, GPT-4o frequently concocted definitions for fake syndromes and lab tests mentioned in the user’s prompt. Given the abundance of medically dubious diagnoses and treatments floating around the internet, these patterns of LLM behavior could contribute to the spread of medical misinformation, particularly if people see LLMs as trustworthy.
OpenAI has reported that the GPT-5 series of models is markedly less sycophantic and prone to hallucination than their predecessors, so the results of these studies might not apply to ChatGPT Health. The company also evaluated the model that powers ChatGPT Health on its responses to health-specific questions, using their publicly available HeathBench benchmark. HealthBench rewards models that express uncertainty when appropriate, recommend that users seek medical attention when necessary, and refrain from causing users unnecessary stress by telling them their condition is more serious that it truly is. It’s reasonable to assume that the model underlying ChatGPT Health exhibited those behaviors in testing, though Bitterman notes that some of the prompts in HealthBench were generated by LLMs, not users, which could limit how well the benchmark translates into the real world.
An LLM that avoids alarmism seems like a clear improvement over systems that have people convincing themselves they have cancer after a few minutes of browsing. And as large language models, and the products built around them, continue to develop, whatever advantage Dr. ChatGPT has over Dr. Google will likely grow. The introduction of ChatGPT Health is certainly a move in that direction: By looking through your medical records, ChatGPT can potentially gain far more context about your specific health situation than could be included in any Google search, although numerous experts have cautioned against giving ChatGPT that access for privacy reasons.
Even if ChatGPT Health and other new tools do represent a meaningful improvement over Google searches, they could still conceivably have a negative effect on health overall. Much as automated vehicles, even if they are safer than human-driven cars, might still prove a net negative if they encourage people to use public transit less, LLMs could undermine users’ health if they induce people to rely on the internet instead of human doctors, even if they do increase the quality of health information available online.
Lederman says that this outcome is plausible. In her research, she has found that members of online communities centered on health tend to put their trust in users who express themselves well, regardless of the validity of the information they are sharing. Because ChatGPT communicates like an articulate person, some people might trust it too much, potentially to the exclusion of their doctor. But LLMs are certainly no replacement for a human doctor—at least not yet.
Correction 1/26: A previous version of this story incorrectly referred to the version of ChatGPT that Rambhatla evaluated. It was GPT-4, not GPT-4o.
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Responsibility Vacuum: Organizational Failure in Scaled Agent Systems
arXiv:2601.15059v1 Announce Type: new Abstract: Modern CI/CD pipelines integrating agent-generated code exhibit a structural failure in responsibility attribution. Decisions are executed through formally correct approval processes, yet no entity possesses both the authority to approve those decisions and the epistemic capacity to meaningfully understand their basis. We define this condition as responsibility vacuum: a state in which decisions occur, but responsibility cannot be attributed bec
The Responsibility Vacuum: Organizational Failure in Scaled Agent Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)
arXiv:2601.14298v1 Announce Type: cross Abstract: The AI era has ushered in Large Language Models (LLM) to the technological forefront, which has been much of the talk in 2023, and is likely to remain as such for many years to come. LLMs are the AI models that are the power house behind generative AI applications such as ChatGPT. These AI models, fueled by vast amounts of data and computational prowess, have unlocked remarkable capabilities, from human-like text generation to assisting with nat
Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)
-
cs.AI, q-bio.NC updates on arXiv.org
-
Towards Execution-Grounded Automated AI Research
arXiv:2601.14525v1 Announce Type: cross Abstract: Automated AI research holds great potential to accelerate scientific discovery. However, current LLMs often generate plausible-looking but ineffective ideas. Execution grounding may help, but it is unclear whether automated execution is feasible and whether LLMs can learn from the execution feedback. To investigate these, we first build an automated executor to implement ideas and launch large-scale parallel GPU experiments to verify their effec
Towards Execution-Grounded Automated AI Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
arXiv:2601.15161v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are particularly challenging as they often manifest as subtle clinical errors that evade detection by generic metrics, while expert-authored fine-grained rubrics remain costly to construct and difficult to scale. In this paper, we propose a retrieval-augmented multi-age
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
arXiv:2307.13658v3 Announce Type: replace-cross Abstract: We propose that future AI transparency and accountability regulations are based on an open global standard for exchanging information about AI systems, which allows co-existence of potentially conflicting local regulations. Then, we discuss key components of a lightweight and effective AI transparency and/or accountability regulation. To prevent overregulation, the proposed approach encourages collaboration between regulators and industr
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
arXiv:2509.19774v2 Announce Type: replace-cross Abstract: Electrocardiography (ECG) is the clinical gold standard for cardiovascular disease (CVD) assessment, yet continuous monitoring is constrained by the need for dedicated hardware and trained personnel. Photoplethysmography (PPG) is ubiquitous in wearable devices and readily scalable, but it lacks electrophysiological specificity, limiting diagnostic reliability. While generative methods aim to translate PPG into clinically useful ECG signa
PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
-
npj Digital Medicine
-
Large language models improve transferability of electronic health record-based predictions across countries and coding systems
npj Digital Medicine, Published online: 22 January 2026; doi:10.1038/s41746-026-02363-5Large language models improve transferability of electronic health record-based predictions across countries and coding systems
Large language models improve transferability of electronic health record-based predictions across countries and coding systems
npj Digital Medicine, Published online: 22 January 2026; doi:10.1038/s41746-026-02363-5
Large language models improve transferability of electronic health record-based predictions across countries and coding systems-
Cell
-
Multimodal AI generates virtual population for tumor microenvironment modeling
GigaTIME leverages multimodal AI to generate virtual multiplex immunofluorescence (mIF) profiles from standard H&E slides, enabling comprehensive tumor immune microenvironment modeling across a large (>14,000) and diverse patient population. This virtual approach unlocks new opportunities for large-scale clinical discoveries that were previously hindered by the scarcity of mIF data.
Multimodal AI generates virtual population for tumor microenvironment modeling
-
(Multiomics OR Omics) AND (Pancreatic)
-
Research progress in diagnosis and treatment of pancreatic cancer with mismatch repair and microsatellite instability
Clin Transl Oncol. 2026 Jan 21. doi: 10.1007/s12094-025-04214-3. Online ahead of print.ABSTRACTPancreatic cancer (PC), predominantly pancreatic ductal adenocarcinoma, remains one of the most lethal malignancies, largely due to late diagnosis and intrinsic resistance to conventional therapies. In recent years, mismatch repair deficiency (dMMR) and microsatellite instability-high (MSI-H) have emerged as clinically actionable biomarkers in a small but distinct subset of PC, accounting for approxima
Research progress in diagnosis and treatment of pancreatic cancer with mismatch repair and microsatellite instability
Clin Transl Oncol. 2026 Jan 21. doi: 10.1007/s12094-025-04214-3. Online ahead of print.
ABSTRACT
Pancreatic cancer (PC), predominantly pancreatic ductal adenocarcinoma, remains one of the most lethal malignancies, largely due to late diagnosis and intrinsic resistance to conventional therapies. In recent years, mismatch repair deficiency (dMMR) and microsatellite instability-high (MSI-H) have emerged as clinically actionable biomarkers in a small but distinct subset of PC, accounting for approximately 1-2% of cases. These tumors display unique molecular characteristics, including a high prevalence of wild-type KRAS and TP53, elevated tumor mutational burden, and recurrent kinase fusions, which together confer enhanced immunogenicity and increased sensitivity to immune checkpoint inhibitors (ICIs). In addition to their therapeutic relevance, dMMR/MSI-H status has important diagnostic implications for the identification of Lynch syndrome-associated pancreatic cancers, informing genetic counseling and familial risk assessment. This review summarizes current understanding of the molecular basis of mismatch repair deficiency and microsatellite instability in PC, evaluates available diagnostic approaches such as immunohistochemistry, polymerase chain reaction, and next-generation sequencing, and discusses the prognostic and predictive significance of dMMR/MSI-H status. Emerging clinical evidence supporting the use of ICIs in selected patients across neoadjuvant, adjuvant, and advanced disease settings is also reviewed, along with challenges related to assay discordance, tumor heterogeneity, and immunotherapy resistance. Finally, future directions are highlighted, emphasizing the need for standardized testing algorithms, integration of multi-omics and spatial profiling technologies, and prospective clinical studies to optimize precision treatment strategies for this rare but clinically meaningful subtype of pancreatic cancer.
PMID:41563663 | DOI:10.1007/s12094-025-04214-3
-
MIT Technology Review
-
Everyone wants AI sovereignty. No one can truly have it.
Governments plan to pour $1.3 trillion into AI infrastructure by 2030 to invest in “sovereign AI,” with the premise being that countries should be in control of their own AI capabilities. The funds include financing for domestic data centers, locally trained models, independent supply chains, and national talent pipelines. This is a response to real shocks: covid-era supply chain breakdowns, rising geopolitical tensions, and the war in Ukraine. But the pursuit of absolute autonomy is runnin
Everyone wants AI sovereignty. No one can truly have it.
Governments plan to pour $1.3 trillion into AI infrastructure by 2030 to invest in “sovereign AI,” with the premise being that countries should be in control of their own AI capabilities. The funds include financing for domestic data centers, locally trained models, independent supply chains, and national talent pipelines. This is a response to real shocks: covid-era supply chain breakdowns, rising geopolitical tensions, and the war in Ukraine.
But the pursuit of absolute autonomy is running into reality. AI supply chains are irreducibly global: Chips are designed in the US and manufactured in East Asia; models are trained on data sets drawn from multiple countries; applications are deployed across dozens of jurisdictions.
If sovereignty is to remain meaningful, it must shift from a defensive model of self-reliance to a vision that emphasizes the concept of orchestration, balancing national autonomy with strategic partnership.
Why infrastructure-first strategies hit walls
A November survey by Accenture found that 62% of European organizations are now seeking sovereign AI solutions, driven primarily by geopolitical anxiety rather than technical necessity. That figure rises to 80% in Denmark and 72% in Germany. The European Union has appointed its first Commissioner for Tech Sovereignty.
This year, $475 billion is flowing into AI data centers globally. In the United States, AI data centers accounted for roughly one-fifth of GDP growth in the second quarter of 2025. But the obstacle for other nations hoping to follow suit isn’t just money. It’s energy and physics. Global data center capacity is projected to hit 130 gigawatts by 2030, and for every $1 billion spent on these facilities, $125 million is needed for electricity networks. More than $750 billion in planned investment is already facing grid delays.
And it’s also talent. Researchers and entrepreneurs are mobile, drawn to ecosystems with access to capital, competitive wages, and rapid innovation cycles. Infrastructure alone won’t attract or retain world-class talent.
What works: An orchestrated sovereignty
What nations need isn’t sovereignty through isolation but through specialization and orchestration. This means choosing which capabilities you build, which you pursue through partnership, and where you can genuinely lead in shaping the global AI landscape.
The most successful AI strategies don’t try to replicate Silicon Valley; they identify specific advantages and build partnerships around them.
Singapore offers a model. Rather than seeking to duplicate massive infrastructure, it invested in governance frameworks, digital-identity platforms, and applications of AI in logistics and finance, areas where it can realistically compete.
Israel shows a different path. Its strength lies in a dense network of startups and military-adjacent research institutions delivering outsize influence despite the country’s small size.
South Korea is instructive too. While it has national champions like Samsung and Naver, these firms still partner with Microsoft and Nvidia on infrastructure. That’s deliberate collaboration reflecting strategic oversight, not dependence.
Even China, despite its scale and ambition, cannot secure full-stack autonomy. Its reliance on global research networks and on foreign lithography equipment, such as extreme ultraviolet systems needed to manufacture advanced chips and GPU architectures, shows the limits of techno-nationalism.
The pattern is clear: Nations that specialize and partner strategically can outperform those trying to do everything alone.
Three ways to align ambition with reality
1. Measure added value, not inputs.
Sovereignty isn’t how many petaflops you own. It’s how many lives you improve and how fast the economy grows. Real sovereignty is the ability to innovate in support of national priorities such as productivity, resilience, and sustainability while maintaining freedom to shape governance and standards.
Nations should track the use of AI in health care and monitor how the technology’s adoption correlates with manufacturing productivity, patent citations, and international research collaborations. The goal is to ensure that AI ecosystems generate inclusive and lasting economic and social value.
2. Cultivate a strong AI innovation ecosystem.
Build infrastructure, but also build the ecosystem around it: research institutions, technical education, entrepreneurship support, and public-private talent development. Infrastructure without skilled talent and vibrant networks cannot deliver a lasting competitive advantage.
3. Build global partnerships.
Strategic partnerships enable nations to pool resources, lower infrastructure costs, and access complementary expertise. Singapore’s work with global cloud providers and the EU’s collaborative research programs show how nations advance capabilities faster through partnership than through isolation. Rather than competing to set dominant standards, nations should collaborate on interoperable frameworks for transparency, safety, and accountability.
What’s at stake
Overinvesting in independence fragments markets and slows cross-border innovation, which is the foundation of AI progress. When strategies focus too narrowly on control, they sacrifice the agility needed to compete.
The cost of getting this wrong isn’t just wasted capital—it’s a decade of falling behind. Nations that double down on infrastructure-first strategies risk ending up with expensive data centers running yesterday’s models, while competitors that choose strategic partnerships iterate faster, attract better talent, and shape the standards that matter.
The winners will be those who define sovereignty not as separation, but as participation plus leadership—choosing who they depend on, where they build, and which global rules they shape. Strategic interdependence may feel less satisfying than independence, but it’s real, it is achievable, and it will separate the leaders from the followers over the next decade.
The age of intelligent systems demands intelligent strategies—ones that measure success not by infrastructure owned, but by problems solved. Nations that embrace this shift won’t just participate in the AI economy; they’ll shape it. That’s sovereignty worth pursuing.
Cathy Li is head of the Centre for AI Excellence at the World Economic Forum.
-
Journal of Medical Internet Research
-
Communication Strategies to Promote Patient Engagement in Telemedicine: Systematic Review
Background: The rapid growth of telemedicine offers convenience, flexibility, and accessibility for patients to have health care services worldwide. To succeed in telemedicine, health care practitioners and telemedicine tools must engage patients through effective communication. However, a research gap exists in understanding the communication strategies used in telemedicine and how they effectively engage patients. Objective: This study aims to identify communication strategies influencing pati
Communication Strategies to Promote Patient Engagement in Telemedicine: Systematic Review
-
Omics in Hepatocellular
-
United multi-omics and machine learning refine regulatory T cell-defined hepatocellular carcinoma subtypes
iScience. 2025 Dec 3;29(1):114328. doi: 10.1016/j.isci.2025.114328. eCollection 2026 Jan 16.ABSTRACTHepatocellular carcinoma (HCC) is highly heterogeneous and aggressive, and the absence of precision individual treatment regimen enables repeated immune escape. Exploiting regulatory T cell (Treg) marker genes as key classifiers, we used 10 clustering algorithms to integrate the multi-omics HCC patient data and combined them with 10 machine learning (ML) algorithms to delineate molecular subtypes
United multi-omics and machine learning refine regulatory T cell-defined hepatocellular carcinoma subtypes
iScience. 2025 Dec 3;29(1):114328. doi: 10.1016/j.isci.2025.114328. eCollection 2026 Jan 16.
ABSTRACT
Hepatocellular carcinoma (HCC) is highly heterogeneous and aggressive, and the absence of precision individual treatment regimen enables repeated immune escape. Exploiting regulatory T cell (Treg) marker genes as key classifiers, we used 10 clustering algorithms to integrate the multi-omics HCC patient data and combined them with 10 machine learning (ML) algorithms to delineate molecular subtypes predictive of prognosis and immune response. We identified two cancer subtypes (CSs) that are associated with prognosis, with the second subtype (CS2) showing the most favorable prognostic outcomes. Subsequently, 9 key genes were screened for HCC model scoring, stratifying patients into low-risk (good prognosis, responsive to immunotherapy) and high-risk (poor outcome, not responsive to immunotherapy) groups. The high-risk group may be effective against the mTOR inhibitor AZD8055. Comprehensive multi-omics data and multiple ML algorithms offer key insights into HCC occurrence and evolution, with model scores guiding patient prognosis and treatment clinically.
PMID:41561382 | PMC:PMC12814435 | DOI:10.1016/j.isci.2025.114328
-
cs.AI, q-bio.NC updates on arXiv.org
-
Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward
arXiv:2601.13122v1 Announce Type: new Abstract: Modern general-purpose AI systems made using large language and vision models, are capable of performing a range of tasks like writing text articles, generating and debugging codes, querying databases, and translating from one language to another, which has made them quite popular across industries. However, there are risks like hallucinations, toxicity, and stereotypes in their output that make them untrustworthy. We review various risks and vuln
Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward
-
cs.AI, q-bio.NC updates on arXiv.org
-
DeepEvidence: Empowering Biomedical Discovery with Deep Knowledge Graph Research
arXiv:2601.11560v1 Announce Type: cross Abstract: Biomedical knowledge graphs (KGs) encode vast, heterogeneous information spanning literature, genes, pathways, drugs, diseases, and clinical trials, but leveraging them collectively for scientific discovery remains difficult. Their structural differences, continual evolution, and limited cross-resource alignment require substantial manual integration, limiting the depth and scale of knowledge exploration. We introduce DeepEvidence, an AI-agent f
DeepEvidence: Empowering Biomedical Discovery with Deep Knowledge Graph Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
arXiv:2601.12471v1 Announce Type: cross Abstract: Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce MedAbstain, a unified benchmark and evaluation protocol for abstention in medical multiple-choice question answering (MCQA) -- a discrete-choice setting that generalizes to agentic action selection -- integra
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Cloud-based Multi-Agentic Workflow for Science
arXiv:2601.12607v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become ubiquitous across various scientific domains, their lack of ability to perform complex tasks like running simulations or to make complex decisions limits their utility. LLM-based agents bridge this gap due to their ability to call external resources and tools and thus are now rapidly gaining popularity. However, coming up with a workflow that can balance the models, cloud providers, and external resources i
A Cloud-based Multi-Agentic Workflow for Science
-
cs.AI, q-bio.NC updates on arXiv.org
-
SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding
arXiv:2601.12805v1 Announce Type: cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. However, their ability to reliably reason from gene-level knowledge to functional understanding, However, their ability to reliably reason from gene-level knowledge to functional understanding, a core requirement for knowledge-enhanced cell atlas interpretation, remains largely underexplored. To address this gap