Normal view
-
TechCrunch
-
Humans& thinks coordination is the next frontier for AI, and they’re building a model to prove it
Humans&, a new startup founded by alumni of Anthropic, Meta, OpenAI, xAI, and Google DeepMind, is building the next generation of foundation models for collaboration, not chat.
-
TechCrunch
-
Former CEO of celeb fav gym Dogpound launches $5M fund to back wellness companies
Jenny Liu is a solo first-time GP looking to back underrepresented wellness founders.
Former CEO of celeb fav gym Dogpound launches $5M fund to back wellness companies
-
MIT Technology Review
-
“Dr. Google” had its issues. Can ChatGPT Health do better?
For the past two decades, there’s been a clear first step for anyone who starts experiencing new medical symptoms: Look them up online. The practice was so common that it gained the pejorative moniker “Dr. Google.” But times are changing, and many medical-information seekers are now using LLMs. According to OpenAI, 230 million people ask ChatGPT health-related queries each week. That’s the context around the launch of OpenAI’s new ChatGPT Health product, which debuted earlier this month.
“Dr. Google” had its issues. Can ChatGPT Health do better?
For the past two decades, there’s been a clear first step for anyone who starts experiencing new medical symptoms: Look them up online. The practice was so common that it gained the pejorative moniker “Dr. Google.” But times are changing, and many medical-information seekers are now using LLMs. According to OpenAI, 230 million people ask ChatGPT health-related queries each week.
That’s the context around the launch of OpenAI’s new ChatGPT Health product, which debuted earlier this month. It landed at an inauspicious time: Two days earlier, the news website SFGate had broken the story of Sam Nelson, a teenager who died of an overdose last year after extensive conversations with ChatGPT about how best to combine various drugs. In the wake of both pieces of news, multiple journalists questioned the wisdom of relying for medical advice on a tool that could cause such extreme harm.
Though ChatGPT Health lives in a separate sidebar tab from the rest of ChatGPT, it isn’t a new model. It’s more like a wrapper that provides one of OpenAI’s preexisting models with guidance and tools it can use to provide health advice—including some that allow it to access a user’s electronic medical records and fitness app data, if granted permission. There’s no doubt that ChatGPT and other large language models can make medical mistakes, and OpenAI emphasizes that ChatGPT Health is intended as an additional support, rather than a replacement for one’s doctor. But when doctors are unavailable or unable to help, people will turn to alternatives.
Some doctors see LLMs as a boon for medical literacy. The average patient might struggle to navigate the vast landscape of online medical information—and, in particular, to distinguish high-quality sources from polished but factually dubious websites—but LLMs can do that job for them, at least in theory. Treating patients who had searched for their symptoms on Google required “a lot of attacking patient anxiety [and] reducing misinformation,” says Marc Succi, an associate professor at Harvard Medical School and a practicing radiologist. But now, he says, “you see patients with a college education, a high school education, asking questions at the level of something an early med student might ask.”
The release of ChatGPT Health, and Anthropic’s subsequent announcement of new health integrations for Claude, indicate that the AI giants are increasingly willing to acknowledge and encourage health-related uses of their models. Such uses certainly come with risks, given LLMs’ well-documented tendencies to agree with users and make up information rather than admit ignorance.
But those risks also have to be weighed against potential benefits. There’s an analogy here to autonomous vehicles: When policymakers consider whether to allow Waymo in their city, the key metric is not whether its cars are ever involved in accidents but whether they cause less harm than the status quo of relying on human drivers. If Dr. ChatGPT is an improvement over Dr. Google—and early evidence suggests it may be—it could potentially lessen the enormous burden of medical misinformation and unnecessary health anxiety that the internet has created.
Pinning down the effectiveness of a chatbot such as ChatGPT or Claude for consumer health, however, is tricky. “It’s exceedingly difficult to evaluate an open-ended chatbot,” says Danielle Bitterman, the clinical lead for data science and AI at the Mass General Brigham health-care system. Large language models score well on medical licensing examinations, but those exams use multiple-choice questions that don’t reflect how people use chatbots to look up medical information.
Sirisha Rambhatla, an assistant professor of management science and engineering at the University of Waterloo, attempted to close that gap by evaluating how GPT-4 responded to licensing exam questions when it did not have access to a list of possible answers. Medical experts who evaluated the responses scored only about half of them as entirely correct. But multiple-choice exam questions are designed to be tricky enough that the answer options don’t give them entirely away, and they’re still a pretty distant approximation for the sort of thing that a user would type into ChatGPT.
A different study, which tested GPT-4o on more realistic prompts submitted by human volunteers, found that it answered medical questions correctly about 85% of the time. When I spoke with Amulya Yadav, an associate professor at Pennsylvania State University who runs the Responsible AI for Social Emancipation Lab and led the study, he made it clear that he wasn’t personally a fan of patient-facing medical LLMs. But he freely admits that, technically speaking, they seem up to the task—after all, he says, human doctors misdiagnose patients 10% to 15% of the time. “If I look at it dispassionately, it seems that the world is gonna change, whether I like it or not,” he says.
For people seeking medical information online, Yadav says, LLMs do seem to be a better choice than Google. Succi, the radiologist, also concluded that LLMs can be a better alternative to web search when he compared GPT-4’s responses to questions about common chronic medical conditions with the information presented in Google’s knowledge panel, the information box that sometimes appears on the right side of the search results.
Since Yadav’s and Succi’s studies appeared online, in the first half of 2025, OpenAI has released multiple new versions of GPT, and it’s reasonable to expect that GPT-5.2 would perform even better than its predecessors. But the studies do have important limitations: They focus on straightforward, factual questions, and they examine only brief interactions between users and chatbots or web search tools. Some of the weaknesses of LLMs—most notably their sycophancy and tendency to hallucinate—might be more likely to rear their heads in more extensive conversations and with people who are dealing with more complex problems. Reeva Lederman, a professor at the University of Melbourne who studies technology and health, notes that patients who don’t like the diagnosis or treatment recommendations that they receive from a doctor might seek out another opinion from an LLM—and the LLM, if it’s sycophantic, might encourage them to reject their doctor’s advice.
Some studies have found that LLMs will hallucinate and exhibit sycophancy in response to health-related prompts. For example, one study showed that GPT-4 and GPT-4o will happily accept and run with incorrect drug information included in a user’s question. In another, GPT-4o frequently concocted definitions for fake syndromes and lab tests mentioned in the user’s prompt. Given the abundance of medically dubious diagnoses and treatments floating around the internet, these patterns of LLM behavior could contribute to the spread of medical misinformation, particularly if people see LLMs as trustworthy.
OpenAI has reported that the GPT-5 series of models is markedly less sycophantic and prone to hallucination than their predecessors, so the results of these studies might not apply to ChatGPT Health. The company also evaluated the model that powers ChatGPT Health on its responses to health-specific questions, using their publicly available HeathBench benchmark. HealthBench rewards models that express uncertainty when appropriate, recommend that users seek medical attention when necessary, and refrain from causing users unnecessary stress by telling them their condition is more serious that it truly is. It’s reasonable to assume that the model underlying ChatGPT Health exhibited those behaviors in testing, though Bitterman notes that some of the prompts in HealthBench were generated by LLMs, not users, which could limit how well the benchmark translates into the real world.
An LLM that avoids alarmism seems like a clear improvement over systems that have people convincing themselves they have cancer after a few minutes of browsing. And as large language models, and the products built around them, continue to develop, whatever advantage Dr. ChatGPT has over Dr. Google will likely grow. The introduction of ChatGPT Health is certainly a move in that direction: By looking through your medical records, ChatGPT can potentially gain far more context about your specific health situation than could be included in any Google search, although numerous experts have cautioned against giving ChatGPT that access for privacy reasons.
Even if ChatGPT Health and other new tools do represent a meaningful improvement over Google searches, they could still conceivably have a negative effect on health overall. Much as automated vehicles, even if they are safer than human-driven cars, might still prove a net negative if they encourage people to use public transit less, LLMs could undermine users’ health if they induce people to rely on the internet instead of human doctors, even if they do increase the quality of health information available online.
Lederman says that this outcome is plausible. In her research, she has found that members of online communities centered on health tend to put their trust in users who express themselves well, regardless of the validity of the information they are sharing. Because ChatGPT communicates like an articulate person, some people might trust it too much, potentially to the exclusion of their doctor. But LLMs are certainly no replacement for a human doctor—at least not yet.
Correction 1/26: A previous version of this story incorrectly referred to the version of ChatGPT that Rambhatla evaluated. It was GPT-4, not GPT-4o.
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Responsibility Vacuum: Organizational Failure in Scaled Agent Systems
arXiv:2601.15059v1 Announce Type: new Abstract: Modern CI/CD pipelines integrating agent-generated code exhibit a structural failure in responsibility attribution. Decisions are executed through formally correct approval processes, yet no entity possesses both the authority to approve those decisions and the epistemic capacity to meaningfully understand their basis. We define this condition as responsibility vacuum: a state in which decisions occur, but responsibility cannot be attributed bec
The Responsibility Vacuum: Organizational Failure in Scaled Agent Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)
arXiv:2601.14298v1 Announce Type: cross Abstract: The AI era has ushered in Large Language Models (LLM) to the technological forefront, which has been much of the talk in 2023, and is likely to remain as such for many years to come. LLMs are the AI models that are the power house behind generative AI applications such as ChatGPT. These AI models, fueled by vast amounts of data and computational prowess, have unlocked remarkable capabilities, from human-like text generation to assisting with nat
Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)
-
cs.AI, q-bio.NC updates on arXiv.org
-
An Optimized Decision Tree-Based Framework for Explainable IoT Anomaly Detection
arXiv:2601.14305v1 Announce Type: cross Abstract: The increase in the number of Internet of Things (IoT) devices has tremendously increased the attack surface of cyber threats thus making a strong intrusion detection system (IDS) with a clear explanation of the process essential towards resource-constrained environments. Nevertheless, current IoT IDS systems are usually traded off with detection quality, model elucidability, and computational effectiveness, thus the deployment on IoT devices. T
An Optimized Decision Tree-Based Framework for Explainable IoT Anomaly Detection
-
cs.AI, q-bio.NC updates on arXiv.org
-
Towards Execution-Grounded Automated AI Research
arXiv:2601.14525v1 Announce Type: cross Abstract: Automated AI research holds great potential to accelerate scientific discovery. However, current LLMs often generate plausible-looking but ineffective ideas. Execution grounding may help, but it is unclear whether automated execution is feasible and whether LLMs can learn from the execution feedback. To investigate these, we first build an automated executor to implement ideas and launch large-scale parallel GPU experiments to verify their effec
Towards Execution-Grounded Automated AI Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
Multimodal system for skin cancer detection
arXiv:2601.14822v1 Announce Type: cross Abstract: Melanoma detection is vital for early diagnosis and effective treatment. While deep learning models on dermoscopic images have shown promise, they require specialized equipment, limiting their use in broader clinical settings. This study introduces a multi-modal melanoma detection system using conventional photo images, making it more accessible and versatile. Our system integrates image data with tabular metadata, such as patient demographics a
Multimodal system for skin cancer detection
-
cs.AI, q-bio.NC updates on arXiv.org
-
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
arXiv:2601.15161v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are particularly challenging as they often manifest as subtle clinical errors that evade detection by generic metrics, while expert-authored fine-grained rubrics remain costly to construct and difficult to scale. In this paper, we propose a retrieval-augmented multi-age
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
arXiv:2505.20310v2 Announce Type: replace Abstract: Meta-analysis is a systematic research methodology that synthesizes data from multiple existing studies to derive comprehensive conclusions. This approach not only mitigates limitations inherent in individual studies but also facilitates novel discoveries through integrated data analysis. Traditional meta-analysis involves a complex multi-stage pipeline including literature retrieval, paper screening, and data extraction, which demands substan
Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
-
cs.AI, q-bio.NC updates on arXiv.org
-
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
arXiv:2307.13658v3 Announce Type: replace-cross Abstract: We propose that future AI transparency and accountability regulations are based on an open global standard for exchanging information about AI systems, which allows co-existence of potentially conflicting local regulations. Then, we discuss key components of a lightweight and effective AI transparency and/or accountability regulation. To prevent overregulation, the proposed approach encourages collaboration between regulators and industr
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
arXiv:2509.19774v2 Announce Type: replace-cross Abstract: Electrocardiography (ECG) is the clinical gold standard for cardiovascular disease (CVD) assessment, yet continuous monitoring is constrained by the need for dedicated hardware and trained personnel. Photoplethysmography (PPG) is ubiquitous in wearable devices and readily scalable, but it lacks electrophysiological specificity, limiting diagnostic reliability. While generative methods aim to translate PPG into clinically useful ECG signa
PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
-
Cell
-
Large-scale spatial profiling of the tumor microenvironment
Multiplex immunofluorescence proteomics is a powerful method in spatial biology to decipher cell types, states, and architecture in tissues. In this issue of Cell, Valanarasu et al. develop GigaTIME, an expansive population-scale analysis of the tumor immune microenvironment of over 14,000 patients across 24 cancer types, enabling clinical discovery and patient stratification.
Large-scale spatial profiling of the tumor microenvironment
-
npj Digital Medicine
-
Large language models improve transferability of electronic health record-based predictions across countries and coding systems
npj Digital Medicine, Published online: 22 January 2026; doi:10.1038/s41746-026-02363-5Large language models improve transferability of electronic health record-based predictions across countries and coding systems
Large language models improve transferability of electronic health record-based predictions across countries and coding systems
npj Digital Medicine, Published online: 22 January 2026; doi:10.1038/s41746-026-02363-5
Large language models improve transferability of electronic health record-based predictions across countries and coding systems-
Cell
-
Multimodal AI generates virtual population for tumor microenvironment modeling
GigaTIME leverages multimodal AI to generate virtual multiplex immunofluorescence (mIF) profiles from standard H&E slides, enabling comprehensive tumor immune microenvironment modeling across a large (>14,000) and diverse patient population. This virtual approach unlocks new opportunities for large-scale clinical discoveries that were previously hindered by the scarcity of mIF data.
Multimodal AI generates virtual population for tumor microenvironment modeling
-
TechCrunch
-
OpenEvidence hits $12B valuation, with new round led by Thrive, DST
The medical info database has doubled in valuation since last raise in October, despite encroachment from model makers.
OpenEvidence hits $12B valuation, with new round led by Thrive, DST
-
Latest Science News -- ScienceDaily
-
A simple blood test mismatch linked to kidney failure and death
A major global study suggests that a hidden mismatch between two common blood tests could quietly signal serious trouble ahead. When results from creatinine and cystatin C—two markers used to assess kidney health—don’t line up, the risk of kidney failure, heart disease, and even death appears to rise sharply. Researchers found that this gap is especially common among hospitalized and older patients, and that relying on just one test may miss early warning signs.
A simple blood test mismatch linked to kidney failure and death
-
STAT

-
STAT+: OpenEvidence raises $250 million, doubling its valuation
OpenEvidence, maker of a popular chatbot that helps doctors search clinical evidence, on Wednesday announced $250 million in new funding. The new round led by Thrive Capital and DST Global values OpenEvidence at $12 billion, and the company has announced $735 million in funding in the last 12 months. OpenEvidence is free to use by any clinician with a national provider identifier number. The company’s primary business model is advertising shown to clinicians. Founded in 2022, OpenEvidence
STAT+: OpenEvidence raises $250 million, doubling its valuation
OpenEvidence, maker of a popular chatbot that helps doctors search clinical evidence, on Wednesday announced $250 million in new funding.
The new round led by Thrive Capital and DST Global values OpenEvidence at $12 billion, and the company has announced $735 million in funding in the last 12 months. OpenEvidence is free to use by any clinician with a national provider identifier number. The company’s primary business model is advertising shown to clinicians.
Founded in 2022, OpenEvidence is one of the most prominent and best-funded companies from a wave of health artificial intelligence companies that emerged since the widespread availability of large language models. Reflecting on the eye-popping fundraising for health AI, a Silicon Valley Bank report released earlier in January raised an eyebrow at the ability of companies like OpenEvidence to deliver on their stratospheric valuations with advertising and software-as-a-service business models. “It won’t be a surprise to see them tap the value of the data they’re already collecting” to offer more services to pharma or other customers, the authors wrote.
Continue to STAT+ to read the full story…


© Adobe
-
(Multiomics OR Omics) AND (Pancreatic)
-
Research progress in diagnosis and treatment of pancreatic cancer with mismatch repair and microsatellite instability
Clin Transl Oncol. 2026 Jan 21. doi: 10.1007/s12094-025-04214-3. Online ahead of print.ABSTRACTPancreatic cancer (PC), predominantly pancreatic ductal adenocarcinoma, remains one of the most lethal malignancies, largely due to late diagnosis and intrinsic resistance to conventional therapies. In recent years, mismatch repair deficiency (dMMR) and microsatellite instability-high (MSI-H) have emerged as clinically actionable biomarkers in a small but distinct subset of PC, accounting for approxima
Research progress in diagnosis and treatment of pancreatic cancer with mismatch repair and microsatellite instability
Clin Transl Oncol. 2026 Jan 21. doi: 10.1007/s12094-025-04214-3. Online ahead of print.
ABSTRACT
Pancreatic cancer (PC), predominantly pancreatic ductal adenocarcinoma, remains one of the most lethal malignancies, largely due to late diagnosis and intrinsic resistance to conventional therapies. In recent years, mismatch repair deficiency (dMMR) and microsatellite instability-high (MSI-H) have emerged as clinically actionable biomarkers in a small but distinct subset of PC, accounting for approximately 1-2% of cases. These tumors display unique molecular characteristics, including a high prevalence of wild-type KRAS and TP53, elevated tumor mutational burden, and recurrent kinase fusions, which together confer enhanced immunogenicity and increased sensitivity to immune checkpoint inhibitors (ICIs). In addition to their therapeutic relevance, dMMR/MSI-H status has important diagnostic implications for the identification of Lynch syndrome-associated pancreatic cancers, informing genetic counseling and familial risk assessment. This review summarizes current understanding of the molecular basis of mismatch repair deficiency and microsatellite instability in PC, evaluates available diagnostic approaches such as immunohistochemistry, polymerase chain reaction, and next-generation sequencing, and discusses the prognostic and predictive significance of dMMR/MSI-H status. Emerging clinical evidence supporting the use of ICIs in selected patients across neoadjuvant, adjuvant, and advanced disease settings is also reviewed, along with challenges related to assay discordance, tumor heterogeneity, and immunotherapy resistance. Finally, future directions are highlighted, emphasizing the need for standardized testing algorithms, integration of multi-omics and spatial profiling technologies, and prospective clinical studies to optimize precision treatment strategies for this rare but clinically meaningful subtype of pancreatic cancer.
PMID:41563663 | DOI:10.1007/s12094-025-04214-3
-
ScienceDirect Publication: Artificial Intelligence in Medicine
-
BRLA-DDI: A novel framework for drug–drug interaction extraction
Publication date: April 2026Source: Artificial Intelligence in Medicine, Volume 174Author(s): Zhu Yuan, Shuailiang Zhang, Zongjin Li, Huiyun Zhang, Huaqi Zhang, Yaxun Jia
BRLA-DDI: A novel framework for drug–drug interaction extraction
Publication date: April 2026
Source: Artificial Intelligence in Medicine, Volume 174
Author(s): Zhu Yuan, Shuailiang Zhang, Zongjin Li, Huiyun Zhang, Huaqi Zhang, Yaxun Jia