❌

Normal view

LongCodeBench: Evaluating Coding LLMs at 1M Context Windows

arXiv:2505.07897v3 Announce Type: replace-cross Abstract: Context lengths for models have grown rapidly, from thousands to millions of tokens in just a few years. The extreme context sizes of modern long-context models have made it difficult to construct realistic long-context benchmarks -- not only due to the cost of collecting million-context tasks but also in identifying realistic scenarios that require significant contexts. We identify code comprehension and repair as a natural testbed and challenge task for long-context models and introduce LongCodeBench (LCB), a benchmark to test LLM coding abilities in long-context scenarios. Our benchmark tests both the comprehension and repair capabilities of LCLMs in realistic and important settings by drawing from real-world GitHub issues and constructing QA (LongCodeQA) and bug fixing (LongSWE-Bench) tasks. We carefully stratify the complexity of our benchmark, enabling us to evaluate models across different scales -- ranging from Qwen2.5 14B Instruct to Google's flagship Gemini model. We find that long-context remains a weakness for all models, with performance drops such as from 29% to 3% for Claude 3.5 Sonnet, or from 70.2% to 40% for Qwen2.5. The LCB dataset is available publicly at https://huggingface.co/datasets/Steefano/LCB and the codebase to replicate the work on this paper at https://github.com/Zteefano/long-code-bench.

Assessing Large Language Models in Building a Structured Dataset From AskDocs Subreddit Data: Methodological Study

Background: In an era marked by the blooming reliance on digital platforms for healthcare consultation, the subreddit r/AskDocs has emerged as a pivotal forum. However, the vast, unstructured nature of forum data presents a formidable challenge; the extraction and meaningful analysis of such data require advanced tools that can navigate the complexities of language and context inherent in user-generated content. Objective: Our objective was to evaluate employing Large Language Models (LLMs) to systematically transform the rich, unstructured textual data from AskDocs into a structured dataset, an approach that aligns more closely with human cognitive processes compared to traditional data extraction methods. Methods: We developed a dataset of Reddit posts from r/AskDocs by extracting key information via human annotators. Then using specially engineered prompts we used state-of-the-art Large Language Models (LLMs) to extract data from posts and compared the results. The variation in the LLMs were further compared to the humans to show similarity. Results: Our findings indicate that LLMs not only match but, in several aspects, surpass even highly educated humans in extracting information, including both demographic and context details, from unstructured texts. Conclusions: This study not only validates the use of LLMs for analyzing digital healthcare communications but also opens new avenues for understanding online behaviors and interactions, signaling a shift towards more sophisticated methodologies in digital research and practice.

Alternatives to animal testing are the future — it’s time that journals, funders and scientists embrace them

Nature, Published online: 20 October 2025; doi:10.1038/d41586-025-03344-6

Biomedical research techniques that don’t involve the use of animals are gaining momentum, but those using innovative approaches still face resistance from some quarters.

Implementing a Digital Mental Health Intervention—the Lumi Nova App—to Support Children With Anxiety in Economically Disadvantaged Areas: Mixed Methods Study

Background: Anxiety is one of the most common mental health problems experienced by children worldwide. In the UK, many children experiencing anxiety do not receive adequate or timely help. Children living in economically-disadvantaged areas experience more mental health problems than those living in high income areas and are less able to engage in activities that can have a positive or protective impact on their mental health. The need for providing low-cost, accessible and engaging mental health interventions for children living in these areas is high. Objective: The study aimed to explore how a digital mental health therapeutic, ‘Lumi Nova: Tales of Courage’, could be used to support children living with anxiety in economically-disadvantaged areas. Methods: A mixed method study design was used to explore the implementation of Lumi Nova using a supported delivery model with mental health teams based in the North of England. Quantitative data collection on recruitment and engagement patterns were collected and analysed. Qualitative research explored children, parent and practitioner views and experiences with the Lumi Nova app. Results: 113 children were consented to use Lumi Nova and 98 (87%) accessed the intervention at least once. Qualitative semi-structured interviews found that children, their parents and practitioners viewed the Lumi Nova app positively. Quantitative analysis of the recruitment data suggested the feasibility of a future larger roll-out. Analysis of usage data demonstrated varied patterns of engagement with the intervention. The frequency and duration of usage varied across children, as did the activities completed within the game: almost half (49%) completed three in-game challenges indicating progression through the treatment pathway. Conclusions: The study demonstrated that a digital mental health intervention could be successfully deployed within economically-disadvantaged areas in the UK to support children experiencing anxiety. Expected barriers to the deployment of digital mental health interventions in economically-disadvantaged areas (e.g. lack of access to smartphones, data plans, lack of technical skills) were not reported. Digital mental health interventions have the potential to address current gaps in mental health provision for disadvantaged individuals and communities.

FUSION: a web-based application for in-depth exploration of multi-omics data with brightfield histology

Nat Commun. 2025 Sep 25;16(1):8388. doi: 10.1038/s41467-025-63050-9.

ABSTRACT

Spatial technologies examining the cell and tissue microenvironment at near single-cell resolution are revealing important molecular insights. However, few tools enable integrated, interactive analysis of spatial-omics with tissue morphology in the same functional tissue unit. Here, we present FUSION (Functional Unit State Identification in Whole Slide Images), a web-based platform for visualizing and analyzing spatial-omics data with high-resolution histology. FUSION provides workflows for assessing cell compositions, quantitative morphometrics, and comparative tissue analyses. We demonstrate applicability across spatial assays, including 10x Visium, Visium HD, 10x Xenium, Cell DIVE, and PhenoCycler, applied to healthy and diseased tissues from kidney, small intestine, lung, and skin in the Human BioMolecular Atlas Program. FUSION is cloud-based, open-source, and accessible at https://fusion.hubmapconsortium.org/ , hosting over 50 paired datasets and tutorials. In a series of use cases, we show its capacity to distinguish renal glomeruli injury states, quantify morphometric changes, and characterize fibrosis with immune infiltration.

PMID:40998789 | PMC:PMC12462499 | DOI:10.1038/s41467-025-63050-9

Expanding care coordination in an integrated health system through causal machine learning

npj Digital Medicine, Published online: 24 September 2025; doi:10.1038/s41746-025-01925-3

Expanding care coordination in an integrated health system through causal machine learning

Opinion: Four reasons why generative AI chatbots could lead to psychosis in vulnerable people

18 September 2025 at 16:30

Three scholars discovered a strange mirror deep in the forest. It spoke to them in a soothing voice and answered all their questions warmly, knowledgeably, and eloquently.

The captivated scholars became obsessed, whispering one secret after another to the mirror. It replied with affection, promise, and meaning that kept them returning to it. They began ignoring one another, each convinced the mirror “understood” them best.

Read the rest…

© Adobe

Functions of the global health system in a new era

Nature Medicine, Published online: 11 September 2025; doi:10.1038/s41591-025-03936-9

In an irrevocably changed landscape, reform of the global health system needs to answer key questions on functions, what should be delivered in different contexts and at different levels, and how the system should operate.

Human interpretable grammar encodes multicellular systems biology models to democratize virtual cell laboratories

We developed a plain text modeling language—a cell behavior hypothesis grammar—to easily build virtual cell models and connect them to data, helping scientists to unlock the hidden dynamics of tissues. We provide examples showing how to use them in virtual experiments exploring how cancer responds to the cells in its environment and how the brain forms layers in development.

Circulating tumour cells & circulating tumour DNA in patients with resectable colorectal liver metastases (MIRACLE): a prospective, observational biomarker study

EClinicalMedicine. 2025 Aug 12;87:103406. doi: 10.1016/j.eclinm.2025.103406. eCollection 2025 Sep.

ABSTRACT

BACKGROUND: Recurrence risk after curative surgery for colorectal liver metastases (CRLM) remains high, underlining the need to identify prognostic markers enabling more individualised treatment approaches.

METHODS: In the MIRACLE, a prospective, observational biomarker study, a total of 188 patients with isolated, resectable CRLM without (neo)adjuvant chemotherapy were included between October 2015 and December 2021. Blood samples were collected before surgery (baseline) and three weeks after surgery. The primary objective was to assess the potential association between postoperative circulating tumour DNA (ctDNA) detection and recurrence of disease for patients with resectable CRLM within one year after resection. The secondary objective was the association between recurrence of disease within one year and detection of circulating tumour cells (CTCs). Baseline ctDNA was measured by next generation sequencing using a targeted panel (Oncomine Colon cell-free DNA assay) and postoperatively by digital PCR on genetic variants found preoperatively with the Oncomine panel. CTCs were enumerated using the FDA-approved CellSearch system.

FINDINGS: ctDNA was detected in 117/187 patients (63%) at baseline, and 28/104 evaluable patients (27%) still had detectable ctDNA postoperatively. CTC enumeration resulted in positivity for 37/183 patients (20%) at baseline and 14/158 patients (9%) postoperatively. No association was found between 1-year recurrence-free survival (RFS) and the presence of CTCs or ctDNA at baseline. In contrast, patients with postoperative undetectable ctDNA had a significantly improved 1-year RFS compared to patients with postoperative ctDNA (54% [95% CI 44%-67%] vs. 25% [95% CI 13%-47%], log-rank p = 0.0011). Similarly, patients with postoperative detectable CTCs had a significantly shorter 1-year RFS compared to patients without postoperative CTCs (15% [95% CI 4%-55%] vs. 53% [95% CI 45%-62%], log-rank p 0.0004). Also in multivariable analysis, detectable ctDNA and CTCs after surgery remained independently associated with a shorter 1-year RFS (HR 2.35; 95% CI 1.34-4.11; p = 0.0028 and HR 2.98; 95% CI 1.56-5.71; p = 0.0010, respectively).

INTERPRETATION: This is the first study conducted in patients with resectable CRLM without (neo)adjuvant chemotherapy, which demonstrates the impact of postoperative detectable circulating tumour load on 1-year RFS. Postoperative ctDNA and CTC detection both represent strong, independent predictors for a shorter RFS after local treatment, as opposed to preoperative detection.

FUNDING: This work was supported by KWF Kankerbestrijding (Dutch Cancer Society, EMCR 2014-6340).

PMID:40838198 | PMC:PMC12361997 | DOI:10.1016/j.eclinm.2025.103406

Redefining druggable targets with artificial intelligence

Nature Biotechnology, Published online: 19 August 2025; doi:10.1038/s41587-025-02770-1

A vast landscape of ‘undruggable’ cancer targets remains beyond the reach of conventional therapeutic agents. Recent advances in artificial intelligence (AI), however, are challenging this paradigm. Synthesizing insights from a Cancer Moonshot workshop, we argue that systemically addressing the undruggable target space with AI requires a new conceptual framework. We highlight the failure of current target taxonomies and the need for benchmarking datasets, and re-evaluate clinical validation for novel AI-driven modalities.

Thor: a platform for cell-level investigation of spatial transcriptomics and histology

Nat Commun. 2025 Aug 5;16(1):7178. doi: 10.1038/s41467-025-62593-1.

ABSTRACT

Spatial transcriptomics links gene expression with tissue morphology, however, current tools often prioritize genomic analysis, lacking integrated image interpretation. To address this, we present Thor, a comprehensive platform for cell-level analysis of spatial transcriptomics and histological images. Thor employs an anti-shrinking Markov diffusion method to infer single-cell spatial transcriptome from spot-level data, effectively combining gene expression and cell morphology. The platform includes 10 modular tools for genomic and image-based analysis, and is paired with Mjolnir, a web-based interface for interactive exploration of gigapixel images. Thor is validated on simulated data and multiple spatial platforms (ISH, MERFISH, Xenium, Stereo-seq). Thor characterizes regenerative signatures in heart failure, screens breast cancer hallmarks, resolves fine layers in mouse olfactory bulb, and annotates fibrotic heart tissue. In high-resolution Visium HD data, it enhances spatial gene patterns aligned with histology. By bridging transcriptomic and histological analysis, Thor enables holistic tissue interpretation in spatial biology.

PMID:40764306 | PMC:PMC12325965 | DOI:10.1038/s41467-025-62593-1

Whole-genome sequencing of 490,640 UK Biobank participants

Nature, Published online: 06 August 2025; doi:10.1038/s41586-025-09272-9

A study reports whole-genome sequences for 490,640 participants from the UK Biobank and combines these data with phenotypic data to provide new insights into the relationship between human variation and sequence variation.

Personalized molecular signatures of insulin resistance and type 2 diabetes

Muscle samples from over 120 people were analyzed to identify molecular patterns linked to insulin resistance, a key feature of type 2 diabetes. The findings reveal new insights that could help tailor more personalized and effective treatments for the disease.

The generative era of medical AI

10 July 2025 at 08:00
Significant progress has been made in recent years in applying large language models and multimodal artificial intelligence to health and medicine, transforming diagnostics, patient interactions, and medical forecasting, although challenges like privacy, regulation, and system integration remain before widespread clinical adoption.
❌