❌

Reading view

Beyond GeneGPT: A Multi-Agent Architecture with Open-Source LLMs for Enhanced Genomic Question Answering

arXiv:2511.15061v1 Announce Type: new Abstract: Genomic question answering often requires complex reasoning and integration across diverse biomedical sources. GeneGPT addressed this challenge by combining domain-specific APIs with OpenAI's code-davinci-002 large language model to enable natural language interaction with genomic databases. However, its reliance on a proprietary model limits scalability, increases operational costs, and raises concerns about data privacy and generalization. In this work, we revisit and reproduce GeneGPT in a pilot study using open source models, including Llama 3.1, Qwen2.5, and Qwen2.5 Coder, within a monolithic architecture; this allows us to identify the limitations of this approach. Building on this foundation, we then develop OpenBioLLM, a modular multi-agent framework that extends GeneGPT by introducing agent specialization for tool routing, query generation, and response validation. This enables coordinated reasoning and role-based task execution. OpenBioLLM matches or outperforms GeneGPT on over 90% of the benchmark tasks, achieving average scores of 0.849 on Gene-Turing and 0.830 on GeneHop, while using smaller open-source models without additional fine-tuning or tool-specific pretraining. OpenBioLLM's modular multi-agent design reduces latency by 40-50% across benchmark tasks, significantly improving efficiency without compromising model capability. The results of our comprehensive evaluation highlight the potential of open-source multi-agent systems for genomic question answering. Code and resources are available at https://github.com/ielab/OpenBioLLM.
  •  

Exploring the use of AI authors and reviewers at Agents4Science

arXiv:2511.15534v1 Announce Type: new Abstract: There is growing interest in using AI agents for scientific research, yet fundamental questions remain about their capabilities as scientists and reviewers. To explore these questions, we organized Agents4Science, the first conference in which AI agents serve as both primary authors and reviewers, with humans as co-authors and co-reviewers. Here, we discuss the key learnings from the conference and their implications for human-AI collaboration in science.
  •  

Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization

arXiv:2411.05034v2 Announce Type: replace-cross Abstract: Embeddings have become a cornerstone in the functionality of large language models (LLMs) due to their ability to transform text data into rich, dense numerical representations that capture semantic and syntactic properties. These embedding vector databases serve as the long-term memory of LLMs, enabling efficient handling of a wide range of natural language processing tasks. However, the surge in popularity of embedding vector databases in LLMs has been accompanied by significant concerns about privacy leakage. Embedding vector databases are particularly vulnerable to embedding inversion attacks, where adversaries can exploit the embeddings to reverse-engineer and extract sensitive information from the original text data. Existing defense mechanisms have shown limitations, often struggling to balance security with the performance of downstream tasks. To address these challenges, we introduce Eguard, a novel defense mechanism designed to mitigate embedding inversion attacks. Eguard employs a transformer-based projection network and text mutual information optimization to safeguard embeddings while preserving the utility of LLMs. Our approach significantly reduces privacy risks, protecting over 95% of tokens from inversion while maintaining high performance across downstream tasks consistent with original embeddings.
  •  

Accelerating Local AI on Consumer GPUs: A Hardware-Aware Dynamic Strategy for YOLOv10s

arXiv:2509.07928v2 Announce Type: replace-cross Abstract: As local AI grows in popularity, there is a critical gap between the benchmark performance of object detectors and their practical viability on consumer-grade hardware. While models like YOLOv10s promise real-time speeds, these metrics are typically achieved on high-power, desktop-class GPUs. This paper reveals that on resource-constrained systems, such as laptops with RTX 4060 GPUs, performance is not compute-bound but is instead dominated by system-level bottlenecks, as illustrated by a simple bottleneck test. To overcome this hardware-level constraint, we introduce a Two-Pass Adaptive Inference algorithm, a model-independent approach that requires no architectural changes. This study mainly focuses on adaptive inference strategies and undertakes a comparative analysis of architectural early-exit and resolution-adaptive routing, highlighting their respective trade-offs within a unified evaluation framework. The system uses a fast, low-resolution pass and only escalates to a high-resolution model pass when detection confidence is low. On a 5000-image COCO dataset, our method achieves a 1.85x speedup over a PyTorch Early-Exit baseline, with a modest mAP loss of 5.51%. This work provides a practical and reproducible blueprint for deploying high-performance, real-time AI on consumer-grade devices by shifting the focus from pure model optimization to hardware-aware inference strategies that maximize throughput.
  •  

Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts

arXiv:2511.11743v2 Announce Type: replace-cross Abstract: Deploying deep neural networks on resource-constrained devices faces two critical challenges: maintaining accuracy under aggressive quantization while ensuring predictable inference latency. We present a curiosity-driven quantized Mixture-of-Experts framework that addresses both through Bayesian epistemic uncertainty-based routing across heterogeneous experts (BitNet ternary, 1-16 bit BitLinear, post-training quantization). Evaluated on audio classification benchmarks (ESC-50, Quinn, UrbanSound8K), our 4-bit quantization maintains 99.9 percent of 16-bit accuracy (0.858 vs 0.859 F1) with 4x compression and 41 percent energy savings versus 8-bit. Crucially, curiosity-driven routing reduces MoE latency variance by 82 percent (p = 0.008, Levene's test) from 230 ms to 29 ms standard deviation, enabling stable inference for battery-constrained devices. Statistical analysis confirms 4-bit/8-bit achieve practical equivalence with full precision (p > 0.05), while MoE architectures introduce 11 percent latency overhead (p
  •  

STAT+: Armed with AI and virtual care, K Health thinks it can make primary care more accessible

In many parts of the United States, patients have gotten used to living without primary care. Nearly 75 million people in the United States live in an area with a shortage of these critical providers, leading to long wait times — if a patient can find primary care at all. 

The scale of the access problem, “it’s like red alert — red, red, red alert level — and it’s been like that for a while,” said physician Rajesh Patel, vice president of digital patient experience at Mass General Brigham.

The situation is only getting worse: By 2037, the nation will be short 87,000 primary care physicians, according to federal estimates. 

Clinical artificial intelligence company K Health thinks it has part of the solution. Over the last two years, it has partnered with five large health systems — Cedars-Sinai, Mayo Clinic, Hackensack Meridian Health, Hartford HealthCare, and Mass General Brigham — to launch round-the-clock virtual primary care platforms enabled by its AI. Today, it announced another partnership with Northwell Health, New York’s largest health system, which began rolling out its platform in October. 

Continue to STAT+ to read the full story…

© Adobe

  •  

Pan-cancer prevalence, risk, and clinical and demographic characteristics of Lynch Syndrome-associated variants in BioBank Japan

Commun Med (Lond). 2025 Nov 13. doi: 10.1038/s43856-025-01231-9. Online ahead of print.

ABSTRACT

BACKGROUND: Although germline testing for DNA mismatch repair (MMR) genes is routinely performed, clinical guidelines highlight evidence gaps due to limited populations and biases. We examined germline pathogenic variants of MMR genes (MLH1, MSH2, MSH6, and PMS2) in 112,927 unselected individuals from BioBank Japan.

METHODS: We analyzed 74,085 cancer patients with 23 cancer types and 38,842 controls matched by sex, age, and hospital area from BioBank Japan, collected between April 2003 and March 2018. Germline pathogenic variants in the coding regions and 2 bp flanking intronic sequences of MMR genes were identified using a multiplex PCR-based target sequencing method. We examined associations with cancer types and demographic characterization of the pathogenic variants, comparing findings to existing clinical guidelines.

RESULTS: Here we show 228 pathogenic variants identified in MMR genes, with pathogenic MSH6 variants most frequently observed in endometrial cancer and 12 other significant associations. Twelve other significant associations are noted across a broad range of odds ratios, whereas pancreatic cancer exhibits no such association. Pathogenic variant carriers are diagnosed up to 12.4 years earlier than non-carriers, and colorectal and gastric cancers are diagnosed up to 16.4 years later than indicated by the guidelines. Higher carrier frequencies are observed in patients with both colorectal and endometrial cancers (24.8%) and in those with endometrial cancer and a family history of endometrial (26.0%) or colorectal (16.1%) cancers.

CONCLUSIONS: This study provides critical insights for clinical guidelines on the associations between cancer types, age at diagnosis, and carrier frequency.

PMID:41258140 | DOI:10.1038/s43856-025-01231-9

  •  

Latent plasticity of the human pancreas across development, health, and disease

bioRxiv [Preprint]. 2025 Oct 3:2025.10.01.679230. doi: 10.1101/2025.10.01.679230.

ABSTRACT

The pancreas plays a central role in major human diseases, yet our understanding of its cellular diversity and plasticity remains incomplete. Here, we present a single-cell multiomics atlas of the human pancreas, profiling over four million cells and nuclei from 57 donors across fetal development, adult homeostasis, and type 2 diabetes (T2D). Integrating sc/snRNA-seq, snATAC-seq, VASA-seq, spatial transcriptomics (Xenium), and multiplexed proteomics (CODEX), we resolve gene expression, chromatin accessibility, and spatial organization at high resolution. We identify transcriptionally plastic centroacinar-like cells (pCACs) in adults with fetal-like features, delineate endocrine and exocrine lineage trajectories during development, and uncover HNF1A-defined beta cell epigenetic states. In T2D, we observe shifts in beta cell subtypes and altered regulatory programs. Glucose perturbation of healthy islets reveals cell-type-specific adaptation and stress responses. This atlas provides a foundational framework to understand pancreas biology and the role of cellular plasticity in regeneration and disease.

PMID:41256699 | PMC:PMC12622017 | DOI:10.1101/2025.10.01.679230

  •  

SMMILe enables accurate spatial quantification in digital pathology using multiple-instance learning

Nature Cancer, Published online: 19 November 2025; doi:10.1038/s43018-025-01060-8

Gao et al. present SMMILe, a multiple-instance learning-based tool that leverages whole-slide images for accurate spatial quantification without compromising on classification performance, and show it outperforms state-of-the-art methods.
  •  

Using a Technology Acceptance Model to Explore the Intention to Use Digital Health Technologies Among People With Disabilities: Cross-Sectional Survey Study

Background: Digital health technologies, including electronic personal health records (e-PHRs), have emerged globally as pivotal tools for enhancing health management, patient autonomy, and healthcare accessibility. Despite the growing significance, people with disabilities (PWDs) often face considerable barriers to digital health adoption due to accessibility issues, limited technological literacy, and inadequate social support. A comprehensive understanding of the factors influencing technology acceptance among PWDs is essential for effectively integrating these digital solutions into healthcare practices. Objective: The objective of this study was to identify the factors influencing the intention to use digital health technologies among PWDs based on the Technology Acceptance Model (TAM). Specifically, this study aimed to analyze the structural relationships among perceived ease of use, perceived usefulness, usage intention, and external factors, including health consciousness, health information consent, content characteristics, information security, eHealth literacy, and effectiveness. Methods: A nationwide survey was conducted in South Korea, reaching a total of 800 PWDs who utilized rehabilitation hospitals, disability welfare centers, public health centers, and other institutions related to disability. Participants were recruited using proportionate stratified sampling and systematic stratified cluster sampling methods. The survey questionnaire was designed based on validated instruments from previous studies and modified to suit the context of e-PHR adoption among PWDs. Data were collected from August 30 to September 30, 2023. Statistical analyses included descriptive statistics, correlation analysis, exploratory factor analysis, and structural equation modeling (SEM) to test hypotheses and examine causal relationships among variables. Results: Perceived usefulness and perceived ease of use significantly influenced participants’ intentions to adopt e-PHR services. The structural model revealed significant paths from external factors to perceived ease of use, including health consciousness, content characteristics, health information consent, information security, and effectiveness. Conversely, eHealth literacy did not significantly impact perceived ease of use. For perceived usefulness, significant influencing factors included content characteristics, health information consent, eHealth literacy, and effectiveness, while health consciousness and information security showed no significant effects. The model demonstrated good overall fit. Conclusions: This study highlights that perceived usefulness and ease of use are crucial mediators driving the adoption of e-PHR among PWDs. Social support, content quality, and consent regarding health information significantly shape technology acceptance. These findings suggest the need for designing user-friendly digital health solutions that integrate robust support systems, address privacy concerns, and deliver high-quality, relevant content tailored to this population. Future research should incorporate longitudinal and mixed-methods studies to further explore the sustained adoption of digital health technologies and validate these results across diverse disability groups and contexts.
  •  

Lightweight LLM powers Japanese enterprise AI deployments

Enterprise AI deployment faces a fundamental tension: organisations need sophisticated language models but baulk at the infrastructure costs and energy consumption of frontier systems.

NTT’s recent launch of tsuzumi 2, a lightweight large language model (LLM) running on a single GPU, demonstrates how businesses are resolving this constraint – with early deployments showing performance matching larger models and running at a fraction of the operational cost.

The business case is straightforward. Traditional large language models require dozens or hundreds of GPUs, creating electricity consumption and operational cost barriers that make AI deployment impractical for many organisations.

(GPU Cost Comparison)

For enterprises operating in markets with constrained power infrastructure or tight operational budgets, these requirements eliminate AI as a viable option. NTT’s press release illustrates the practical considerations driving lightweight LLM adoption with Tokyo Online University’s deployment.

The university operates an on-premise platform keeping student and staff data in its campus network – a data sovereignty requirement common in educational institutions and regulated industries.

After validating that tsuzumi 2 handles complex context understanding and long-document processing at production-ready levels, the university deployed it for course Q&A enhancement, teaching material creation support, and personalised student guidance.

The single-GPU operation means the university avoids both capital expenditure for GPU clusters and ongoing electricity costs. More significantly, on-premise deployment addresses data privacy concerns that prevent many educational institutions from using cloud-based AI services that process sensitive student information.

Performance without scale: The technical economics

NTT’s internal evaluation for financial-system inquiry handling showed tsuzumi 2 matching or exceeding leading external models despite dramatically smaller infrastructure requirements. The performance-to-resource ratio determines AI adoption feasibility for enterprises where the total cost of ownership drives decisions.

The model delivers what NTT characterises as “world-top results among models of comparable size” in Japanese language performance, with particular strength in business domains prioritising knowledge, analysis, instruction-following, and safety.

For enterprises operating primarily in Japanese markets, this language optimisation reduces the need to deploy larger multilingual models requiring significantly more computational resources.

Reinforced knowledge in financial, medical, and public sectors – developed based on customer demand – enables domain-specific deployments without extensive fine-tuning.

The model’s RAG (Retrieval-Augmented Generation) and fine-tuning capabilities allow efficient development of specialised applications for enterprises with proprietary knowledge bases or industry-specific terminology where generic models underperform.

Data sovereignty and security as business drivers

Beyond cost considerations, data sovereignty drives lightweight LLM adoption in regulated industries. Organisations handling confidential information face risk exposure when processing data through external AI services subject to foreign jurisdiction.

NTT positions tsuzumi 2 as a “purely domestic model” developed from scratch in Japan, operating on-premises or in private clouds. This addresses concerns prevalent in Asia-Pacific markets about data residency, regulatory compliance, and information security.

FUJIFILM Business Innovation’s partnership with NTT DOCOMO BUSINESS demonstrates how enterprises combine lightweight models with existing data infrastructure. FUJIFILM’s REiLI technology converts unstructured corporate data – contracts, proposals, mixed text and images – into structured information.

Integrating tsuzumi 2’s generative capabilities enables advanced document analysis without transmitting sensitive corporate information to external AI providers. This architectural approach – combining lightweight models with on-premise data processing – represents a practical enterprise AI strategy balancing capability requirements with security, compliance, and cost constraints.

Multimodal capabilities and enterprise workflows

tsuzumi 2 includes built-in multimodal support handling text, images, and voice in enterprise applications. Thematters for business workflows requiring AI to process multiple data types without deploying separate specialised models.

Manufacturing quality control, customer service operations, and document processing workflows typically involve text, images, and sometimes voice inputs. Single models handling all three reduce integration complexity compared to managing multiple specialised systems with different operational requirements.

Market context and implementation considerations

NTT’s lightweight approach contrasts with hyperscaler strategies emphasising massive models with broad capabilities. For enterprises with substantial AI budgets and advanced technical teams, frontier models from OpenAI, Anthropic, and Google provide cutting-edge performance.

However, this approach excludes organisations lacking these resources – a significant portion of the enterprise market, particularly in Asia-Pacific regions with varying infrastructure quality. Regional considerations matter.

Power reliability, internet connectivity, data centre availability, and regulatory frameworks vary significantly in markets. Lightweight models enabling on-premise deployment accommodate these variations better than approaches requiring consistent cloud infrastructure access.

Organisations evaluating lightweight LLM deployment should consider several factors:

Domain specialisation: tsuzumi 2’s reinforced knowledge in financial, medical, and public sectors addresses specific domains, but organisations in other industries should evaluate whether available domain knowledge meets their requirements.

Language considerations: Optimisation for Japanese language processing benefits Japanese-market operations but may not suit multilingual enterprises requiring consistent cross-language performance.

Integration complexity: On-premise deployment requires internal technical capabilities for installation, maintenance, and updates. Organisations lacking these capabilities may find cloud-based alternatives operationally simpler despite higher costs.

Performance tradeoffs: While tsuzumi 2 matches larger models in specific domains, frontier models may outperform in edge cases or novel applications. Organisations should evaluate whether domain-specific performance suffices or whether broader capabilities justify higher infrastructure costs.

The practical path forward?

NTT’s tsuzumi 2 deployment demonstrates that sophisticated AI implementation doesn’t require hyperscale infrastructure – at least for organisations whose requirements align with lightweight model capabilities. Early enterprise adoptions show practical business value: reduced operational costs, improved data sovereignty, and production-ready performance for specific domains.

As enterprises navigate AI adoption, the tension between capability requirements and operational constraints increasingly drives demand for efficient, specialised solutions rather than general-purpose systems requiring extensive infrastructure.

For organisations evaluating AI deployment strategies, the question isn’t whether lightweight models are “better” than frontier systems – it’s whether they’re sufficient for specific business requirements while addressing cost, security, and operational constraints that make alternative approaches impractical.

The answer, as Tokyo Online University and FUJIFILM Business Innovation deployments demonstrate, is increasingly yes.

See also: How Levi Strauss is using AI for its DTC-first business model

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Lightweight LLM powers Japanese enterprise AI deployments appeared first on AI News.

  •  

PathMind: A Retrieve-Prioritize-Reason Framework for Knowledge Graph Reasoning with Large Language Models

arXiv:2511.14256v1 Announce Type: new Abstract: Knowledge graph reasoning (KGR) is the task of inferring new knowledge by performing logical deductions on knowledge graphs. Recently, large language models (LLMs) have demonstrated remarkable performance in complex reasoning tasks. Despite promising success, current LLM-based KGR methods still face two critical limitations. First, existing methods often extract reasoning paths indiscriminately, without assessing their different importance, which may introduce irrelevant noise that misleads LLMs. Second, while many methods leverage LLMs to dynamically explore potential reasoning paths, they require high retrieval demands and frequent LLM calls. To address these limitations, we propose PathMind, a novel framework designed to enhance faithful and interpretable reasoning by selectively guiding LLMs with important reasoning paths. Specifically, PathMind follows a "Retrieve-Prioritize-Reason" paradigm. First, it retrieves a query subgraph from KG through the retrieval module. Next, it introduces a path prioritization mechanism that identifies important reasoning paths using a semantic-aware path priority function, which simultaneously considers the accumulative cost and the estimated future cost for reaching the target. Finally, PathMind generates accurate and logically consistent responses via a dual-phase training strategy, including task-specific instruction tuning and path-wise preference alignment. Extensive experiments on benchmark datasets demonstrate that PathMind consistently outperforms competitive baselines, particularly on complex reasoning tasks with fewer input tokens, by identifying essential reasoning paths.
  •  

Synthetic Clinical Notes for Rare ICD Codes: A Data-Centric Framework for Long-Tail Medical Coding

arXiv:2511.14112v1 Announce Type: cross Abstract: Automatic ICD coding from clinical text is a critical task in medical NLP but remains hindered by the extreme long-tail distribution of diagnostic codes. Thousands of rare and zero-shot ICD codes are severely underrepresented in datasets like MIMIC-III, leading to low macro-F1 scores. In this work, we propose a data-centric framework that generates high-quality synthetic discharge summaries to mitigate this imbalance. Our method constructs realistic multi-label code sets anchored on rare codes by leveraging real-world co-occurrence patterns, ICD descriptions, synonyms, taxonomy, and similar clinical notes. Using these structured prompts, we generate 90,000 synthetic notes covering 7,902 ICD codes, significantly expanding the training distribution. We fine-tune two state-of-the-art transformer-based models, PLM-ICD and GKI-ICD, on both the original and extended datasets. Experiments show that our approach modestly improves macro-F1 while maintaining strong micro-F1, outperforming prior SOTA. While the gain may seem marginal relative to the computational cost, our results demonstrate that carefully crafted synthetic data can enhance equity in long-tail ICD code prediction.
  •  

Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning

arXiv:2511.14445v1 Announce Type: cross Abstract: We present Tell Me, a mental well-being system that leverages advances in large language models to provide accessible, context-aware support for users and researchers. The system integrates three components: (i) a retrieval-augmented generation (RAG) assistant for personalized, knowledge-grounded dialogue; (ii) a synthetic client-therapist dialogue generator conditioned on client profiles to facilitate research on therapeutic language and data augmentation; and (iii) a Well-being AI crew, implemented with CrewAI, that produces weekly self-care plans and guided meditation audio. The system is designed as a reflective space for emotional processing rather than a substitute for professional therapy. It illustrates how conversational assistants can lower barriers to support, complement existing care, and broaden access to mental health resources. To address the shortage of confidential therapeutic data, we introduce synthetic client-therapist dialogue generation conditioned on client profiles. Finally, the planner demonstrates an innovative agentic workflow for dynamically adaptive, personalized self-care, bridging the limitations of static well-being tools. We describe the architecture, demonstrate its functionalities, and report evaluation of the RAG assistant in curated well-being scenarios using both automatic LLM-based judgments and a human-user study. This work highlights opportunities for interdisciplinary collaboration between NLP researchers and mental health professionals to advance responsible innovation in human-AI interaction for well-being.
  •  
❌