❌

Reading view

Making LLMs Reliable When It Matters Most: A Five-Layer Architecture for High-Stakes Decisions

arXiv:2511.07669v1 Announce Type: new Abstract: Current large language models (LLMs) excel in verifiable domains where outputs can be checked before action but prove less reliable for high-stakes strategic decisions with uncertain outcomes. This gap, driven by mutually reinforcing cognitive biases in both humans and artificial intelligence (AI) systems, threatens the defensibility of valuations and sustainability of investments in the sector. This report describes a framework emerging from systematic qualitative assessment across 7 frontier-grade LLMs and 3 market-facing venture vignettes under time pressure. Detailed prompting specifying decision partnership and explicitly instructing avoidance of sycophancy, confabulation, solution drift, and nihilism achieved initial partnership state but failed to maintain it under operational pressure. Sustaining protective partnership state required an emergent 7-stage calibration sequence, built upon a 4-stage initialization process, within a 5-layer protection architecture enabling bias self-monitoring, human-AI adversarial challenge, partnership state verification, performance degradation detection, and stakeholder protection. Three discoveries resulted: partnership state is achievable through ordered calibration but requires emergent maintenance protocols; reliability degrades when architectural drift and context exhaustion align; and dissolution discipline prevents costly pursuit of fundamentally wrong directions. Cross-model validation revealed systematic performance differences across LLM architectures. This approach demonstrates that human-AI teams can achieve cognitive partnership capable of preventing avoidable regret in high-stakes decisions, addressing return-on-investment expectations that depend on AI systems supporting consequential decision-making without introducing preventable cognitive traps when verification arrives too late.
  •  

Toward Practical BCI: A Real-time Wireless Imagined Speech EEG Decoding System

arXiv:2511.07936v1 Announce Type: new Abstract: Brain-computer interface (BCI) research, while promising, has largely been confined to static and fixed environments, limiting real-world applicability. To move towards practical BCI, we introduce a real-time wireless imagined speech electroencephalogram (EEG) decoding system designed for flexibility and everyday use. Our framework focuses on practicality, demonstrating extensibility beyond wired EEG devices to portable, wireless hardware. A user identification module recognizes the operator and provides a personalized, user-specific service. To achieve seamless, real-time operation, we utilize the lab streaming layer to manage the continuous streaming of live EEG signals to the personalized decoder. This end-to-end pipeline enables a functional real-time application capable of classifying user commands from imagined speech EEG signals, achieving an overall 4-class accuracy of 62.00 % on a wired device and 46.67 % on a portable wireless headset. This paper demonstrates a significant step towards truly practical and accessible BCI technology, establishing a clear direction for future research in robust, practical, and personalized neural interfaces.
  •  

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v1 Announce Type: new Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.
  •  

Targeted inhibition of gastric adenocarcinoma by nano-curcumin liposomes: Insights from combined machine learning and experimental analyses into the mechanisms of cuproptosis and metabolic reprogramming

Int J Pharm. 2025 Nov 9:126368. doi: 10.1016/j.ijpharm.2025.126368. Online ahead of print.

ABSTRACT

PURPOSE: Gastric adenocarcinoma is a highly aggressive malignancy characterized by a complex tumor microenvironment. Nano-curcumin liposomes hold great potential in inhibiting tumor growth and survival, as well as inducing cuproptosis and oxidative stress. Although the anticancer properties of curcumin have been demonstrated, the specific mechanisms by which curcumin inhibites gastric adenocarcinoma through cuproptosis remains unclear. This study investigated how nano-curcumin liposomes mediated the inhibition of gastric adenocarcinoma cell proliferation and survival via cuproptosis.

METHODS: This study utilized the gastric adenocarcinoma cell line AGS to establish 2D and 3D in vitro gastric adenocarcinoma models. Furthermore, we prepared nano-curcumin liposomes to investigate their effects and regulatory mechanisms on AGS gastric adenocarcinoma models. A series of in vitro assays, including flow cytometry, CCK-8, scratch assays and morphological assessments, were performed to evaluate the effects of nano-curcumin liposomes on cell apoptosis, proliferation and migration. Additionally, bioinformatics and machine learning methods were employed to identify key targets that inhibited gastric adenocarcinoma growth and survival associated with nano-curcumin liposomes, which were further validated through RT-qPCR and omics analysis. Computer simulations were also conducted to assess the stability of binding interactions between curcumin and key target proteins.

RESULTS: Cellular experiments demonstrated that nano-curcumin liposomes significantly inhibited proliferation and invasive capacity of gastric adenocarcinoma cells while promoting cellular oxidative stress. Bioinformatics and machine learning analyses identified FDX1, GPX4, SERPINE1 and SLC27A5 as key targets. RT-qPCR results confirmed that nano-curcumin liposomes significantly downregulated the expression of these targets. Molecular dynamics simulations indicated that curcumin could form stable binding interactions with key protein targets.

CONCLUSION: This study revealed that nano-curcumin liposomes inhibited growth and survival of gastric adenocarcinoma cells by interfering with the expression of FDX1, GPX4, SERPINE1 and SLC27A5, which were closely linked to copper-induced oxidative stress. Nano-curcumin liposomes downregulated the expression of FDX1 and GPX4, disrupted mitochondrial energy metabolism, and induced oxidative stress, thereby promoting tumor-associated programmed cell death linked to cuproptosis. Furthermore, by downregulating SERPINE1, nano-curcumin liposomes modulated cell adhesion and migration, inhibiting the invasive and metastatic potential of tumor cells. Finally, downregulation of SLC27A5 altered tumor metabolism and cellular homeostasis, induced oxidative stress, and disrupted intracellular environmental stability, thereby suppressing the growth of gastric adenocarcinoma.

PMID:41218732 | DOI:10.1016/j.ijpharm.2025.126368

  •  

HybridGuard: Enhancing Minority-Class Intrusion Detection in Dew-Enabled Edge-of-Things Networks

arXiv:2511.07793v1 Announce Type: cross Abstract: Securing Dew-Enabled Edge-of-Things (EoT) networks against sophisticated intrusions is a critical challenge. This paper presents HybridGuard, a framework that integrates machine learning and deep learning to improve intrusion detection. HybridGuard addresses data imbalance through mutual information based feature selection, ensuring that the most relevant features are used to improve detection performance, especially for minority attack classes. The framework leverages Wasserstein Conditional Generative Adversarial Networks with Gradient Penalty (WCGAN-GP) to further reduce class imbalance and enhance detection precision. It adopts a two-phase architecture called DualNetShield to support advanced traffic analysis and anomaly detection, improving the granular identification of threats in complex EoT environments. HybridGuard is evaluated on the UNSW-NB15, CIC-IDS-2017, and IOTID20 datasets, where it demonstrates strong performance across diverse attack scenarios and outperforms existing solutions in adapting to evolving cybersecurity threats. This approach establishes HybridGuard as an effective tool for protecting EoT networks against modern intrusions.
  •  

Reliable and Private Utility Signaling for Data Markets

arXiv:2511.07975v1 Announce Type: cross Abstract: The explosive growth of data has highlighted its critical role in driving economic growth through data marketplaces, which enable extensive data sharing and access to high-quality datasets. To support effective trading, signaling mechanisms provide participants with information about data products before transactions, enabling informed decisions and facilitating trading. However, due to the inherent free-duplication nature of data, commonly practiced signaling methods face a dilemma between privacy and reliability, undermining the effectiveness of signals in guiding decision-making. To address this, this paper explores the benefits and develops a non-TCP-based construction for a desirable signaling mechanism that simultaneously ensures privacy and reliability. We begin by formally defining the desirable utility signaling mechanism and proving its ability to prevent suboptimal decisions for both participants and facilitate informed data trading. To design a protocol to realize its functionality, we propose leveraging maliciously secure multi-party computation (MPC) to ensure the privacy and robustness of signal computation and introduce an MPC-based hash verification scheme to ensure input reliability. In multi-seller scenarios requiring fair data valuation, we further explore the design and optimization of the MPC-based KNN-Shapley method with improved efficiency. Rigorous experiments demonstrate the efficiency and practicality of our approach.
  •  

Self-Correction Distillation for Structured Data Question Answering

arXiv:2511.07998v1 Announce Type: cross Abstract: Structured data question answering (QA), including table QA, Knowledge Graph (KG) QA, and temporal KG QA, is a pivotal research area. Advances in large language models (LLMs) have driven significant progress in unified structural QA frameworks like TrustUQA. However, these frameworks face challenges when applied to small-scale LLMs since small-scale LLMs are prone to errors in generating structured queries. To improve the structured data QA ability of small-scale LLMs, we propose a self-correction distillation (SCD) method. In SCD, an error prompt mechanism (EPM) is designed to detect errors and provide customized error messages during inference, and a two-stage distillation strategy is designed to transfer large-scale LLMs' query-generation and error-correction capabilities to small-scale LLM. Experiments across 5 benchmarks with 3 structured data types demonstrate that our SCD achieves the best performance and superior generalization on small-scale LLM (8B) compared to other distillation methods, and closely approaches the performance of GPT4 on some datasets. Furthermore, large-scale LLMs equipped with EPM surpass the state-of-the-art results on most datasets.
  •  

Clinical Uncertainty Impacts Machine Learning Evaluations

arXiv:2509.22242v2 Announce Type: replace Abstract: Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on medical imaging benchmarks, accounting for the confidence in binary labels significantly impacts model rankings. We therefore argue that machine-learning evaluations should explicitly account for annotation uncertainty using probabilistic metrics that directly operate on distributions. These metrics can be applied independently of the annotations' generating process, whether modeled by simple counting, subjective confidence ratings, or probabilistic response models. They are also computationally lightweight, as closed-form expressions have linear-time implementations once examples are sorted by model score. We thus urge the community to release raw annotations for datasets and to adopt uncertainty-aware evaluation so that performance estimates may better reflect clinical data.
  •  

SCoTT: Strategic Chain-of-Thought Tasking for Wireless-Aware Robot Navigation in Digital Twins

arXiv:2411.18212v3 Announce Type: replace-cross Abstract: Path planning under wireless performance constraints is a complex challenge in robot navigation. However, naively incorporating such constraints into classical planning algorithms often incurs prohibitive search costs. In this paper, we propose SCoTT, a wireless-aware path planning framework that leverages vision-language models (VLMs) to co-optimize average path gains and trajectory length using wireless heatmap images and ray-tracing data from a digital twin (DT). At the core of our framework is Strategic Chain-of-Thought Tasking (SCoTT), a novel prompting paradigm that decomposes the exhaustive search problem into structured subtasks, each solved via chain-of-thought prompting. To establish strong baselines, we compare classical A* and wireless-aware extensions of it, and derive DP-WA*, an optimal, iterative dynamic programming algorithm that incorporates all path gains and distance metrics from the DT, but at significant computational cost. In extensive experiments, we show that SCoTT achieves path gains within 2% of DP-WA* while consistently generating shorter trajectories. Moreover, SCoTT's intermediate outputs can be used to accelerate DP-WA* by reducing its search space, saving up to 62% in execution time. We validate our framework using four VLMs, demonstrating effectiveness across both large and small models, thus making it applicable to a wide range of compact models at low inference cost. We also show the practical viability of our approach by deploying SCoTT as a ROS node within Gazebo simulations. Finally, we discuss data acquisition pipelines, compute requirements, and deployment considerations for VLMs in 6G-enabled DTs, underscoring the potential of natural language interfaces for wireless-aware navigation in real-world applications.
  •  

Explaining the Unexplainable: A Systematic Review of Explainable AI in Finance

arXiv:2503.05966v3 Announce Type: replace-cross Abstract: Practitioners and researchers trying to strike a balance between accuracy and transparency center Explainable Artificial Intelligence (XAI) at the junction of finance. This paper offers a thorough overview of the changing scene of XAI applications in finance together with domain-specific implementations, methodological developments, and trend mapping of research. Using bibliometric and content analysis, we find topic clusters, significant research, and most often used explainability strategies used in financial industries. Our results show a substantial dependence on post-hoc interpretability techniques; attention mechanisms, feature importance analysis and SHAP are the most often used techniques among them. This review stresses the need of multidisciplinary approaches combining financial knowledge with improved explainability paradigms and exposes important shortcomings in present XAI systems.
  •  

UniSite: The First Cross-Structure Dataset and Learning Framework for End-to-End Ligand Binding Site Detection

arXiv:2506.03237v3 Announce Type: replace-cross Abstract: The detection of ligand binding sites for proteins is a fundamental step in Structure-Based Drug Design. Despite notable advances in recent years, existing methods, datasets, and evaluation metrics are confronted with several key challenges: (1) current datasets and methods are centered on individual protein-ligand complexes and neglect that diverse binding sites may exist across multiple complexes of the same protein, introducing significant statistical bias; (2) ligand binding site detection is typically modeled as a discontinuous workflow, employing binary segmentation and subsequent clustering algorithms; (3) traditional evaluation metrics do not adequately reflect the actual performance of different binding site prediction methods. To address these issues, we first introduce UniSite-DS, the first UniProt (Unique Protein)-centric ligand binding site dataset, which contains 4.81 times more multi-site data and 2.08 times more overall data compared to the previously most widely used datasets. We then propose UniSite, the first end-to-end ligand binding site detection framework supervised by set prediction loss with bijective matching. In addition, we introduce Average Precision based on Intersection over Union (IoU) as a more accurate evaluation metric for ligand binding site prediction. Extensive experiments on UniSite-DS and several representative benchmark datasets demonstrate that IoU-based Average Precision provides a more accurate reflection of prediction quality, and that UniSite outperforms current state-of-the-art methods in ligand binding site detection. The dataset and codes will be made publicly available at https://github.com/quanlin-wu/unisite.
  •  

Digital Health Technologies for Screening and Identifying Unmet Social Needs: Scoping Review

Background: Social determinants of health (SDOH) strongly influence clinical outcomes. Social needs are the individual-level, actionable facets of the broader SDOH framework, including food security, stable housing, and access to essential services. When these needs go unmet, they adversely affect wellbeing and quality of care. Systematically detecting social needs is therefore critical, and emerging digital tools now offer efficient, scalable approaches for screening and identification. Objective: This scoping review aims to examine digital health technology (DHT) use or interventions documented for screening and identifying unmet social needs within high-need populations. We explore trends, effects, challenges, and limitations of identified technologies. Methods: Following PRISMA-ScR guidelines, we searched databases including MEDLINE, Embase, Scopus, ACM Digital Library, and Web of Science for studies published from 2010 to 2025. Eligible studies used technology to screen for and identify unmet social needs in populations with health and socioeconomic challenges. Data extraction focused on the types of technology, screening processes, and social needs identified. Results: Our findings highlight a limited yet evolving landscape of technological applications. We identified 14 studies using tools like self-assessment surveys, tablet-based systems, and electronic portals. These tools were applied across diverse groups, such as refugees and patients in emergency departments. Innovative approaches, such as chatbots and multi-dimensional risk appraisal systems for older adults, showed potential. However, challenges included single-site studies, small samples, and integration issues with medical records. The effectiveness of these tools in screening for unmet social needs shows mixed outcomes. Conclusions: DHTs play a pivotal role in improving the identification of unmet social needs. The findings underscore the need for broader, more integrated research to fully understand the impact of technology-based assessments and screening processes for social needs. Future efforts should focus on facilitated screening using technology both within and outside of the visit, ensuring the linkage to appropriate resources and care.
  •  

STAT+: Chinese government’sΒ support for biotech fuels huge rally

Want to stay on top of the science and politics driving biotech today?Β Sign upΒ to get our biotech newsletter in your inbox.

Good morning, we just had our first snow of the season in Chicago, I just ordered a pie for Thanksgiving, and I’m still in denial that the year is almost ending.

Onto the news today.

Continue to STAT+ to read the full story…

Β© PHILIPPE LOPEZ/AFP/Getty Images

  •  
❌