❌

Normal view

Inositol Metabolism Modulates Inflammatory Injury in Acute Pancreatitis via the ISYNA1-NETs Axis

J Inflamm Res. 2026 Sep 22;19:606503. doi: 10.2147/JIR.S606503. eCollection 2026.

ABSTRACT

BACKGROUND: Neutrophil extracellular traps (NETs) were key factors mediating inflammatory injury in acute pancreatitis (AP). To this end, there was an urgent need to identify precise and effective therapeutic targets that modulate NETs formation, providing new ideas for the prevention and treatment of AP pancreatitis injury.

GAP: To address this gap, we investigated the potential involvement of the myo-inositol metabolism in modulating NETs and inflammatory damage during AP.

METHODS: Multi-omics analysis identified myo-inositol metabolism as critical. We then established the in vitro NETs model using phorbol-12-myristate-13-acetate (PMA) to investigate the role and regulatory mechanism of inositol-3-phosphate synthase 1 (ISYNA1) on NETs formation. Finally, the findings were validated in the classic AP mouse model to verify the correlation between myo-inositol metabolism and AP pathogenesis.

RESULTS: Multiple omics analyses showed that the myo-inositol metabolic pathway is the most significant, and the key enzyme ISYNA1 involved in myo-inositol synthesis was significantly reduced. ISYNA1 was significantly downregulated in both the in vitro NETs model and in neutrophils infiltrating the pancreatic tissue of AP mice. Meanwhile, exogenous supplementation of ISYNA1 or myo-inositol significantly inhibited the NETs formation in vitro and inflammatory injury in AP mice. Mechanistically, downregulation of ISYNA1 led to reduced myo-inositol synthesis, thereby promoting NETs formation via modulation of the PI3K/AKT pathway.

CONCLUSION: ISYNA1 and myo-inositol metabolism were among the key links that regulated NETs formation and inflammatory injury in AP. Therefore, enhancing ISYNA1 and myo-inositol metabolism might serve as a potential intervention target for treating acute organ injury in AP.

PMID:42801157 | PMC:PMC13615823 | DOI:10.2147/JIR.S606503

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

arXiv:2609.13009v1 Announce Type: new Abstract: Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.

SCOPE-OPSD: Fisher-Conditioned Privileged Subspaces for On-Policy Self-Distillation

arXiv:2609.12579v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) scores student-generated prefixes with a solution-conditioned self-teacher, yet transfers supervision only through next-token probabilities. We ask whether the aligned final-layer discrepancy offers a useful second channel, and how to test that channel without confusing its geometry with auxiliary strength. SCOPE-OPSD projects the privileged teacher-student residual onto a frozen rank-64 factor estimated from residual covariance and language-model-head Fisher sensitivity. It reuses the forwards already required by OPSD and adds neither rollouts nor inference-time modules. A matched Random control preserves the structured factor's rank and nonzero spectrum and uses per-arm gradient-RMS calibration, isolating the effect of the data-dependent orientation. Across the complete 25/50/75/100-step trajectories for Qwen3-1.7B, 4B, and 8B, Structured is never below Pure OPSD, with strict gains in 11 of the 12 model-checkpoint combinations and an exact tie at 4B step 25. Structured also exceeds matched Random in 10 of the 12 combinations. At step 75 on Qwen3-1.7B, Structured exceeds matched Random by 1.39 Macro Avg@12 points in each of two independent training reruns. A cross-fitted diagnostic also shows 4.40 times greater held-out privileged-gap capture than the matched random orientation. The results support a compact, Fisher-conditioned privileged subspace for short-budget OPSD.

FEAT: A Linear-Complexity Foundation Model for Extremely Large Structured Data

arXiv:2603.16513v4 Announce Type: replace-cross Abstract: Structured data is widely used in domains such as healthcare, finance, and scientific data management. Recent studies on structured data foundation models (SFMs) aim to support data analysis and mining tasks over such data, but still face scalability and generalization challenges when applied to real-world enterprise databases. First, many SFMs rely on full self-attention, which introduces an O(N^2) computational bottleneck and limits the number of tuples that can be processed jointly. Second, directly replacing attention with linear-complexity sequence models may conflict with the permutation-invariant nature of structured data, introducing artificial order bias and degrading representation quality. Moreover, models trained only on synthetic data may struggle to generalize to the heavy-tailed and heterogeneous distributions commonly found in real-world databases. To address these challenges, we propose FEAT, a linear-complexity foundation model for extremely large structured data. FEAT replaces quadratic attention with a multi-layer dual-axis encoding architecture. It integrates an adaptive-fusion bidirectional state-space model (AFBM) with convolutional gated linear attention (Conv-GLA), enabling cross-tuple contextualization in O(N) time while supporting permutation-invariant representation learning. To improve robustness under real-world data skewness, FEAT further adopts a hybrid structural causal pre-training pipeline with a robust reconstruction objective. Experiments on 12 real-world database benchmarks show that FEAT consistently outperforms representative SFMs on zero-shot tasks and scales linearly with structured-data sample length, achieving up to 50x faster inference latency.

Tumor-derived CTHRC1 mediates ITGB3-dependent osteoclast differentiation to promote prostate cancer bone metastasis

Oncogene, Published online: 11 September 2026; doi:10.1038/s41388-026-03976-6

Tumor-derived CTHRC1 mediates ITGB3-dependent osteoclast differentiation to promote prostate cancer bone metastasis

Impact of LLM-supported patient education on patient perspectives and patient-reported outcomes: a mixed-methods systematic review

npj Digital Medicine, Published online: 10 September 2026; doi:10.1038/s41746-026-03228-7

Impact of LLM-supported patient education on patient perspectives and patient-reported outcomes: a mixed-methods systematic review

Distilling Image Prototypes for Guided Test-Time Adaptation

arXiv:2609.09737v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) enhances the robustness of models against distribution shifts but faces two critical challenges: error accumulation from noisy pseudo-labels and catastrophic forgetting of source knowledge. Uncertainty-based approaches designed to mitigate error accumulation often yield overconfident or computationally expensive estimates, while strategies intended to prevent forgetting via prototype replay rely on static representations that easily become misaligned as the model adapts. To address these issues, this paper proposes a novel framework, Distilling Image Prototype for Guided Test-Time Adaptation (DIPTTA). The core of the proposed approach is the introduction of a Distill Image Prototype (DIP), a compact set of synthetic images that serves as a dynamic and regenerative anchor of source knowledge. This prototype enables a dynamic feature replay mechanism that continuously generates feature prototypes aligned with the current state of the model, thus effectively preventing catastrophic forgetting. Furthermore, the DIP anchors a source-calibrated uncertainty estimation method, which provides a less biased measure of sample reliability by leveraging stable source knowledge, thereby robustly suppressing error accumulation. Extensive experiments on multiple benchmarks demonstrate that DIPTTA significantly outperforms state-of-the-art methods, particularly under severe domain shifts. The source code is available at https://github.com/LiwenWang919/DIPTTA.

FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models

arXiv:2609.09905v1 Announce Type: cross Abstract: Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains unclear. In particular, existing forward-process alignment methods require fresh samples from the current model, while offline methods based on fixed preference pairs rely primarily on positive-only fine-tuning or DPO-style likelihood-ratio surrogates. We organize these approaches through a divergence-based framework and introduce FlowCPO, an offline forward-KL objective that uses both preferred and dispreferred samples without online rollouts. For linear interpolation, we show under explicit regularity conditions that the forward-KL objective is bounded by a contrastive flow matching loss, yielding a tractable surrogate on fixed data. We further show that this loss is nonnegative, whereas the signed regression loss of simplified FlowDPO can be unbounded below. In the in-domain setting, FlowCPO achieves higher mean GenEval and OCR scores than the evaluated baselines, reaching 0.84 and 0.87 versus 0.81 and 0.74 for FlowDPO at CFG 3.0. In the out-of-domain setting, the results are mixed, with the best GenEval result but lower reward scores than RFT on several metrics.

Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization

arXiv:2609.10464v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn physics and generate physically realistic dynamics remains hitherto untested. In this work, we introduce SemiGroup-JEPA (SG-JEPA), which extends the LeWorldModel framework by supplying the parameter governing the physics to the temporal model via action-conditioning and jointly training an encoder and predictor through an autoregressive latent rollout. To evaluate the model's ability to generalize out of distribution, we design dynamical tasks under different gravitational fields that, despite obeying the same physical law, exhibit qualitatively different dynamics, ranging from floating motion in weak gravitational fields to rapid bouncing in strong ones. In contrast to DINO-WM, SG-JEPA reduces open-loop prediction error by up to 2 times on two-dimensional datasets, and increases control success rate up to 2.5 times for three-dimensional robotic datasets, for which we train independent diffusion policies. To explain this advantage, we develop a linear feature model that separates local law-conditioned error from its recursive amplification under rollout. Guided by this model, we find that back-propagating the multi-step rollout loss into the representation trains the encoder to keep the features that the predictor can carry forward, and that those are the features the dynamics depend on, so most of the gain comes from the encoder learning better features rather than from the predictor learning better dynamics. See project page at https://sg-jepa.github.io.

MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation

arXiv:2510.05124v3 Announce Type: replace-cross Abstract: We propose MADS (Multi-Agent Dialogue Simulation), a scalable framework for generating persuasive multi-turn dialogues via agent self-play. MADS employs three coordinated agents: User Agents designed to simulate diverse persona-driven behaviors by leveraging personality signifiers such as Zodiac Signs and MBTI types, a Dialog Agent executing task-oriented persuasion strategies and an Optimization Agent evaluating and refining dialogue outcomes. We further validate its effectiveness through users' Chain-of-Attitude (CoA) modeling and dedicated LLMs' persuasion assessment. This approach enables low-cost generation of training data without human annotation, addressing key industry challenges such as lack of user data, cold-start evaluation difficulties, and prompt inefficiency. Applied to a real-world marketing scenario, MADS significantly improved the persuasion capacity of small LLMs, increasing the organic traffic conversion rate by 22.4% (from 1.83% to 2.24%) , demonstrating clear business value.

Transmembrane glycoprotein BSG serves a dual role as a prognostic and immunological modulator in the tumor microenvironment of lung adenocarcinoma

Transl Oncol. 2026 Sep 8;73:102990. doi: 10.1016/j.tranon.2026.102990. Online ahead of print.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD) is a predominant and lethal subtype of non-small cell lung cancer, with a lack of reliable prognostic biomarkers to guide clinical management. Basigin (BSG) has been implicated in tumor progression across multiple cancers, yet its expression pattern, prognostic significance, and underlying mechanisms in LUAD remain incompletely elucidated.

METHODS: We integrated multi-omics data from TCGA, GTEx, CCLE, and GEO databases to analyze BSG expression profiles. Clinical correlations were assessed via Kruskal-Wallis tests. Prognostic value was determined using Kaplan-Meier survival analysis, univariate/multivariate Cox regression, and nomogram construction with calibration curves. Functional enrichment (GO/KEGG) and immune infiltration analyses were performed to explore BSG-related mechanisms, followed by immunohistochemical (IHC) validation in A549 cells and clinical LUAD tissue microarrays.

RESULTS: BSG was significantly upregulated in LUAD tissues versus normal/paired adjacent tissues, correlating with advanced T/N/pathologic stages. High BSG expression predicted worse survival outcomes in TCGA-LUAD, which was validated in GEO datasets. Multivariate Cox regression identified BSG as an independent prognostic factor, with a well-calibrated nomogram for survival prediction. Functional exploration indicated that BSG mainly participated in tumor-associated and immunological pathways. Immune infiltration analysis indicated that BSG was significantly correlated with the infiltration of various immune cells. Moreover, BSG exhibited a strong association with immune checkpoint proteins, chemokines, chemokine receptors, and MHC genes. IHC further confirmed its cytoplasmic/membranous localization and prognostic relevance.

CONCLUSION: BSG serves as an independent prognostic biomarker and potential therapeutic target in LUAD, shedding light on its regulatory roles in tumor progression and immune microenvironment remodeling.

PMID:42710246 | DOI:10.1016/j.tranon.2026.102990

Transmembrane glycoprotein BSG serves a dual role as a prognostic and immunological modulator in the tumor microenvironment of lung adenocarcinoma

Transl Oncol. 2026 Sep 8;73:102990. doi: 10.1016/j.tranon.2026.102990. Online ahead of print.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD) is a predominant and lethal subtype of non-small cell lung cancer, with a lack of reliable prognostic biomarkers to guide clinical management. Basigin (BSG) has been implicated in tumor progression across multiple cancers, yet its expression pattern, prognostic significance, and underlying mechanisms in LUAD remain incompletely elucidated.

METHODS: We integrated multi-omics data from TCGA, GTEx, CCLE, and GEO databases to analyze BSG expression profiles. Clinical correlations were assessed via Kruskal-Wallis tests. Prognostic value was determined using Kaplan-Meier survival analysis, univariate/multivariate Cox regression, and nomogram construction with calibration curves. Functional enrichment (GO/KEGG) and immune infiltration analyses were performed to explore BSG-related mechanisms, followed by immunohistochemical (IHC) validation in A549 cells and clinical LUAD tissue microarrays.

RESULTS: BSG was significantly upregulated in LUAD tissues versus normal/paired adjacent tissues, correlating with advanced T/N/pathologic stages. High BSG expression predicted worse survival outcomes in TCGA-LUAD, which was validated in GEO datasets. Multivariate Cox regression identified BSG as an independent prognostic factor, with a well-calibrated nomogram for survival prediction. Functional exploration indicated that BSG mainly participated in tumor-associated and immunological pathways. Immune infiltration analysis indicated that BSG was significantly correlated with the infiltration of various immune cells. Moreover, BSG exhibited a strong association with immune checkpoint proteins, chemokines, chemokine receptors, and MHC genes. IHC further confirmed its cytoplasmic/membranous localization and prognostic relevance.

CONCLUSION: BSG serves as an independent prognostic biomarker and potential therapeutic target in LUAD, shedding light on its regulatory roles in tumor progression and immune microenvironment remodeling.

PMID:42710246 | DOI:10.1016/j.tranon.2026.102990

ITGA5 promotes homologous recombination mediated radioresistance in esophageal squamous cell carcinoma by upregulating RAD51AP1 expression

Oncogene, Published online: 04 September 2026; doi:10.1038/s41388-026-03966-8

ITGA5 promotes homologous recombination mediated radioresistance in esophageal squamous cell carcinoma by upregulating RAD51AP1 expression

Genomics and social practices at Mogou and other Gansu sites during prehistoric trans-Eurasian exchange

Ancient DNA from 149 individuals at 11 sites in Gansu, China, dated to around 4,700–3,000 years ago, reveals human population history during early transcontinental exchanges of agriculture and technology, as well as contemporary social practices, at the large Mogou cemetery.

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

arXiv:2605.24636v2 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated questions across three formats (multiple-choice, short-answer, and case-based questions) and assesses three progressive reasoning levels: knowledge recall (L1), routine reasoning (L2), and individualized reasoning (L3). To ensure data quality, the automated construction framework was calibrated by six senior dentists, achieving expert agreement rates of 99.98% for multiple-choice and short-answer questions and 96.78% for the more complex case-based questions. Evaluation of 12 frontier LLMs on GlobalDentBench revealed a sharp, stepwise performance degradation with increasing reasoning complexity. Specifically, accuracy plummeted from 81.34% on multiple-choice to 64.53% on short-answer and 22.34% on case-based questions, while declining markedly from 74.01% at L1 to 55.64% at L2 and 35.71% at L3. More critically, risk analysis of real-world dental cases demonstrated an alarming overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm and risks particularly pronounced in specialties such as orthodontics. These findings expose fundamental limitations in the medical reasoning and safety of current LLMs. Consequently, GlobalDentBench provides a scalable foundation for trustworthy clinical AI evaluation, underscoring the urgent need for rigorous validation before the safe deployment of these models in healthcare.

DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMs

arXiv:2605.25188v1 Announce Type: new Abstract: Multi-agent LLM systems improve reasoning by combining outputs from multiple agents, but interaction-heavy methods can introduce error propagation and high communication overhead. When agents exchange raw responses or reasoning traces, incorrect intermediate reasoning may be adopted and amplified, leading to confident but wrong consensus; multi-round communication also increases token consumption, latency, and inference cost. In this paper, we propose a controlled-communication coordination framework named DarkForest. DarkForest first keeps agents independent, so each agent produces an answer without seeing the others' outputs. It then parses the raw responses into structured candidate records, groups semantically equivalent candidates into clusters, and estimates a calibrated belief distribution over these clusters using agent reliability, confidence, parse quality, support-pattern reliability, and independence corrections. A coordinator receives only policy-permitted evidence from this belief state with controlled communication. Experiments on six reasoning benchmarks show that DarkForest achieves leading overall quality, improves the strongest baseline by up to 30.7\% on benchmark metrics, and reduces token consumption by up to $6.5\times$ compared with communication-heavy baselines.

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

arXiv:2605.25246v2 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines. Existing benchmarks are limited to small or simplified examples far below real-world scale and complexity. We introduce FrontierOR, among the first benchmarks to systematically evaluate LLM-based efficient algorithm design for realistic large-scale optimization problems. FrontierOR includes 180 tasks derived from methodologically diverse papers published in top-tier operations research venues, each with standardized instances and a hidden, expert-verified evaluation suite. We evaluate seven LLMs spanning frontier, cost-effective, and open-source models both in one-shot and test-time evolution settings. The results reveal that frontier models still struggle to move from executable formulations to efficient optimization algorithms: the strongest one-shot model outperforms Gurobi in only 31% of cases in both solution quality and computational efficiency, and even strong coding agents with test-time evolution achieve only 50% on selected hard tasks. FrontierOR establishes a practical evaluation platform for LLM-based optimization algorithm design, which enables future LLMs and agents to be systematically tested on whether they can move beyond correct formulation toward a feasible, high-quality, and efficient algorithm.

Towards end-to-end LLM-based censoring-aware survival analysis

arXiv:2605.25399v1 Announce Type: new Abstract: Objective: Survival analysis is central to medical prediction, yet large language models (LLMs) are rarely used as end-to-end survival models because censoring prevents straightforward supervised fine-tuning. Here we present LLMSurvival, a framework that enables censoring-aware survival analysis with unmodified LLMs operating directly on tabular clinical data. Materials and Methods: LLMSurvival reformulates time-to-event prediction as pairwise ranking among comparable subjects, and derives test-time risk by aggregating comparisons against anchor individuals from the training cohort. Results: Across two clinical tasks (ICU mortality prediction in MIMIC-IV and fragility fracture prediction in a NewYork-Presbyterian/Weill Cornell Medicine cohort), LLMSurvival improves overall concordance over Cox proportional hazards modeling by 3.1% for ICU mortality and 0.5% for fracture risk, 2.1% on average for ICU mortality and 2.8% for fracture risk over three established deep learning survival models. Discussion: The results show that survival modeling with censoring can be made compatible with LLM fine-tuning through comparison-based reformulation. The framework demonstrates high portability and superior performance over expert curated scores like SAPS-II and FRAX scores across diverse clinical context. Furthermore, the framework supports local deployment, as compact, publicly available base models provide sufficient performance. Conclusion: The LLMSurvival framework serves as a proof of concept for an integrated, censoring-conscious approach to survival analysis via LLMs.

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digital workflows. However, real-world execution environments are far from ideal: pop-ups, resolution changes, and competing applications frequently interfere with agent perception and control. We introduce AgentHijack, a benchmark designed to evaluate the robustness of computer-use agents under common corruptions, where the uncertainties in dynamic environment disrupt the execution flow without direct adversarial intent. Specifically, AgentHijack introduces 9 configurable common corruptions to replicate realistic imperfect scenarios. We evaluate a variety of desktop tasks that utilize MLLM-based agents and discover that even minor instances of corruption can result in substantial performance degradation, which emphasizes the fragility of agents and underscores the necessity of robustness evaluation. Afterward, we propose AgentHijack-Agent, a framework that integrates an action generator with enhanced grounding capabilities and an onlooker responsible for behavior summarization and environment checking. Extensive experiments validate its effectiveness. Our code, environment, baseline models and data are publicly available at: https://AgentHijack.github.io.
❌