❌

Normal view

SPIRIT-CONSORT-ELM: element-level annotated dataset and large language model approach for assessing randomized controlled trial reporting

npj Digital Medicine, Published online: 06 October 2026; doi:10.1038/s41746-026-03318-6

SPIRIT-CONSORT-ELM: element-level annotated dataset and large language model approach for assessing randomized controlled trial reporting

Levetiracetam therapeutically targets GABAergic synapses in diffuse midline glioma

Nature Medicine, Published online: 17 September 2026; doi:10.1038/s41591-026-04646-6

Results of this study show in experimental models and data from patient cohorts that the antiseizure medication levetiracetam is associated with longer survival and reduced tumor growth in diffuse midline glioma, but not hemispheric high-grade glioma, by selectively dampening GABAergic synaptic signaling, independently of its canonical SV2A-mediated primary antiseizure mechanism.

Liquid biopsy for early detection of pancreatic ductal adenocarcinoma

Nature Medicine, Published online: 16 September 2026; doi:10.1038/s41591-026-04625-x

In a prospective study involving 1,785 individuals from four countries, the PANXEON exosome-based biomarker, combined with carbohydrate antigen 19-9 levels, achieves high sensitivity for the detection of early-stage pancreatic cancer.

Fifteen challenges for generative AI applications to cell biology

Drawing inspiration from Hilbert’s list of 23 mathematical problems that have focused the mathematical community’s attention for more than a century, we propose fifteen grand AI challenges to focus the biomedical community’s attention on critically relevant questions, most of which still lack effective predictive methodologies.

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

arXiv:2609.13009v1 Announce Type: new Abstract: Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.

Recommendation Retrievers Need Verifiers: Universal Generative Reranking for Sequential Recommendations

arXiv:2609.12270v1 Announce Type: cross Abstract: First-stage recommenders in multi-stage systems produce a ranked candidate list from which a limited prefix is forwarded to downstream rankers. Because each forwarded item must be processed by more expensive ranking stages, this shortlist cannot be arbitrarily large. The first-stage objective is therefore high coverage of relevant items within the forwarded prefix, commonly measured by Recall@$k$. A relevant item may be available deeper in the retrieved list but absent from the shorter prefix that is actually consumed. This paper studies post-hoc verification for promoting such candidates into the consumed shortlist without retraining or replacing the retriever. We introduce a lightweight generative verifier for retrieval models. Given a retriever state and a candidate item, the verifier scores the item through the likelihood of its identifier tokens. It is trained post hoc with next-token cross entropy, requires no sampled negatives or candidate pool during training, and scores only the retriever's top-$K$ candidates at inference. The interface is minimal: the retriever supplies a query state and candidate items, and the item representation can use any fixed tokenization. Across Amazon product recommendation and YaMBDa music recommendation, the same verifier training recipe improves Recall@10 for SASRec, GRU4Rec, NextItNet, and MiniOneRec. Ablations show that the improvements are not explained solely by injecting item-content features into the retriever, supporting verification as a post-hoc output-side adaptation mechanism.

Safe Learning Under Irreversible Dynamics via Asking for Help

arXiv:2502.14043v4 Announce Type: replace-cross Abstract: Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot be recovered from. Instead, we allow the learning agent to ask for help from a mentor and to transfer knowledge between similar states. We show that this combination enables the agent to learn both safely and effectively. Under standard online learning assumptions, we provide an algorithm whose regret and number of mentor queries are both sublinear in the time horizon for Markov decision processes with irreversible dynamics and infinite state spaces. Our proof involves a sequence of three reductions, making our result more general than a single algorithm. Conceptually, our result may be the first formal proof that it is possible for an agent to obtain high reward while becoming self-sufficient in an unknown, unbounded, and high-stakes environment without resets.

Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC

Nature Medicine, Published online: 13 September 2026; doi:10.1038/s41591-026-04488-2

In a large international real-world study of non-small cell lung cancer, a multimodal explainable AI model outperformed established biomarkers for immunotherapy outcome prediction and improved physician decision-making.

Teclistamab versus lenalidomide-dexamethasone in high-risk smoldering multiple myeloma: a randomized phase 2 trial

Nature Medicine, Published online: 11 September 2026; doi:10.1038/s41591-026-04642-w

In the randomized phase 2 ImmunoPRISM trial, patients with high-risk smoldering multiple myeloma (MM) showed higher rates of complete clinical responses in response to treatment with teclistamab compared with lenalidomide–dexamethasone, although longer follow-up is required to determine durable prevention of progression to MM.

Integrated in vitro transcription and oligo-dT affinity chromatography enable multi-cycle reagent recycling for mRNA manufacturing

Kis and colleagues report an integrated sequential-batch IVT–oligo-dT process that links RNA synthesis and affinity capture through a shared buffer, enabling direct crude-IVT loading and flowthrough recycling. The workflow improves cap-analog utilization and raw-material efficiency while preserving functional mRNA expression across five cycles.

Targeting of the oncogenic fusion EWSR1-FLI1 in Ewing Sarcoma by CRISPR/dCas9 silencers

Blancafort and colleagues describe a non-viral polymeric system for the delivery of dCas9-KRAB silencers as ribonucleoprotein (RNP) payloads for EWSR1-FLI1 repression. They demonstrate highly efficient RNP delivery and robust silencing of EWSR1-FLI1 in both cell line and patient-derived xenografts of Ewing sarcoma, accompanied by potent anti-tumor effects.

ACC1 inhibition enhances BCG-induced trained immunity by reprogramming acetyl-CoA metabolism

The efficacy of vaccines remains suboptimal in many settings, underscoring the need for new strategies. Baydemir and colleagues show that modulation of acetyl-CoA metabolism reshapes metabolic and epigenetic programs underlying Bacille Calmette-Guérin-induced trained immunity, enhancing cellular innate immune responses and identifying immunometabolic targeting as a promising approach to improve vaccine efficacy.

Quantum-well metasurface for free-space-accessible enhanced nonlinear polarization

Nature Nanotechnology, Published online: 02 September 2026; doi:10.1038/s41565-026-02268-0

Interband-engineered quantum wells integrated with a resonant metasurface yield giant near-infrared-to-visible optical nonlinearities.

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

arXiv:2609.09240v1 Announce Type: cross Abstract: Ultra-low-bit language models promise reductions in storage and memory traffic, but a nominal "1.58-bit" label does not specify the deployed representation or its execution cost. We study a scale-up of an aggressive post-training conversion pipeline from Qwen3-4B to Qwen3-8B. The conversion uses KOTMS rotation, E2M-ATQ adaptive ternarisation, and GPTQ-style error compensation in a weight-only A16 configuration. We do not claim these algorithms as new. Our contribution is the end-to-end scale-up characterisation: an external reproduction gate, matched 4B/8B capability analysis, cross-corpus perplexity, effective-bit accounting, lossless lattice-aware packing, and direct packed execution. The 8B model reaches a three-corpus perplexity ratio of 1.361x, with WikiText-2, C4, and PTB ratios of 1.318x, 1.393x, and 1.371x. On eight zero-shot tasks at n = 500, mean accuracy is 64.6% versus 72.4% for FP16, corresponding to 78.5% chance-corrected retention and a 7.8-point absolute cost. The matched 4B run retains 69.6%, yielding an 8.9-point 8B advantage. The packed checkpoint is 8.24 GiB and preserves the recorded perplexity to measurement precision. Direct packed execution reaches 15.52 tokens/s in 7.35 GiB, while a preliminary packed GEMV remains slower than FP16 cuBLAS. The result is a validated scale-up baseline: model size improves robustness to aggressive post-training discretisation, actual serialisation is solved for the measured artefact, and direct execution is feasible, while broader seeds, calibration distributions, and kernel optimisation remain open.

Safe Learning Under Irreversible Dynamics via Asking for Help

arXiv:2502.14043v3 Announce Type: replace-cross Abstract: Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot be recovered from. Instead, we allow the learning agent to ask for help from a mentor and to transfer knowledge between similar states. We show that this combination enables the agent to learn both safely and effectively. Under standard online learning assumptions, we provide an algorithm whose regret and number of mentor queries are both sublinear in the time horizon for Markov decision processes with irreversible dynamics and infinite state spaces. Our proof involves a sequence of three reductions, making our result more general than a single algorithm. Conceptually, our result may be the first formal proof that it is possible for an agent to obtain high reward while becoming self-sufficient in an unknown, unbounded, and high-stakes environment without resets.

Omni Interaction Agent Technical Report

arXiv:2609.08977v2 Announce Type: replace-cross Abstract: In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities, including video, speech, and text, enabling natural full-duplex interaction in both everyday conversations and complex workflow-oriented agent scenarios. Users can interrupt the model at any time, while the model can also proactively provide intermediate feedback or ask follow up questions. To natively support these capabilities, Gander adopts two key architectural designs: 1) It employs a Cerebellum-Brain collaborative framework, in which the Cerebellum is responsible for realtime interaction and omni conversational capabilities, while the Brain handles complex reasoning and higher-level agentic tasks. The two components interact continuously through tool calling and the agent orchestration runtime. 2) The Cerebellum is built upon a streaming Thinker-Talker architecture, user inputs and model outputs are further flattened into an ordered token stream at the chunk level, providing a unified representation for low latency, continuous interaction. We conduct comprehensive evaluations of Gander across four dimensions: conversational ability, omni understanding, interactive capability, and agentic intelligence. Internal human evaluations demonstrate that Gander maintains the natural and expressive spoken dialogue capabilities of SOTA open source models while achieving competitive performance in omni interaction. Gander also demonstrates robustness in challenging real-world scenarios, including background noise interference, multi-party interactions, and backchannel communication. We release Gander together with its models, code, and data to facilitate further research and development in the community.

CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures

arXiv:2605.25338v1 Announce Type: cross Abstract: Large language model (LLM) agents frequently fail on multi-step tasks involving reasoning, tool use, and environment interaction. While such failures are typically logged or retried heuristically, they contain structured signals about where execution broke down. We introduce CausalFlow, an interventional framework that converts failed agent traces into minimal counterfactual repairs and reusable supervision. CausalFlow models execution traces as sequential chains of dependent steps and computes Causal Responsibility Scores(CRS) via step-level counterfactual intervention to identify failure-inducing steps. For these steps, we generate minimally edited repairs that flip the final outcome to success, producing validated contrastive pairs of the form (wrong step, corrected step). CausalFlow supports two complementary uses: targeted test-time repair that recovers from failures with minimal behavioral drift, and training-time supervision suitable for offline preference optimization or reward modeling. Across four benchmarks spanning mathematical reasoning, code generation, question answering, and medical browsing, CausalFlow converts failed executions into validated minimal repairs with high minimality and causal-consensus scores, and demonstrates that causal attribution is necessary for reliable improvement across diverse agent tasks, outperforming heuristic refinement in complex retrieval settings while producing more localized repairs throughout. These results demonstrate that interventional analysis over structured execution traces provides a principled and scalable mechanism for transforming agent failures into reliability gains and learning-ready supervision.

De novo design of quasisymmetric two-component protein cages

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10464-0

Researchers designed two-component proteins forming quasisymmetric cages via geometric frustration, enabling tunable virus-like assemblies for cargo delivery, cellular uptake and studying intracellular diffusion and protein localization.
❌