❌

Reading view

Language Is an Insufficient Substrate for Quantitative Reasoning, and Consequential Domains Need Large Quantitative Models

arXiv:2609.12105v1 Announce Type: new Abstract: The prevailing assumption in applied machine learning is that progress on consequential quantitative decisions such as pricing risk, allocating capital, triaging patients, or containing a network intrusion will follow from progress in large language models (LLMs). A language model is trained on a representation of the world that was produced by human description; description is a lossy encoding of the quantitative record, and the loss is irreversible: no downstream model, at any scale, can recover from a description what the description did not encode. We formalize this as a property of the representation on which a model is trained rather than of the model capacity, and we identify three further properties that consequential settings demand of a model and that a language substrate cannot supply by construction: reproducibility, lineage from every output back to the source records that produced. it, and calibrated uncertainty. We argue that these properties define a distinct model class, which we call the Large Quantitative Model (LQM).
  •  

Mined from Scientific Literature: Process Schemas for Atomic Layer Deposition and Etching in Materials Science

arXiv:2609.12139v1 Announce Type: new Abstract: Atomic layer deposition (ALD) and atomic layer etching (ALE) are reported heterogeneously across experimental and simulation literature in materials science, hindering comparison and machine-actionable reuse. We present four domain-expert-reviewed JSON Schemas for ALD and ALE experimental and simulation processes. Curated with schema-miner and grounded in QUDT using schema-miner pro, the schemas structure materials, process conditions, configurations , and measured or predicted results. We compare their scope, structure, and semantic grounding, and demonstrate their use for schema-guided literature extraction and publication of structured records through ORKG templates.
  •  

When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation

arXiv:2609.12718v1 Announce Type: new Abstract: Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading approach for assessing LLMs in medicine, but it is unclear whether rubric scores reflect such errors. We first study this in a controlled setting using MedHallu, finding that more specific rubrics better distinguish correct from hallucinated responses. To test this systematically, we develop a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that creates matched correct and error-injected responses. Across HealthBench, HealthBench Professional, and LiveMedBench, our clinically relevant hallucinations are missed by rubrics, often leaving scores unchanged. We find that rubrics are most effective when explicitly checking facts, and are less effective for additional or unexpected errors they do not anticipate. A preliminary retrieval-based factuality check recovers some of the rubric-blind errors, suggesting a complementary approach. These findings reveal systematic blind spots in current medical evaluation of LLMs and suggest that rubric scores alone are insufficient to establish clinical reliability, potentially undermining clinician trust and confidence in clinical deployment.
  •  

Assisted Spatial Cognition Through Vision-Language Models

arXiv:2609.12747v1 Announce Type: new Abstract: Multimodal AI, powered by Large Language Models (LLMs) and Vision-Language Models (VLMs), is transforming assistive technologies by enabling simultaneous processing of visual and textual data. This advancement holds significant promise for over 43 million visually impaired and neuro-divergent individuals worldwide who face persistent challenges in navigating indoor and outdoor environments due to limited spatial awareness and insufficient environmental cues. Existing navigation aids often lack comprehensive 3D scene understanding, relying on constrained route-based strategies that hinder user autonomy. In this paper, we introduce a novel end-to-end framework that integrates LLMs, VLMs and digital twin technologies to deliver a spatially cognitive navigation support for visually impaired and neuro-divergent users. Our system captures video input via standard mobile phone cameras, and employs SLAM3R to generate dense 3D point clouds from monocular RGB sequences in real-time. Our custom post-processing algorithm ensures accurate point cloud alignment across multiple viewpoints without requiring predefined reference points. This enhances the capabilities of SpatialLM to produce structured 3D representations, including architectural elements and oriented object bounding boxes. The enriched spatial data is then processed by a locally deployed LLM, which interprets 3D contexts to generate detailed scene descriptions and precise distance measurements between users and surrounding objects. We evaluated our approach across diverse video scenarios featuring various perspectives, looped walking views and captured in multiple environments. The evaluation results demonstrate consistent accuracy in 3D scene interpretation and object localisation, underscoring the potential of our system as a transformative assistive navigation solution that combines advanced visual perception with spatial reasoning
  •  

Scaling Clinical Judgment to Evaluate Medical AI

arXiv:2609.12822v1 Announce Type: new Abstract: Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific questions investigated and makes it unclear whether findings would be reproduced with a different set of evaluators. To more rigorously and scalably study clinical reasoning in AI models, here we introduce PrecepTron, an LLM fine-tuned for physician-level evaluation of open-ended responses. PrecepTron was trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples. We also release GRAND-ROUNDS, a new large-scale physician-annotated benchmark of 9,217 scores by 11 physicians across seven studies. We show that frontier LLMs in typical "LLM-as-a-judge" approaches often disagree with physicians and with each other, but fine-tuning PrecepTron on a small number of cases enables physician-level consistent scoring across tasks. We use PrecepTron to reproduce headline findings from five influential studies assessing LLMs for clinical care in JAMA, Science, and Nature Medicine without new human grading. Using PrecepTron, we then pose new questions about how LLMs reason in medicine that would have been infeasible with human grading alone, including measuring the diagnostic accuracy of frontier LLMs when clinical cases are provided piecemeal, even token by token. Together, PrecepTron and GRAND-ROUNDS provide a foundation for reproducible, large-scale study of how LLMs reason in medicine. All code, data, and labels are made freely available for researchers.
  •  

Diffusion Models and Concept Formation

arXiv:2609.13047v1 Announce Type: new Abstract: Humans organize knowledge into a taxonomy of concepts with nested levels of abstraction and a \emph{basic level} at which people recognize and name objects with the least cognitive effort. Cobweb is a classic cognitive account of this ability, an incremental learner that builds a probabilistic concept hierarchy by maximizing category utility. We argue that diffusion models, although designed for image synthesis, implicitly perform the same computation. The noisy marginals of a diffusion model are Gaussian smoothings of the data distribution, and the modes of these marginals form a hierarchy that corresponds to a Cobweb tree of probabilistic prototypes in four respects. Both are hierarchical density models, both are hierarchical-Bayesian models with Gaussian prototypes, both treat categorization as score-following that reduces uncertainty, and in both a basic level emerges. We locate this basic level for a diffusion model at an intermediate noise level, where recent analyses show that the reverse process commits to the class identity of a sample. The two models differ mainly in how they represent and learn the taxonomy. Cobweb learns a discrete tree incrementally, whereas a diffusion model encodes a continuous, interpolable hierarchy in a single learned score field fit to the data distribution. We test the correspondence on MNIST and Fashion-MNIST by recovering the diffusion hierarchy through mode-finding and comparing the basic levels of the two models. This reframes diffusion as a cognitive model of concept formation and offers Cobweb a continuous, scalable instantiation.
  •  

Demonstration of on-chip all-optical switching of magnetization in integrated photonics

Nature Nanotechnology, Published online: 14 September 2026; doi:10.1038/s41565-026-02281-3

This study demonstrates on-chip all-optical switching by integrating magnetic memory elements with photonic circuits, establishing a route towards fully integrated magneto-photonic systems for ultrafast and energy-efficient information technologies.
  •  

Management of acute infusion-related reactions in AAV gene therapy guided by mechanistic insight

Infusion-related reactions (IRRs) can occur with AAV gene therapy. IRRs were observed in two participants who received AAV8 gene therapy and in one who received AAV9 gene therapy. AAV gene therapy IRRs may be rate-related, possibly driven by complement activation. In our studies, a slow, staged infusion mitigated additional IRRs.
  •  

Integrated in vitro transcription and oligo-dT affinity chromatography enable multi-cycle reagent recycling for mRNA manufacturing

Kis and colleagues report an integrated sequential-batch IVT–oligo-dT process that links RNA synthesis and affinity capture through a shared buffer, enabling direct crude-IVT loading and flowthrough recycling. The workflow improves cap-analog utilization and raw-material efficiency while preserving functional mRNA expression across five cycles.
  •  

An mRNA–lipid nanoparticle vaccine targeting the Plasmodium vivax E140 antigen

Immunization with a nucleoside-modified mRNA-LNP vaccine encoding the conserved malaria antigen E140 induced durable humoral immunity in mice and generated invasion-blocking antibodies against Plasmodium vivax. The results highlight E140 as a promising candidate for next-generation malaria vaccines.
  •  

Functional analysis of TTN uORFs reveals context-dependent translational regulation

TTN truncating variants cause dilated cardiomyopathy and may be amenable to therapeutic upregulation. Disrupting TTN upstream open reading frames (uORFs) increased luciferase reporter expression, however disrupting endogenous uORFs in hiPSC-derived cardiomyocytes did not increase titin protein. This highlights the importance of validating regulatory mechanisms in disease-relevant cellular contexts.
  •  

Novel lipid nanoparticle in a mRNA cancer vaccine drives tumor control via type I IFNs and effector CD8+ T cells

Ionizable lipid components of LNP vaccine formulations play an integral role in providing the necessary immune stimulatory signals required for effective anti-tumor T cell responses. This study finds that inclusion of a novel ionizable lipid confers effective tumor control in part through the generation of type I interferons.
  •  

BRD4 Inhibition Mitigates Acute and Chronic Corneal Injury Following Topical Nitrogen Mustard Exposure

Lu and colleagues identify BRD4 as a central epigenetic driver of vesicant-induced corneal injury. Using reproducible mouse and rabbit models, they show that short-term topical BRD4 inhibition suppresses acute inflammation and provides durable protection of corneal clarity, stromal organization, endothelial integrity, and neovascularization, supporting translational therapy for chemical eye injuries.
  •  

Targeting of the oncogenic fusion EWSR1-FLI1 in Ewing Sarcoma by CRISPR/dCas9 silencers

Blancafort and colleagues describe a non-viral polymeric system for the delivery of dCas9-KRAB silencers as ribonucleoprotein (RNP) payloads for EWSR1-FLI1 repression. They demonstrate highly efficient RNP delivery and robust silencing of EWSR1-FLI1 in both cell line and patient-derived xenografts of Ewing sarcoma, accompanied by potent anti-tumor effects.
  •  

A renewed role for DNA vaccines for tuberculosis

Tuberculosis (TB) remains a leading cause of infection-related mortality worldwide, with an estimated 10.7 million incident cases and 1.23 million deaths recorded in 2024.1 The Bacillus Calmette-Guérin vaccine (BCG), despite over a century of use, provides unreliable protection against pulmonary disease in adolescents and adults, and the identification of antigens capable of inducing durable protective immunity in these populations remains an unsolved challenge. Against this backdrop, the emergence of plasmid DNA as a vaccine platform in the mid-1990s generated considerable scientific optimism.
  •  

In vivo engineering of T cells with a synthetic cytokine receptor enables selective enrichment and expansion of anti-CD22 CAR T cells

In preclinical studies, UB-VV400, an off-the-shelf, investigational lentiviral drug product, generates fully human anti-CD22 CAR T cells in vivo without the need for lymphodepletion. Activation of the synthetic rapamycin-activated cytokine receptor drives selective CAR T cell expansion and enrichment, resulting in complete tumor clearance and B cell depletion.
  •  
❌