Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
arXiv:2609.13009v1 Announce Type: new Abstract: Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluat
-
cs.AI, q-bio.NC updates on arXiv.org
-
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
arXiv:2605.21384v2 Announce Type: replace-cross Abstract: As long-horizon coding agents produce more code than any developer can review, oversight collapses onto a single surface: the automated test suite. Reward hacking naturally arises in this setup, as the agent optimizes for passing tests while deviating from the users true goal. We study this reward hacking phenomenon by decompose software engineering tasks into three parts: (i) a natural language description of the specification (ii) visi
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
arXiv:2605.04906v2 Announce Type: replace Abstract: While Large Language Models (LLMs) excel in certain reasoning tasks, they struggle in multi-agent games where the final outcome depends on the joint strategies of all agents. In multi-agent games, the non-stationarity of other agents brings significant challenges on the evaluation of the reasoning process and the credit assignment over multiple reasoning steps. Existing single-agent reinforcement learning (RL) approaches and their multi-agent
Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
-
Nature - Issue - nature.com science feeds
-
A SAUR gene enhances maize drought resilience by promoting silk elongation
Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10566-9The Small Auxin Up RNA (SAUR) protein ZmSAUR72 in maize (Zea mays) promotes silk growth via regulation of H+-ATPase activity, and is a key determinant of the anthesis-silking interval and thus resilience to drought.
A SAUR gene enhances maize drought resilience by promoting silk elongation
Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10566-9
The Small Auxin Up RNA (SAUR) protein ZmSAUR72 in maize (Zea mays) promotes silk growth via regulation of H+-ATPase activity, and is a key determinant of the anthesis-silking interval and thus resilience to drought.-
Nature - Issue - nature.com science feeds
-
Engineered immunosuppressive dendritic cells protect against cardiac remodelling
Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10346-5Lesion-targeted immune modulation is a feasible strategy to control cardiac fibrosis, and engineered dendritic cells are a promising therapeutic platform for treating cardiac remodelling and heart failure.
Engineered immunosuppressive dendritic cells protect against cardiac remodelling
Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10346-5
Lesion-targeted immune modulation is a feasible strategy to control cardiac fibrosis, and engineered dendritic cells are a promising therapeutic platform for treating cardiac remodelling and heart failure.-
cs.AI, q-bio.NC updates on arXiv.org
-
Search, Do not Guess: Teaching Small Language Models to Be Effective Search Agents
arXiv:2604.04651v1 Announce Type: new Abstract: Agents equipped with search tools have emerged as effective solutions for knowledge-intensive tasks. While Large Language Models (LLMs) exhibit strong reasoning capabilities, their high computational cost limits practical deployment for search agents. Consequently, recent work has focused on distilling agentic behaviors from LLMs into Small Language Models (SLMs). Through comprehensive evaluation on complex multi-hop reasoning tasks, we find that
Search, Do not Guess: Teaching Small Language Models to Be Effective Search Agents
-
MRD
-
Dynamic Targetable Extracellular Vesicle Surface Proteins Monitor Depth of Response to CAR T Therapy
Res Sq [Preprint]. 2026 Mar 18:rs.3.rs-8913641. doi: 10.21203/rs.3.rs-8913641/v1.ABSTRACTExtracellular vesicles (EVs) represent a promising liquid biopsy platform in multiple myeloma (MM). We developed an MM EV Surface Protein Assay to quantify and dynamically monitor four MM EV subpopulations defined by targetable MM surface proteins (BCMA, CD38, GPRC5D, and CD319) across 336 serial blood samples from 45 relapsed/refractory MM (RRMM) patients treated with anti-BCMA chimeric antigen receptor (CA
Dynamic Targetable Extracellular Vesicle Surface Proteins Monitor Depth of Response to CAR T Therapy
Res Sq [Preprint]. 2026 Mar 18:rs.3.rs-8913641. doi: 10.21203/rs.3.rs-8913641/v1.
ABSTRACT
Extracellular vesicles (EVs) represent a promising liquid biopsy platform in multiple myeloma (MM). We developed an MM EV Surface Protein Assay to quantify and dynamically monitor four MM EV subpopulations defined by targetable MM surface proteins (BCMA, CD38, GPRC5D, and CD319) across 336 serial blood samples from 45 relapsed/refractory MM (RRMM) patients treated with anti-BCMA chimeric antigen receptor (CAR) T-cell therapy. All four MM EV subpopulations significantly decreased in 43 patients with initial response, while BCMA+, GPRC5D+, and CD319+ MM EVs increased in 19 patients with progression, and antigen escape was detected by BCMA+ MM EVs. MM EV subpopulations differentiated minimal residual disease (MRD) status and complemented MRD for detecting early relapse before clinical progression. Notably, CD319+ MM EVs were early predictors of progression-free and overall survival in MRD-negative patients. This assay enables noninvasive monitoring of deep response, progression, and antigen escape, and stratifies survival in MRD-negative patients with RRMM.
PMID:41890853 | PMC:PMC13015583 | DOI:10.21203/rs.3.rs-8913641/v1
-
cs.AI, q-bio.NC updates on arXiv.org
-
Purify Once, Edit Freely: Breaking Image Protections under Model Mismatch
arXiv:2603.13028v1 Announce Type: cross Abstract: Diffusion models enable high-fidelity image editing but can also be misused for unauthorized style imitation and harmful content generation. To mitigate these risks, proactive image protection methods embed small, often imperceptible adversarial perturbations into images before sharing to disrupt downstream editing or fine-tuning. However, in realistic post-release scenarios, content owners cannot control downstream processing pipelines, and pro
Purify Once, Edit Freely: Breaking Image Protections under Model Mismatch
-
cs.AI, q-bio.NC updates on arXiv.org
-
Deconstructing Multimodal Mathematical Reasoning: Towards a Unified Perception-Alignment-Reasoning Paradigm
arXiv:2603.08291v1 Announce Type: new Abstract: Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems that involve both textual and visual modalities. However, current models still face significant challenges in real-world visual math tasks. They often misinterpret diagrams, fail to align mathematical symbols with visual evidence, and produce inconsistent reasoning steps. Moreover, existing evaluations mainly focus
Deconstructing Multimodal Mathematical Reasoning: Towards a Unified Perception-Alignment-Reasoning Paradigm
-
cs.AI, q-bio.NC updates on arXiv.org
-
Do Deployment Constraints Make LLMs Hallucinate Citations? An Empirical Study across Four Models and Five Prompting Regimes
arXiv:2603.07287v1 Announce Type: cross Abstract: LLMs are increasingly used to draft academic text and to support software engineering (SE) evidence synthesis, but they often hallucinate bibliographic references that look legitimate. We study how deployment-motivated prompting constraints affect citation verifiability in a closed-book setting. Using 144 claims (24 in SE&CS) and a deterministic verification pipeline (Crossref + Semantic Scholar), we evaluate two proprietary models (Claude S
Do Deployment Constraints Make LLMs Hallucinate Citations? An Empirical Study across Four Models and Five Prompting Regimes
-
cs.AI, q-bio.NC updates on arXiv.org
-
RLJP: Legal Judgment Prediction via First-Order Logic Rule-enhanced with Large Language Models
arXiv:2505.21281v2 Announce Type: replace Abstract: Legal Judgment Prediction (LJP) is a pivotal task in legal AI. Existing semantic-enhanced LJP models integrate judicial precedents and legal knowledge for high performance. But they neglect legal reasoning logic, a critical component of legal judgments requiring rigorous logical analysis. Although some approaches utilize legal reasoning logic for high-quality predictions, their logic rigidity hinders adaptation to case-specific logical framewo
RLJP: Legal Judgment Prediction via First-Order Logic Rule-enhanced with Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
APRES: An Agentic Paper Revision and Evaluation System
arXiv:2603.03142v1 Announce Type: cross Abstract: Scientific discoveries must be communicated clearly to realize their full potential. Without effective communication, even the most groundbreaking findings risk being overlooked or misunderstood. The primary way scientists communicate their work and receive feedback from the community is through peer review. However, the current system often provides inconsistent feedback between reviewers, ultimately hindering the improvement of a manuscript an
APRES: An Agentic Paper Revision and Evaluation System
-
cs.AI, q-bio.NC updates on arXiv.org
-
Model-agnostic Selective Labeling with Provable Statistical Guarantees
arXiv:2510.14581v3 Announce Type: replace-cross Abstract: Obtaining high-quality labels for large datasets is expensive, requiring massive annotations from human experts. While AI models offer a cost-effective alternative by predicting labels, their label quality is compromised by the unavoidable labeling errors. Existing methods mitigate this issue through selective labeling, where AI labels a subset and human labels the remainder. However, these methods lack theoretical guarantees on the qual