Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Analysis of LLM Performance on AWS Bedrock: Receipt-item Categorisation Case Study
arXiv:2604.01615v1 Announce Type: new Abstract: This paper presents a systematic, cost-aware evaluation of large language models (LLMs) for receipt-item categorisation within a production-oriented classification framework. We compare four instruction-tuned models available through AWS Bedrock: Claude 3.7 Sonnet, Claude 4 Sonnet, Mixtral 8x7B Instruct, and Mistral 7B Instruct. The aim of the study was (1) to assess performance across accuracy, response stability, and token-level cost, and (2) to
-
cs.AI, q-bio.NC updates on arXiv.org
-
Predicting LLM Output Length via Entropy-Guided Representations
arXiv:2602.11812v2 Announce Type: replace Abstract: The long-tailed distribution of sequence lengths in LLM serving and reinforcement learning (RL) sampling causes significant computational waste due to excessive padding in batched inference. Existing methods rely on auxiliary models for static length prediction, but they incur high overhead, generalize poorly, and fail in stochastic "one-to-many" sampling scenarios. We introduce a lightweight framework that reuses the main model's internal hid
Predicting LLM Output Length via Entropy-Guided Representations
-
Cell
-
Genetically encoded fluorescent reporters to visualize α-synuclein pathology in live brain
The development of genetically encoded fluorescent reporters, along with their corresponding knock-in mouse lines for labeling α-Syn inclusions, enables diverse applications in studying the propagation and pathological effects of α-Syn inclusions in the live brain.
Genetically encoded fluorescent reporters to visualize α-synuclein pathology in live brain
-
MRD
-
Dynamic Targetable Extracellular Vesicle Surface Proteins Monitor Depth of Response to CAR T Therapy
Res Sq [Preprint]. 2026 Mar 18:rs.3.rs-8913641. doi: 10.21203/rs.3.rs-8913641/v1.ABSTRACTExtracellular vesicles (EVs) represent a promising liquid biopsy platform in multiple myeloma (MM). We developed an MM EV Surface Protein Assay to quantify and dynamically monitor four MM EV subpopulations defined by targetable MM surface proteins (BCMA, CD38, GPRC5D, and CD319) across 336 serial blood samples from 45 relapsed/refractory MM (RRMM) patients treated with anti-BCMA chimeric antigen receptor (CA
Dynamic Targetable Extracellular Vesicle Surface Proteins Monitor Depth of Response to CAR T Therapy
Res Sq [Preprint]. 2026 Mar 18:rs.3.rs-8913641. doi: 10.21203/rs.3.rs-8913641/v1.
ABSTRACT
Extracellular vesicles (EVs) represent a promising liquid biopsy platform in multiple myeloma (MM). We developed an MM EV Surface Protein Assay to quantify and dynamically monitor four MM EV subpopulations defined by targetable MM surface proteins (BCMA, CD38, GPRC5D, and CD319) across 336 serial blood samples from 45 relapsed/refractory MM (RRMM) patients treated with anti-BCMA chimeric antigen receptor (CAR) T-cell therapy. All four MM EV subpopulations significantly decreased in 43 patients with initial response, while BCMA+, GPRC5D+, and CD319+ MM EVs increased in 19 patients with progression, and antigen escape was detected by BCMA+ MM EVs. MM EV subpopulations differentiated minimal residual disease (MRD) status and complemented MRD for detecting early relapse before clinical progression. Notably, CD319+ MM EVs were early predictors of progression-free and overall survival in MRD-negative patients. This assay enables noninvasive monitoring of deep response, progression, and antigen escape, and stratifies survival in MRD-negative patients with RRMM.
PMID:41890853 | PMC:PMC13015583 | DOI:10.21203/rs.3.rs-8913641/v1
-
Nature Medicine
-
In vivo generation of anti-BCMA CAR-T cells in relapsed or refractory multiple myeloma: a phase 1 study
Nature Medicine, Published online: 25 March 2026; doi:10.1038/s41591-026-04244-6In a phase 1 trial, the in vivo generation of anti-BCMA CAR-T cells by lentiviral delivery was feasible and did not lead to dose-limiting toxicities in five patients with relapsed or refractory multiple myeloma.
In vivo generation of anti-BCMA CAR-T cells in relapsed or refractory multiple myeloma: a phase 1 study
Nature Medicine, Published online: 25 March 2026; doi:10.1038/s41591-026-04244-6
In a phase 1 trial, the in vivo generation of anti-BCMA CAR-T cells by lentiviral delivery was feasible and did not lead to dose-limiting toxicities in five patients with relapsed or refractory multiple myeloma.-
cs.AI, q-bio.NC updates on arXiv.org
-
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
arXiv:2603.12298v1 Announce Type: cross Abstract: Activation engineering enables precise control over Large Language Models (LLMs) without the computational cost of fine-tuning. However, existing methods deriving vectors from static activation differences are susceptible to high-dimensional noise and layer-wise semantic drift, often capturing spurious correlations rather than the target intent. To address this, we propose Global Evolutionary Refined Steering (GER-steer), a training-free framewo
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
-
cs.AI, q-bio.NC updates on arXiv.org
-
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
arXiv:2602.12670v3 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 86 tasks across 11 domains paired with curated Skills and deterministic verifiers. Each task is evaluated under three conditions: no Skills, curated Skills, and self-generated Skills. We test 7 agent-model configurat
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
-
cs.AI, q-bio.NC updates on arXiv.org
-
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
arXiv:2602.07075v4 Announce Type: replace-cross Abstract: Chemical large language models (LLMs) predominantly rely on explicit Chain-of-Thought (CoT) in natural language to perform complex reasoning. However, chemical reasoning is inherently continuous and structural, and forcing it into discrete linguistic tokens introduces a fundamental representation mismatch that constrains both efficiency and performance. We introduce LatentChem, a latent reasoning interface that decouples chemical computa
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images
arXiv:2603.08486v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods require explicit safety labels or contrastive data; yet, threat-related concepts are concrete and visually depictable, while safety concepts, like helpfulness, are abstract and lack visual referents. Inspired by the Self-Fulfilling mechanism underlying emergent misalignment, we propose Visual Self-Fulfi
Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images
-
cs.AI, q-bio.NC updates on arXiv.org
-
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
arXiv:2602.12670v2 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 86 tasks across 11 domains paired with curated Skills and deterministic verifiers. Each task is evaluated under three conditions: no Skills, curated Skills, and self-generated Skills. We test 7 agent-model configurat
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
-
cs.AI, q-bio.NC updates on arXiv.org
-
In-Run Data Shapley for Adam Optimizer
arXiv:2602.00329v3 Announce Type: replace-cross Abstract: Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley value serving as the theoretical gold standard. While recent "In-Run" methods bypass the prohibitive cost of retraining by estimating contributions dynamically, they heavily rely on the linear structure of Stochastic Gradient Descent (SGD) and fail to capture the complex dynamics of adaptive optimizers
In-Run Data Shapley for Adam Optimizer
-
cs.AI, q-bio.NC updates on arXiv.org
-
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
arXiv:2602.14201v1 Announce Type: cross Abstract: The "thinking-with-images" paradigm enables multimodal large language models (MLLMs) to actively explore visual scenes via zoom-in tools. This is essential for ultra-high-resolution (UHR) remote sensing VQA, where task-relevant cues are sparse and tiny. However, we observe a consistent failure mode in existing zoom-enabled MLLMs: Tool Usage Homogenization, where tool calls collapse into task-agnostic patterns, limiting effective evidence acquisi