❌

Reading view

Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning

arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amplifying learning pressure on uncertain, low-confidence tokens, leading to suboptimal training dynamics. We propose Trimmed Logit-Gap SFT (TrimSFT), a simple token-level reweighting method that scales the SFT loss according to the logit gap between the gold token and its strongest competitor. TrimSFT trims supervision away from both extremes: tokens already mastered (large logit gap) and tokens weakly supported by the current model (small or negative logit gap), concentrating learning within an intermediate logit-gap region between them. We instantiate this principle with a Gaussian weight centered at margin m with bandwidth {\tau}, requiring no reference model or additional forward pass. We evaluate TrimSFT on six base models from the Llama, Qwen, and DeepMath families across five mathematical reasoning benchmarks. TrimSFT consistently improves over standard SFT, achieving the best average performance on five out of six models, with gains of up to +26.9 points over SFT on MATH500. Further analyses show that the bandwidth {\tau} matters more than the exact margin location, and that half-trim variants that remove supervision pressure from only one side yield inferior trade-offs. A token-level logit-gap distribution analysis suggests that TrimSFT reshapes model confidence in a more balanced way than uniform SFT or monotonic reweighting methods. These results suggest that reasoning SFT can benefit from trimming both extremes rather than treating all tokens uniformly.
  •  

Early stage nonsmall cell lung cancer: Toward a risk-adaptive paradigm in the era of biologic precision

CA Cancer J Clin. 2026 Sep-Oct;76(5):e70100. doi: 10.3322/caac.70100.

ABSTRACT

The clinical landscape of early stage nonsmall cell lung cancer is at transformative crossroads. Driven by the widespread adoption of low-dose computed tomography screening, the frequent detection of ground-glass opacities, and a rising incidence among never-smokers, the diagnostic center of gravity has shifted toward earlier, potentially curable disease. This shift has been accompanied by equally important therapeutic advances, including parenchyma-sparing surgical techniques, minimally invasive platforms enhanced by digital navigation, and the transformative integration of perioperative immunotherapy and targeted agents. Concurrently, noninvasive monitoring approaches, such as liquid biopsy, have emerged as powerful tools to guide precision management. Despite this progress, substantial barriers to achieving a universal cure persist. Clinicians continue to face uncertainty in the management of ground-glass opacities, the anatomy-based TNM staging system fails to capture the biologic heterogeneity of early tumors, and global disparities in access to innovation remain unresolved. To address these challenges, the authors propose a shift toward a risk-adaptive management paradigm that harnesses artificial intelligence-driven analytics and multi-omics profiling to tailor treatment intensity according to each patient's biologic risk. Such an approach would enable appropriate escalation for high-risk individuals while permitting safe de-escalation for those at low risk. This holistic, lifespan-oriented strategy must be embraced to deliver equitable and durable cures for patients with early stage nonsmall cell lung cancer.

PMID:42713910 | PMC:PMC13555834 | DOI:10.3322/caac.70100

  •  

Early stage nonsmall cell lung cancer: Toward a risk-adaptive paradigm in the era of biologic precision

CA Cancer J Clin. 2026 Sep-Oct;76(5):e70100. doi: 10.3322/caac.70100.

ABSTRACT

The clinical landscape of early stage nonsmall cell lung cancer is at transformative crossroads. Driven by the widespread adoption of low-dose computed tomography screening, the frequent detection of ground-glass opacities, and a rising incidence among never-smokers, the diagnostic center of gravity has shifted toward earlier, potentially curable disease. This shift has been accompanied by equally important therapeutic advances, including parenchyma-sparing surgical techniques, minimally invasive platforms enhanced by digital navigation, and the transformative integration of perioperative immunotherapy and targeted agents. Concurrently, noninvasive monitoring approaches, such as liquid biopsy, have emerged as powerful tools to guide precision management. Despite this progress, substantial barriers to achieving a universal cure persist. Clinicians continue to face uncertainty in the management of ground-glass opacities, the anatomy-based TNM staging system fails to capture the biologic heterogeneity of early tumors, and global disparities in access to innovation remain unresolved. To address these challenges, the authors propose a shift toward a risk-adaptive management paradigm that harnesses artificial intelligence-driven analytics and multi-omics profiling to tailor treatment intensity according to each patient's biologic risk. Such an approach would enable appropriate escalation for high-risk individuals while permitting safe de-escalation for those at low risk. This holistic, lifespan-oriented strategy must be embraced to deliver equitable and durable cures for patients with early stage nonsmall cell lung cancer.

PMID:42713910 | DOI:10.3322/caac.70100

  •  

Human-AI Collaboration in Science at Scale: A Global Large-scale Randomized Field Experiment

arXiv:2605.24180v1 Announce Type: cross Abstract: Collaboration is the defining mode of modern science, yet its core mechanism -- feedback -- remains hard to observe, difficult to scale, and unequally distributed. Here we test whether large language models (LLMs) can contribute to this hidden but vital practice and reallocate scientific feedback, an essential yet scarce resource for knowledge production. In a global large-scale randomized field experiment, we delivered customized LLM-generated feedback for over 31,000 arXiv preprints across 150 fields and more than 45,000 researchers from 133 geographic regions. Relative to controls, authors who received feedback had a significantly higher likelihood of revising their manuscripts, corresponding to a 12.55% relative increase over the baseline revision rate. Exposure to AI feedback also increased authors' subsequent use of LLM tools in their future papers, suggesting longer-run shifts in scientific practice. These effects were strongest among authors from non-English-dominant research regions, manuscripts less embedded in the scholarly literature, and teams with lower h-indexes and earlier career stages, consistent with the idea that AI feedback may provide the greatest benefit where access to timely critique is otherwise limited. Together, these findings provide causal evidence that structured AI-based interventions can transform access to scientific feedback from a largely private advantage into a more widely distributed resource, with broader implications for productivity, equity, and capacity across the global research system.
  •  

STING signaling modulation by COPII cargo recognition

Lyu et al. identify the STING-ER-exit motif and the mechanism of its recognition by the COPII vesicle cargo-binding protein SEC24C. This study reveals how STING achieves controlled rather than constitutive ER exit and how COPII cargo recognition of STING can be modulated to control STING signaling.
  •  

METTL16 enhances proteasome inhibitor resistance in multiple myeloma by inhibiting eIF2α-PERK interaction and promoting PSMB5 translation

Oncogene, Published online: 13 March 2026; doi:10.1038/s41388-026-03706-y

METTL16 enhances proteasome inhibitor resistance in multiple myeloma by inhibiting eIF2α-PERK interaction and promoting PSMB5 translation
  •  

OSExpert: Computer-Use Agents Learning Professional Skills via Exploration

arXiv:2603.07978v1 Announce Type: new Abstract: General-purpose computer-use agents have shown impressive performance across diverse digital environments. However, our new benchmark, OSExpert-Eval, indicates they remain far less helpful than human experts. Although inference-time scaling enables adaptation, these agents complete complex tasks inefficiently with degraded performance, transfer poorly to unseen UIs, and struggle with fine-grained action sequences. To solve the problem, we introduce a GUI-based depth-first search (GUI-DFS) exploration algorithm to comprehensively explore and verify an environment's unit functions. The agent then exploits compositionality between unit skills to self-construct a curriculum for composite tasks. To support fine-grained actions, we curate a database of action primitives for agents to discover during exploration; these are saved as a skill set once the exploration is complete. We use the learned skills to improve the agent's performance and efficiency by (1) enriching agents with ready-to-use procedural knowledge, allowing them to plan only once for long trajectories and generate accurate actions, and (2) enabling them to end inference-time scaling earlier by realizing their boundary of capabilities. Extensive experiments show that our environment-learned agent takes a meaningful step toward expert-level computer use, achieving a around 20 percent performance gain on OSExpert-Eval and closing the efficiency gap to humans by around 80 percent
  •  

Long-Short Term Agents for Pure-Vision Bronchoscopy Robotic Autonomy

arXiv:2603.07909v1 Announce Type: cross Abstract: Accurate intraoperative navigation is essential for robot-assisted endoluminal intervention, but remains difficult because of limited endoscopic field of view and dynamic artifacts. Existing navigation platforms often rely on external localization technologies, such as electromagnetic tracking or shape sensing, which increase hardware complexity and remain vulnerable to intraoperative anatomical mismatch. We present a vision-only autonomy framework that performs long-horizon bronchoscopic navigation using preoperative CT-derived virtual targets and live endoscopic video, without external tracking during navigation. The framework uses hierarchical long-short agents: a short-term reactive agent for continuous low-latency motion control, and a long-term strategic agent for decision support at anatomically ambiguous points. When their recommendations conflict, a world-model critic predicts future visual states for candidate actions and selects the action whose predicted state best matches the target view. We evaluated the system in a high-fidelity airway phantom, three ex vivo porcine lungs, and a live porcine model. The system reached all planned segmental targets in the phantom, maintained 80\% success to the eighth generation ex vivo, and achieved in vivo navigation performance comparable to the expert bronchoscopist. These results support the preclinical feasibility of sensor-free autonomous bronchoscopic navigation.
  •  

EndoSERV: A Vision-based Endoluminal Robot Navigation System

arXiv:2603.08324v1 Announce Type: cross Abstract: Robot-assisted endoluminal procedures are increasingly used for early cancer intervention. However, the intricate, narrow and tortuous pathways within the luminal anatomy pose substantial difficulties for robot navigation. Vision-based navigation offers a promising solution, but existing localization approaches are error-prone due to tissue deformation, in vivo artifacts and a lack of distinctive landmarks for consistent localization. This paper presents a novel EndoSERV localization method to address these challenges. It includes two main parts, \textit{i.e.}, \textbf{SE}gment-to-structure and \textbf{R}eal-to-\textbf{V}irtual mapping, and hence the name. For long-range and complex luminal structures, we divide them into smaller sub-segments and estimate the odometry independently. To cater for label insufficiency, an efficient transfer technique maps real image features to the virtual domain to use virtual pose ground truth. The training phases of EndoSERV include an offline pretraining to extract texture-agnostic features, and an online phase that adapts to real-world conditions. Extensive experiments based on both public and clinical datasets have been performed to demonstrate the effectiveness of the method even without any real pose labels.
  •  

Towards Effective Orchestration of AI x DB Workloads

arXiv:2603.03772v1 Announce Type: cross Abstract: AI-driven analytics are increasingly crucial to data-centric decision-making. The practice of exporting data to machine learning runtimes incurs high overhead, limits robustness to data drift, and expands the attack surface, especially in multi-tenant, heterogeneous data systems. Integrating AI directly into database engines, while offering clear benefits, introduces challenges in managing joint query processing and model execution, optimizing end-to-end performance, coordinating execution under resource contention, and enforcing strong security and access-control guarantees. This paper discusses the challenges of joint DB-AI, or AIxDB, data management and query processing within AI-powered data systems. It presents various challenges that need to be addressed carefully, such as query optimization, execution scheduling, and distributed execution over heterogeneous hardware. Database components such as transaction management and access control need to be re-examined to support AI lifecycle management, mitigate data drift, and protect sensitive data from unauthorized AI operations. We present a design and preliminary results to demonstrate what may be key to the performance for serving AIxDB queries.
  •  

GeoSeg: Training-Free Reasoning-Driven Segmentation in Remote Sensing Imagery

arXiv:2603.03983v1 Announce Type: cross Abstract: Recent advances in MLLMs are reframing segmentation from fixed-category prediction to instruction-grounded localization. While reasoning based segmentation has progressed rapidly in natural scenes, remote sensing lacks a generalizable solution due to the prohibitive cost of reasoning-oriented data and domain-specific challenges like overhead viewpoints. We present GeoSeg, a zero-shot, training-free framework that bypasses the supervision bottleneck for reasoning-driven remote sensing segmentation. GeoSeg couples MLLM reasoning with precise localization via: (i) bias-aware coordinate refinement to correct systematic grounding shifts and (ii) a dual-route prompting mechanism to fuse semantic intent with fine-grained spatial cues. We also introduce GeoSeg-Bench, a diagnostic benchmark of 810 image--query pairs with hierarchical difficulty levels. Experiments show that GeoSeg consistently outperforms all baselines, with extensive ablations confirming the effectiveness and necessity of each component.
  •  

Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews

arXiv:2403.07183v3 Announce Type: replace-cross Abstract: We present an approach for estimating the fraction of text in a large corpus which is likely to be substantially modified or produced by a large language model (LLM). Our maximum likelihood model leverages expert-written and AI-generated reference texts to accurately and efficiently examine real-world LLM-use at the corpus level. We apply this approach to a case study of scientific peer review in AI conferences that took place after the release of ChatGPT: ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023. Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates. The circumstances in which generated text occurs offer insight into user behavior: the estimated fraction of LLM-generated text is higher in reviews which report lower confidence, were submitted close to the deadline, and from reviewers who are less likely to respond to author rebuttals. We also observe corpus-level trends in generated text which may be too subtle to detect at the individual level, and discuss the implications of such trends on peer review. We call for future interdisciplinary work to examine how LLM use is changing our information and knowledge practices.
  •  

TabTracer: Monte Carlo Tree Search for Complex Table Reasoning with Large Language Models

arXiv:2602.14089v1 Announce Type: cross Abstract: Large language models (LLMs) have emerged as powerful tools for natural language table reasoning, where there are two main categories of methods. Prompt-based approaches rely on language-only inference or one-pass program generation without step-level verification. Agent-based approaches use tools in a closed loop, but verification is often local and backtracking is limited, allowing errors to propagate and increasing cost. Moreover, they rely on chain- or beam-style trajectories that are typically combinatorially redundant, leading to high token costs. In this paper, we propose TabTracer, an agentic framework that coordinates multi-step tool calls over intermediate table states, with explicit state tracking for verification and rollback. First, it enforces step-level verification with typed operations and lightweight numeric and format checks to provide reliable rewards and suppress hallucinations. Second, execution-feedback Monte Carlo Tree Search maintains a search tree of candidate table states and uses backpropagated reflection scores to guide UCB1 selection and rollback via versioned snapshots. Third, it reduces redundancy with budget-aware pruning, deduplication, and state hashing with a monotonicity gate to cut token cost. Comprehensive evaluation on TabFact, WikiTQ, and CRT datasets shows that TabTracer outperforms state-of-the-art baselines by up to 6.7% in accuracy while reducing token consumption by 59--84%.
  •  

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

arXiv:2601.23232v3 Announce Type: replace-cross Abstract: In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Open-domain video shot retrieval, which involves richer temporal structure and more complex semantics, still lacks systematic benchmarks and analysis. To fill this gap, we introduce ShotFinder, a benchmark that formalizes editing requirements as keyframe-oriented shot descriptions and introduces five types of controllable single-factor constraints: Temporal order, Color, Visual style, Audio, and Resolution. We curate 1,210 high-quality samples from YouTube across 20 thematic categories, using large models for generation with human verification. Based on the benchmark, we propose ShotFinder, a text-driven three-stage retrieval and localization pipeline: (1) query expansion via video imagination, (2) candidate video retrieval with a search engine, and (3) description-guided temporal localization. Experiments on multiple closed-source and open-source models reveal a significant gap to human performance, with clear imbalance across constraints: temporal localization is relatively tractable, while color and visual style remain major challenges. These results reveal that open-domain video shot retrieval is still a critical capability that multimodal large models have yet to overcome.
  •  
❌