❌

Reading view

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

arXiv:2609.13005v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provide limited evidence of reliability on real-world tasks that involve following manuals spanning hundreds of pages with complex, interdependent guidelines. In this paper, we introduce Tasks over Application Manuals (TAM), a benchmark for evaluating long-horizon procedural reasoning. We construct TAM by curating real-world tasks from two domains: ICD-10-CM clinical coding (mapping medical conditions to diagnostic codes) and U.S. federal sentencing (computing crime sentencing guideline outcomes, specifically offense levels), with human-validated labels. Each task requires following an authoritative manual with tens of thousands of rules and executing a sequence of interdependent steps across different sections to produce an exact answer. We evaluate general-purpose prompting approaches, including retrieval-augmented generation, ReAct-style prompting, and an agent-harness baseline on GPT-5, and find that the best exact-match performance remains extremely low: 1% on ICD-10-CM coding and 15.5% on sentencing tasks. These results show that current benchmarks may overestimate LLM reasoning ability and miss a key challenge: reliably following long, rule-based procedures. The complete TAM data and code are publicly available.
  •  

CDO1 as a prognostic biomarker and therapeutic target in gastric cancer: Mechanistic insights into the PI3K/AKT-THBS1 axis and epigenetic reactivation by decitabine

Clin Transl Med. 2026 Sep;16(9):e70784. doi: 10.1002/ctm2.70784.

ABSTRACT

BACKGROUND: As a pivotal metabolic enzyme, cysteine dioxygenase type 1 (CDO1) exerts tumour-suppressive effects across diverse tumour types, and its expression is strongly correlated with clinical prognosis. However, the molecular mechanisms underlying CDO1-mediated tumour suppression in gastric cancer (GC), its relationship with the tumour-associated immune microenvironment, and pharmacological strategies to restore its expression remain poorly understood.

METHODS: CDO1 expression and prognosis were evaluated by multi-omics and tissue microarray analyses. Tumour microenvironment and immune infiltration were analyzed using ESTIMATE and ssGSEA. Downstream pathways and interacting proteins were identified by transcriptomics, co-immunoprecipitation, and GST pull-down. CDO1 function was assessed by proliferation, apoptosis, and migration assays in gain- and loss-of-function models. In vivo tumorigenesis and CDO1-dependent decitabine efficacy were evaluated by subcutaneous xenografts. Patient-derived organoids were used to assess decitabine sensitivity and 5-FU synergy.

RESULTS: Compared with normal controls, CDO1 expression was notably decreased in GC tissues, and its low expression was strongly linked to unfavourable prognosis, supporting its utility as a biomarker for prognosis. Elevated CDO1 levels correlated with an immune-active tumour microenvironment and reduced metastatic signatures. Mechanistically, CDO1 directly bound to PI3K p85α, disrupting p85α-p110α dimerization, thereby attenuating PI3K/AKT phosphorylation and downregulating THBS1 expression. CDO1 overexpression led to reduced proliferation, invasiveness, and EMT, accompanied by increased apoptosis. These effects were reversed by PI3K activation or THBS1 co-overexpression. Decitabine was identified as an agent that epigenetically restores CDO1 expression. Critically, CDO1 knockdown significantly attenuated the anti-tumour efficacy of decitabine in vivo, confirming that decitabine acts primarily through CDO1 reactivation. Decitabine synergized with 5-FU in both organoids and xenografts.

CONCLUSIONS: Our data identify CDO1 as both a biomarker for prognosis and a tumour suppressor in gastric cancer. They reveal a CDO1-PI3K/AKT-THBS1 signalling axis and support the epigenetic reactivation of CDO1 by decitabine as a translatable therapeutic strategy.

KEY POINTS: CDO1 is frequently downregulated in gastric cancer and serves as an independent favourable prognostic biomarker. CDO1 directly binds PI3K p85α, disrupting p85α-p110α dimerization to suppress the PI3K/AKT-THBS1 signalling axis. Decitabine epigenetically restores CDO1 expression, and its anti-tumour activity is critically CDO1-dependent in vivo. Combining decitabine with 5-FU synergistically overcomes gastric cancer growth in patient-derived organoids and subcutaneous xenograft models.

PMID:42670236 | PMC:PMC13527532 | DOI:10.1002/ctm2.70784

  •  

Protein glycosylation profiling in lung adenocarcinoma and precursor lesions: analysis of FFPE tissue sections

Anal Bioanal Chem. 2026 Jul 27. doi: 10.1007/s00216-026-06702-z. Online ahead of print.

ABSTRACT

Protein glycosylation is a major post-translational modification that regulates tumor initiation and progression; however, its dynamic modeling during multistep evolution of lung adenocarcinoma (LUAD) remains poorly understood, particularly in clinically archived tissues. Here, we established an integrated multi-omics workflow combining global proteomes, N-glycans, and site-specific intact N-glycopeptides to comprehensively characterize glycosylation in formalin-fixed paraffin-embedded (FFPE) specimens spanning four pathological stages of LUAD progression: inflammatory nodules (IN), atypical adenomatous hyperplasia (AAH), adenocarcinoma in situ (AIS), and invasive adenocarcinoma (IAC). Using optimized protein extraction, hydrophilic interaction liquid chromatography (HILIC)-based glycopeptide enrichment, and high-resolution LC-MS/MS, we achieved large-scale identification of proteins, N-glycans, and intact glycopeptides from archival clinical samples. Integrated analyses revealed progressive remodeling of site-specific N-glycosylation during malignant transformation, characterized by increased glycan branching, fucosylation, and sialylation during the transition from premalignant lesions to invasive cancer. Sialylated glycans reached their highest abundance in the premalignant AAH stage, whereas highly branched and fucosylated complex N-glycans predominated in invasive adenocarcinoma, indicating stage-dependent glycan remodeling throughout disease progression. Functional enrichment analyses linked these glycosylation alterations to extracellular matrix organization, neutrophil degranulation, and immune-associated pathways, while representative glycoproteins, including CEACAM6 and FGB, exhibited coordinated changes in protein abundance and site-specific glycoform micro-heterogeneity across pathological stages. Collectively, this study demonstrates the feasibility of deep glycoproteomic profiling using archived FFPE tissues and provides a comprehensive molecular atlas of glycosylation remodeling during LUAD progression. These findings establish a valuable resource for elucidating disease mechanisms and identifying stage-specific glycosylation biomarkers and potential glycan-targeted therapeutic candidates for early lung adenocarcinoma.

PMID:42509285 | DOI:10.1007/s00216-026-06702-z

  •  

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation

arXiv:2605.24675v1 Announce Type: cross Abstract: Translating text embedded in Web images is crucial for improving content accessibility and cross-lingual information retrieval, particularly within social media and e-commerce domains. Although Large Vision-Language Models (LVLMs) have advanced multimodal understanding, applying them to Web image translation remains challenging due to the visual representation gap: standard encoders often prioritize high-level semantics over the fine-grained visual details required for recognizing diverse character morphologies. To address this challenge, we propose VaaWIT, an end-to-end framework that adapts Large Language Models for multilingual Web image translation. The framework introduces two key technical contributions: (1) a Dual-Stream Attention Module (DSAM), which facilitates bidirectional interaction between multilingual semantic features and detailed visual representations, thereby synthesizing unified features robust to textual variations; and (2) a Visual-Aware Adapter (VAA), a parameter-efficient fine-tuning strategy that dynamically injects these fused visual cues into the frozen LLM backbone. This design enables the model to align the visual context with linguistic reasoning effectively while minimizing computational costs. Extensive experiments on eight tasks on three public benchmarks demonstrate that VaaWIT significantly outperforms state-of-the-art (SOTA) open-source baselines and achieves competitive performance against proprietary models. These results validate the efficacy of integrating fine-grained visual perception into LLMs for complex Web content analysis.
  •  

Kaempferol functionally reprograms CD47 signaling to promote cytoprotection and attenuate oxeiptosis in severe acute pancreatitis

Phytomedicine. 2026 May 15;157:158305. doi: 10.1016/j.phymed.2026.158305. Online ahead of print.

ABSTRACT

BACKGROUND: Severe acute pancreatitis (SAP) lacks targeted therapies, and massive loss of functional pancreatic acinar cells (PAC) drives mortality. Kaempferol (KA) possesses well-established anti-inflammatory and cytoprotective activities and is derived from herbal medicinal plants, but its direct molecular targets and mechanism of action in SAP remain undefined.

PURPOSE: To evaluate the protective effects of KA against SAP and to elucidate its molecular mechanism of specific action, with a focus on identifying the direct cellular target through which KA exerts its cytoprotective effects.

STUDY DESIGN: Gain‑/loss‑of‑function in vitro and PAC‑specific CD47 SAP mouse models, combined with multi‑omics screening and biophysical assays.

METHODS: CD47 manipulation (siRNA/overexpression) was performed in primary PACs and cell lines, combined with WT/CD47-/-/Mist1‑CD47‑iOE (PAC‑specific) mouse models. Network pharmacology, transcriptomics and proteomics were integrated to screen and validate KA's protective effects. Computational‑experimental approaches (molecular docking/dynamics, CETSA, SPR, co‑IP, pharmacological epistasis) characterized KA's allosteric modulation of CD47 signaling.

RESULTS: CD47 was upregulated in SAP; its knockout reduced PAC death via KEAP1/PGAM5/AIFM1-driven oxeiptosis. KA reduced PAC death across genotypes, afforded no extra benefit in CD47-KO, and was not overridden by CD47‑OE. Mechanistically, KA allosterically binds CD47 ectodomain, stabilizes the CD47‑ UBQLN1 complex, and redirects signaling from Gαi‑mediated death to Gβγ/ ERK/NRF2‑mediated survival. ERK inhibition attenuated KA's protection. KA's action was CD47‑dependent.

CONCLUSION: This study identifies anti-oxeiptosis as a novel pharmacological activity of KA in SAP. This is achieved through allosteric modulation of CD47, redirecting its signaling from death‑promoting to a protective axis via activating Gβγ/ERK/NRF2 to suppress oxeiptosis. These findings reveal the CD47‑oxeiptosis axis as a therapeutic target and position KA as a promising candidate for SAP therapy, adding a new mechanistic dimension to KA's known pharmacological profile.

PMID:42184499 | DOI:10.1016/j.phymed.2026.158305

  •  

AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems

arXiv:2604.03425v1 Announce Type: cross Abstract: Fully Homomorphic Encryption (FHE) enables privacy-preserving Transformer inference, but long-sequence encrypted Transformers quickly exceed single-GPU memory capacity because encoded weights are already large and encrypted activations grow rapidly with sequence length. Multi-GPU execution therefore becomes unavoidable, yet scaling remains challenging because communication is jointly induced by application-level aggregation and encryption-level RNS coupling. Existing approaches either synchronize between devices frequently or replicate encrypted tensors across devices, leading to excessive communication and latency. We present AEGIS, an Application-Encryption Guided Inference System for scalable long-sequence encrypted Transformer inference on multi-GPU platforms. AEGIS derives device placement from ciphertext dependencies jointly induced by Transformer dataflow and CKKS polynomial coupling, co-locating modulus-coherent and token-coherent data so that communication is introduced only when application dependencies require it, while reordering polynomial operators to overlap the remaining collectives with computation. On 2048-token inputs, AEGIS reduces inter-GPU communication by up to 57.9% in feed-forward networks and 81.3% in self-attention versus prior state-of-the-art designs. On four GPUs, it achieves up to 96.62% scaling efficiency, 3.86x end-to-end speedup, and 69.1% per-device memory reduction. These results establish coordinated application-encryption parallelism as a practical foundation for scalable homomorphic Transformer inference.
  •  

AromaGen: Interactive Generation of Rich Olfactory Experiences with Multimodal Language Models

arXiv:2604.01650v1 Announce Type: cross Abstract: Smell's deep connection with food, memory, and social experience has long motivated researchers to bring olfaction into interactive systems. Yet most olfactory interfaces remain limited to fixed scent cartridges and pre-defined generation patterns, and the scarcity of large-scale olfactory datasets has further constrained AI-based approaches. We present AromaGen, an AI-powered wearable interface capable of real-time, general-purpose aroma generation from free-form text or visual inputs. AromaGen is powered by a multimodal LLM that leverages latent olfactory knowledge to map semantic inputs to structured mixtures of 12 carefully selected base odorants, released through a neck-worn dispenser. Users can iteratively refine generated aromas through natural language feedback via in-context learning. Through a controlled user study ($N = 26$), AromaGen matches human-composed mixtures in zero-shot generation and significantly surpasses them after iterative refinement, achieving a median similarity of 8/10 to real food aromas and reducing perceived artificiality to levels comparable to real food. AromaGen is a step towards real-world interactive aroma generation, opening new possibilities for communication, wellbeing, and immersive technologies.
  •  

DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models

arXiv:2601.04823v5 Announce Type: replace Abstract: Mixture-of-Experts (MoE) has become a prominent paradigm for scaling Large Language Models (LLMs). Parameter-efficient fine-tuning methods, such as LoRA, are widely adopted to adapt pretrained MoE LLMs to downstream tasks. However, existing approaches typically assign identical LoRA ranks to all expert modules, ignoring the heterogeneous specialization of pretrained experts. This uniform allocation leads to a resource mismatch: task-relevant experts are under-provisioned, while less relevant ones receive redundant parameters. To address this, we propose DR-LoRA, a Dynamic Rank LoRA framework for fine-tuning pretrained MoE models. Specifically, DR-LoRA initializes all expert LoRA modules with a small active rank and uses an expert saliency score, which combines routing frequency and gradient-based rank importance, to identify which experts would benefit most from additional capacity. It then periodically expands the active ranks of the task-critical expert LoRA, progressively constructing a heterogeneous rank distribution tailored to the target task. Experiments on three MoE models across six tasks show that DR-LoRA consistently outperforms LoRA and other strong baselines, demonstrating that task-adaptive heterogeneous rank allocation is an effective strategy to improve active capacity utilization in MoE fine-tuning.
  •  

Scalable AI-assisted Workflow Management for Detector Design Optimization Using Distributed Computing

arXiv:2603.30014v1 Announce Type: cross Abstract: The Production and Distributed Analysis (PanDA) system, originally developed for the ATLAS experiment at the CERN Large Hadron Collider (LHC), has evolved into a robust platform for orchestrating large-scale workflows across distributed computing resources. Coupled with its intelligent Distributed Dispatch and Scheduling (iDDS) component, PanDA supports AI/ML-driven workflows through a scalable and flexible workflow engine. We present an AI-assisted framework for detector design optimization that integrates multi-objective Bayesian optimization with the PanDA--iDDS workflow engine to coordinate iterative simulations across heterogeneous resources. The framework addresses the challenge of exploring high-dimensional parameter spaces inherent in modern detector design. We demonstrate the framework using benchmark problems and realistic studies of the ePIC and dRICH detectors for the Electron-Ion Collider (EIC). Results show improved automation, scalability, and efficiency in multi-objective optimization. This work establishes a flexible and extensible paradigm for AI-driven detector design and other computationally intensive scientific applications.
  •  

QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation

arXiv:2507.13266v4 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a central paradigm for training large language models (LLMs) in reasoning tasks. Yet recent studies question RL's ability to incentivize reasoning capacity beyond the base model. This raises a key challenge: how can RL be adapted to solve harder reasoning problems more effectively? To address this challenge, we propose a simple yet effective strategy via Question Augmentation: introduce partial solutions during training to reduce problem difficulty and provide more informative learning signals. Our method, QuestA, when applied during RL training on math reasoning tasks, not only improves pass@1 but also pass@k-particularly on problems where standard RL struggles to make progress. This enables continual improvement over strong open-source models such as DeepScaleR and OpenMath Nemotron, further enhancing their reasoning capabilities. We achieve new state-of-the-art results on math benchmarks using 1.5B-parameter models: 72.50% (+10.73%) on AIME24, 62.29% (+12.79%) on AIME25, and 41.67% (+10.11%) on HMMT25. Code, data and model are available at https://github.com/foreverlasting1202/QuestA.
  •  

mAVE: A Watermark for Joint Audio-Visual Generation Models

arXiv:2603.07090v1 Announce Type: cross Abstract: As Joint Audio-Visual Generation Models see widespread commercial deployment, embedding watermarks has become essential for protecting vendor copyright and ensuring content provenance. However, existing techniques suffer from an architectural mismatch by treating modalities as decoupled entities, exposing a critical Binding Vulnerability. Adversaries exploit this via Swap Attacks by replacing authentic audio with malicious deepfakes while retaining the watermarked video. Because current detectors rely on independent verification ($Video_{wm}\vee Audio_{wm}$), they incorrectly authenticate the manipulated content, falsely attributing harmful media to the original vendor and severely damaging their reputation. To address this, we propose mAVE (Manifold Audio-Visual Entanglement), the first watermarking framework natively designed for joint architectures. mAVE cryptographically binds audio and video latents at initialization without fine-tuning, defining a Legitimate Entanglement Manifold via Inverse Transform Sampling. Experiments on state-of-the-art models (LTX-2, MOVA) demonstrate that mAVE guarantees performance-losslessness and provides an exponential security bound against Swap Attacks. Achieving near-perfect binding integrity ($>99\%$), mAVE offers a robust cryptographic defense for vendor copyright.
  •  

Contextualized Privacy Defense for LLM Agents

arXiv:2603.02983v1 Announce Type: cross Abstract: LLM agents increasingly act on users' personal information, yet existing privacy defenses remain limited in both design and adaptability. Most prior approaches rely on static or passive defenses, such as prompting and guarding. These paradigms are insufficient for supporting contextual, proactive privacy decisions in multi-step agent execution. We propose Contextualized Defense Instructing (CDI), a new privacy defense paradigm in which an instructor model generates step-specific, context-aware privacy guidance during execution, proactively shaping actions rather than merely constraining or vetoing them. Crucially, CDI is paired with an experience-driven optimization framework that trains the instructor via reinforcement learning (RL), where we convert failure trajectories with privacy violations into learning environments. We formalize baseline defenses and CDI as distinct intervention points in a canonical agent loop, and compare their privacy-helpfulness trade-offs within a unified simulation framework. Results show that our CDI consistently achieves a better balance between privacy preservation (94.2%) and helpfulness (80.6%) than baselines, with superior robustness to adversarial conditions and generalization.
  •  

A Very Big Video Reasoning Suite

arXiv:2602.20159v1 Announce Type: cross Abstract: Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindered by the lack of large-scale training data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks following a principled taxonomy and over one million video clips, approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, benchmark toolkit, and models are publicly available at https://video-reason.com/ .
  •  

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

arXiv:2510.10689v2 Announce Type: replace Abstract: Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed benchmark dedicated to assessing synergistic audio-visual understanding, with a strong emphasis on modality complementarity and logical consistency. Specifically, OmniVideoBench comprises 1000 high-quality question-answer(QA) pairs, each annotated with step-by-step reasoning traces, derived from 628 diverse videos ranging from several seconds to 30 minutes, and manually verified to guarantee complete correctness and uniqueness. Moreover, OmniVideoBench encompasses 13 carefully designed question types, covering temporal reasoning, spatial localization, counting, causal inference, summarization, and beyond, thereby capturing the essential challenges of video understanding. Evaluation of multiple MLLMs on OmniVideoBench reveals a pronounced gap between model performance and human reasoning, with open-source models lagging significantly behind their closed-source counterparts, underscoring the inherent difficulty of genuine audio-visual reasoning. We will release OmniVideoBench to foster the development of MLLMs with stronger and more generalizable reasoning capabilities.
  •  

Single Image Reflection Separation via Dual Prior Interaction Transformer

arXiv:2505.12641v3 Announce Type: replace-cross Abstract: Single image reflection separation aims to separate the transmission and reflection layers from a mixed image. Existing methods typically combine general priors from pre-trained models with task-specific priors such as text prompts and reflection detection. However, the transmission prior, as the most direct task-specific prior for the target transmission layer, has not been effectively modeled or fully utilized, limiting performance in complex scenarios. To address this issue, we propose a dual-prior interaction framework based on lightweight transmission prior generation and effective prior fusion. First, we design a Local Linear Correction Network (LLCN) that finetunes pre-trained models based on the physical constraint T=SI+B, where S and B represent pixel-wise and channel-wise scaling and bias transformations. LLCN efficiently generates high-quality transmission priors with minimal parameters. Second, we construct a Dual-Prior Interaction Transformer (DPIT) that employs a dual-stream channel reorganization attention mechanism. By reorganizing features from general and transmission priors for attention computation, DPIT achieves deep fusion of both priors, fully exploiting their complementary information. Experimental results on multiple benchmark datasets demonstrate that the proposed method achieves state-of-the-art performance.
  •  

Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design

arXiv:2602.10016v2 Announce Type: replace-cross Abstract: Deriving predictable scaling laws that govern the relationship between model performance and computational investment is crucial for designing and allocating resources in massive-scale recommendation systems. While such laws are established for large language models, they remain challenging for recommendation systems, especially those processing both user history and context features. We identify poor scaling efficiency as the main barrier to predictable power-law scaling, stemming from inefficient modules with low Model FLOPs Utilization (MFU) and suboptimal resource allocation. We introduce Kunlun, a scalable architecture that systematically improves model efficiency and resource allocation. Our low-level optimizations include Generalized Dot-Product Attention (GDPA), Hierarchical Seed Pooling (HSP), and Sliding Window Attention. Our high-level innovations feature Computation Skip (CompSkip) and Event-level Personalization. These advances increase MFU from 17% to 37% on NVIDIA B200 GPUs and double scaling efficiency over state-of-the-art methods. Kunlun is now deployed in major Meta Ads models, delivering significant production impact.
  •  
❌