❌

Normal view

SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning

arXiv:2609.12599v1 Announce Type: cross Abstract: Multi-task learning (MTL) requires navigating unavoidable trade-offs among competing objectives. This paradigm is frequently formulated as multi-objective optimization (MOO), where the scalarization is favored to reduce an MOO problem to a single objective. We empirically find that existing merit-function-based scalarization approaches are sensitive to the relative scales of different objectives in practical MTL, where task losses commonly differ by orders of magnitude. The optimization process often favors objectives with larger scales even though the underlying Pareto optimal solutions remains invariant to rescaling (i.e., multiplying an objective by a positive constant). To address this issue, we propose Scale-Invariant Merit-function-based Scalarization (SIMS) for MTL. Specifically, SIMS adopts a transformation-induced merit function to convert the MOO problem of MTL to a single objective that renders optimization invariant to the magnitudes of losses. Theoretically, we prove that the requirement for scale invariance uniquely determines this transformation to be logarithmic. We further show that this general transformation-induced merit function preserves weak Pareto optimality and admits a smooth surrogate with controllable approximation error. Extensive experiments on representative multi-task benchmarks demonstrate that SIMS consistently outperforms existing scalarization methods and achieves state-of-the-art performance.

Demystifying the Privacy-Utility Trade-off in LLM Interactions

arXiv:2609.10992v2 Announce Type: replace Abstract: The integration of Large Language Models into daily tasks relies on context-rich instructions, inevitably exposing sensitive user information. Current privacy-preserving methods typically employ context-agnostic static rules, causing severe utility degradation. However, the specific mechanisms governing how sanitization impacts downstream performance remain largely underexplored. To address this, we conduct a systematic analysis to deconstruct the privacy-utility trade-off, uncovering three underlying mechanisms: (1) Context-Dependent Utility, which first establishes when to sanitize by revealing that data value shifts from critical constraints to dispensable noise based on user intent; (2) Strategic Adaptation, which subsequently determines how to sanitize by dictating that the choice between removal and replacement depends on the task's reliance on factual integrity versus structural coherence; and (3) Combinatorial Interplay, which finally extends the protection scope by demonstrating that attributes form a semantic web of synergistic dependencies or antagonistic redundancies. Guided by these insights, we introduce an intent-driven local protection framework. By distilling a lightweight model Veilmind-4B to drive a dynamic extraction-sanitization-restoration pipeline, our approach reaches a low-leakage privacy point while preserving substantially higher response utility than existing privacy-oriented baselines, advancing the privacy-utility trade-off toward the Pareto frontier.

Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges

arXiv:2609.01210v2 Announce Type: replace-cross Abstract: Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is critical in Chinese harmful-content evaluation, where linguistic variation and adversarial transformations can obscure risky intent. We introduce C-SafeQA, a policy-grounded benchmark for response-level Chinese safety evaluation. It comprises 538 base queries and 8,877 adversarial queries answered by four full-model LLM deployments, yielding 37,660 query-response records labeled safe, unsafe, or disputed. Reference labels are generated through agreement-aware multi-model adjudication and blind audits of stratified subsets by three safety experts. C-SafeQA supports both evaluation of target-model safety and auditing of seven automated safety judges against shared reference labels. Unsafe-response rates range from 0.93% to 3.35% on base queries and from 11.68% to 30.05% on adversarial queries. On the adversarial subset, judges show substantial trade-offs between unsafe-response recall and risk-query-conditioned safe-response false positive rate, and no judge dominates all metrics. Both acrostic transformations reduce unsafe recall for all seven judges, revealing mechanism-specific evaluator weaknesses. Dataset records, metadata, verification code, and judge scripts are publicly released to support recomputation, while benchmark construction, target-response generation, and private adjudication remain outside the release boundary.

BRD4 Inhibition Mitigates Acute and Chronic Corneal Injury Following Topical Nitrogen Mustard Exposure

Lu and colleagues identify BRD4 as a central epigenetic driver of vesicant-induced corneal injury. Using reproducible mouse and rabbit models, they show that short-term topical BRD4 inhibition suppresses acute inflammation and provides durable protection of corneal clarity, stromal organization, endothelial integrity, and neovascularization, supporting translational therapy for chemical eye injuries.

A helicase-fused Cas9 improves large-size fragment knockin

By fusing MCM5, a subunit of the eukaryotic MCM2–7 helicase complex, to the N terminus of spCas9 (MCCas), the MCCas fusion protein enhances large-size fragment knockin via homologous recombination, reduces insertions and deletions (indels), and enables efficient large-size fragment insertions in human cells and rabbit embryos.

EBV strain interacts with host HLA to drive nasopharyngeal carcinoma risk

Nature, Published online: 15 April 2026; doi:10.1038/s41586-026-10416-8

A genome-to-genome association study identifies host and viral risk factors that interact to drive nasopharyngeal carcinoma endemicity in southern China.

Effect of the Maxing Huoqiao granule on nonsevere community-acquired pneumonia: A multicenter, double-blind, placebo-controlled randomized trial

Pharmacol Res. 2026 Apr 9:108186. doi: 10.1016/j.phrs.2026.108186. Online ahead of print.

ABSTRACT

Community-acquired pneumonia (CAP) remains a major global public health challenge with substantial morbidity and mortality. Although preclinical studies suggest that Maxing Huoqiao (MXHQ) granule may have therapeutic potential for pneumonia, high-quality clinical evidence is still limited. We conducted a multicenter, double-blind, randomized, placebo-controlled trial at two tertiary hospitals in China to evaluate the clinical efficacy of MXHQ as adjunctive therapy and to explore its potential mechanisms in adults with nonsevere CAP receiving standard moxifloxacin treatment. A total of 96 patients were enrolled and randomized (1:1:1) to receive standard-dose MXHQ, low-dose MXHQ, or placebo in addition to moxifloxacin for 7 days, with a 14-day follow-up. The primary endpoint was clinical cure, defined as composite recovery of major respiratory symptoms, lung rales, and fever; secondary endpoints included symptom relief, radiographic improvement, and safety. Compared with placebo, standard-dose MXHQ was associated with a higher day-14 clinical cure rate (30.78% vs. 68.97%; RR = 0.45, 95% CI = 0.24-0.83; P < 0.01). Furthermore, the standard-dose intervention was correlated with a shorter time to relief and recovery of cough and sputum (P < 0.05), as well as improvements in symptom scores (P < 0.05) and promoting lesion absorption on chest CT (P < 0.05). Low-dose MXHQ showed no significant clinical benefit, whereas safety profiles were comparable across all groups. Transcriptomic analyses of peripheral blood mononuclear cells, complemented by a Streptococcus pneumonia animal model, indicated that the clinical benefits of MXHQ are linked to the modulation of inflammation and innate immunity. These omics and in vivo observations suggest a potential mechanism underlying the protective effects of MXHQ against inflammatory injury and promotion of tissue repair, involving the regulation of anti-inflammatory mediators and tissue repair-related factors. (Chictr.org.cn, ID Number: ChiCTR2400082095).

PMID:41966499 | DOI:10.1016/j.phrs.2026.108186

AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning

arXiv:2602.13685v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) excel at perception but struggle with complex reasoning requiring precise acoustic measurements. While external tools can extract fine-grained features like exact tempo or pitch, effective integration remains challenging: naively using all tools causes information overload, while prompt-based selection fails to assess context-dependent utility. To address this, we propose AuTAgent (Audio Tool Agent), a reinforcement learning framework that learns when and which tools to invoke. By employing a sparse-feedback training strategy with a novel Differential Reward mechanism, the agent learns to filter out irrelevant tools and invokes external assistance only when it yields a net performance gain over the base model. Experimental results confirm that AuTAgent complements the representation bottleneck of LALMs by providing verifiable acoustic evidence. It improves accuracy by 4.20% / 6.20% and 9.80% / 8.00% for open-source and closed-source backbones on the MMAU Test-mini and the MMAR benchmarks, respectively. In addition, further experiments demonstrate exceptional transferability. We highlight the complementary role of external tools in augmenting audio model reasoning.
❌