❌

Normal view

REFA: Reference Free Alignment for multi-preference optimization

arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic content. To address this, we introduce REFA, a new alignment framework that proposes probabilistic control on a structural token that controls termination. Our core innovation is a new class of regularizers that operate directly on the probability of the End-of-Sequence (EOS) token, a previously unexploited control lever. This token-level intervention provides a principled solution to the URSLA shortcut, ensuring genuine quality improvements. Furthermore, it unlocks a versatile mechanism for managing the alignment-efficiency tradeoff, enabling practitioners to fine-tune models that adhere to specific token budgets. Empirically, REFA achieves a 60.29% win rate and a 52.17% length-controlled win rate on AlpacaEval2 with Llama-3-8B-Instruct, demonstrating the power of our token-level control paradigm.

AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents

arXiv:2510.21031v1 Announce Type: cross Abstract: The emergence of foundation models (FMs) has enabled the development of highly capable and autonomous agents, unlocking new application opportunities across a wide range of domains. Evaluating the architecture of agents is particularly important as the architectural decisions significantly impact the quality attributes of agents given their unique characteristics, including compound architecture, autonomous and non-deterministic behaviour, and continuous evolution. However, these traditional methods fall short in addressing the evaluation needs of agent architecture due to the unique characteristics of these agents. Therefore, in this paper, we present AgentArcEval, a novel agent architecture evaluation method designed specially to address the complexities of FM-based agent architecture and its evaluation. Moreover, we present a catalogue of agent-specific general scenarios, which serves as a guide for generating concrete scenarios to design and evaluate the agent architecture. We demonstrate the usefulness of AgentArcEval and the catalogue through a case study on the architecture evaluation of a real-world tax copilot, named Luna.

Preclinical application of a CD155 targeting chimeric antigen receptor T cell therapy for digestive system cancers

Oncogene, Published online: 01 March 2025; doi:10.1038/s41388-025-03322-2

Preclinical application of a CD155 targeting chimeric antigen receptor T cell therapy for digestive system cancers

Multiscale drug screening for cardiac fibrosis identifies MD2 as a therapeutic target

A multiscale drug discovery platform integrating human induced pluripotent stem cells, 3D-engineered heart tissues, and animal models identifies artesunate as a safe and potent antifibrotic compound.

Label-free detection and profiling of individual solution-phase molecules

Nature, Published online: 08 May 2024; doi:10.1038/s41586-024-07370-8

Enhanced light–molecule interactions in high-finesse fibre-based Fabry–Pérot microcavities are used to detect and profile individual unlabelled solution-phase biomolecules, leading to potential applications in the life and chemical sciences.
❌