❌

Reading view

From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Models

arXiv:2511.10788v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have made reasoning a central benchmark for evaluating intelligence. While prior surveys focus on efficiency by examining how to shorten reasoning chains or reduce computation, this view overlooks a fundamental challenge: current LLMs apply uniform reasoning strategies regardless of task complexity, generating long traces for trivial problems while failing to extend reasoning for difficult tasks. This survey reframes reasoning through the lens of {adaptivity}: the capability to allocate reasoning effort based on input characteristics such as difficulty and uncertainty. We make three contributions. First, we formalize deductive, inductive, and abductive reasoning within the LLM context, connecting these classical cognitive paradigms with their algorithmic realizations. Second, we formalize adaptive reasoning as a control-augmented policy optimization problem balancing task performance with computational cost, distinguishing learned policies from inference-time control mechanisms. Third, we propose a systematic taxonomy organizing existing methods into training-based approaches that internalize adaptivity through reinforcement learning, supervised fine-tuning, and learned controllers, and training-free approaches that achieve adaptivity through prompt conditioning, feedback-driven halting, and modular composition. This framework clarifies how different mechanisms realize adaptive reasoning in practice and enables systematic comparison across diverse strategies. We conclude by identifying open challenges in self-evaluation, meta-reasoning, and human-aligned reasoning control.
  •  

Latent Principle Discovery for Language Model Self-Improvement

arXiv:2505.16927v2 Announce Type: replace-cross Abstract: When language model (LM) users aim to improve the quality of its generations, it is crucial to specify concrete behavioral attributes that the model should strive to reflect. However, curating such principles across many domains, even non-exhaustively, requires a labor-intensive annotation process. To automate this process, we propose eliciting these latent attributes that guide model reasoning toward human-preferred responses by explicitly modeling them in a self-correction setting. Our approach mines new principles from the LM itself and compresses the discovered elements to an interpretable set via clustering. Specifically, we employ a form of posterior-regularized Monte Carlo Expectation-Maximization to both identify a condensed set of the most effective latent principles and teach the LM to strategically invoke them in order to intrinsically refine its responses. We demonstrate that bootstrapping our algorithm over multiple iterations enables smaller language models (7-8B parameters) to self-improve, achieving +8-10% in AlpacaEval win-rate, an average of +0.3 on MT-Bench, and +19-23% in principle-following win-rate on IFEval. We also show that clustering the principles yields interpretable and diverse model-generated constitutions while retaining model performance. The gains that our method achieves highlight the potential of automated, principle-driven post-training recipes toward continual self-improvement.
  •  

An Adaptive Multi Agent Bitcoin Trading System

arXiv:2510.08068v2 Announce Type: replace-cross Abstract: This paper presents a Multi Agent Bitcoin Trading system that utilizes Large Language Models (LLMs) for alpha generation and portfolio management in the cryptocurrencies market. Unlike equities, cryptocurrencies exhibit extreme volatility and are heavily influenced by rapidly shifting market sentiments and regulatory announcements, making them difficult to model using static regression models or neural networks trained solely on historical data. The proposed framework overcomes this by structuring LLMs into specialised agents for technical analysis, sentiment evaluation, decision-making, and performance reflection. The agents improve over time via a novel verbal feedback mechanism where a Reflect agent provides daily and weekly natural-language critiques of trading decisions. These textual evaluations are then injected into future prompts of the agents, allowing them to adjust allocation logic without weight updates or finetuning. Back-testing on Bitcoin price data from July 2024 to April 2025 shows consistent outperformance across market regimes: the Quantitative agent delivered over 30\% higher returns in bullish phases and 15\% overall gains versus buy-and-hold, while the sentiment-driven agent turned sideways markets from a small loss into a gain of over 100\%. Adding weekly feedback further improved total performance by 31\% and reduced bearish losses by 10\%. The results demonstrate that verbal feedback represents a new, scalable, and low-cost approach of tuning LLMs for financial goals.
  •  

Digital Lifestyle Interventions to Support Healthy Gestational Weight Gain: Scoping Review

Background: Digital lifestyle interventions hold promise in supporting healthy gestational weight gain (GWG) during pregnancy. However, clarity on their key design and implementation features remains limited. The prevalence of excessive GWG and its associated maternal and infant health risks makes understanding the landscape of digital intervention characteristics critical. Objective: This scoping review aimed to map current literature on digital lifestyle interventions designed to promote healthy GWG and to identify intervention characteristics, including behavior change techniques (BCTs), employed across these interventions, with particular attention to patterns in design and implementation features across studies reporting positive outcomes. Methods: Following PRISMA-ScR guidelines, we systematically searched PubMed, Embase, Cochrane, and Web of Science for peer-reviewed studies published between 2014 and 2024. Studies were included if they described interventions with at least one digital component targeting GWG. Studies on high-risk pregnancies, non-human subjects, protocols without results, abstracts, gray literature, and non-English publications were excluded. Data extraction covered study characteristics, theoretical frameworks, timing, duration, frequency, delivery modes, and BCTs applied. The landscape of intervention characteristics was mapped, including descriptive analysis of features that appeared across different study outcomes. Results: A total of 44 studies met inclusion criteria: 23 primary data articles (pilot studies, randomized controlled trials, etc) and 21 secondary data articles (meta-analyses, systematic reviews, etc). Primary studies showed that interventions were more likely to achieve intended outcomes when they started earlier, lasted longer and combined digital and in-person components. Five BCTs were commonly present across interventions achieving positive outcomes: Goal setting (outcome) (71%), Discrepancy between current behavior and goal (43%), Self-monitoring of behavior (86%), Social support (unspecified) (71%), and Credible source (71%). Secondary studies supported these findings, identifying several helpful features: starting before mid-pregnancy, long duration with high intensity, in-person contact, and BCTs related to goal setting, action planning, feedback on and monitoring of behavior. However, primary studies showed gaps in reporting practices, with many details lacking about design and implementation features, such as BCTs. This converged with secondary studies reporting insufficient detail in reviewed primary literature, limiting interpretation and replication potential. Conclusions: This scoping review maps digital interventions for GWG and identifies key patterns in intervention design and implementation. Evidence suggests that interventions may be more promising when combining digital delivery with in-person components and incorporating BCTs related to goal setting, self-monitoring, and social support. This review provides comprehensive mapping of BCT usage and other intervention features highlighting approaches associated with positive outcomes. However, significant gaps in reporting practices limit evidence synthesis. The findings can inform design of digital interventions for managing GWG by identifying potentially successful design and implementation features. Future research should prioritize standardized reporting practices and evaluate interventions in underserved populations, including healthcare desert communities, to enhance the evidence base.
  •  
❌