❌

Reading view

Suicidal Thoughts and Behaviors Among Chinese Adolescents in Relation to Negative Life Events, Internet Addiction, and Sexual Abuse: Cross-Sectional Study

Background: Increasing suicidal thoughts and behaviors (STB) among adolescents raise social concerns and have a well-recognized association with sexual abuse (SA). However, research regarding the mechanisms explaining the association between SA and STB remains limited. Objective: This study aims to examine the chained mediating effects of negative life events (NLE) and internet addiction (IA) between SA and STB among adolescents in China. Methods: This cross-sectional study used data from the Science Database of the People Mental Health survey conducted between March 2013 and December 2022 by the National Population Health Data Center of the National Research Institute for Family Planning. Through stratified sampling, 20,893 adolescents were recruited from 16 Chinese provinces. After excluding samples with missing relevant variables, 10,664 (55.89%; aged 16-17.9 y; n=5826, 54.63% women) adolescents were included in the final analysis. STB was the outcome variable, with NLE and IA as mediators, all assessed via a questionnaire that was uniformly administered by trained investigators in school settings. The Pearson Ο‡ test was used to analyze the association between SA and STB. Using a combination of multiple linear regression and bootstrap testing, the study constructed a chain mediation model to explore how SA influences STB in adolescents through NLE and IA. Results: The scores for SA, NLE, IA, and STB were 1.330 (SD 1.714), 51.960 (SD 23.822), 34.88 (SD 13.852), and 0.690 (SD 1.396), respectively. Multiple linear regression analysis indicated SA was associated with NLE (Ξ²=2.382, 95% CI 2.112‐2.653;
  •  

Improving Retrieval Augmented Generation for Health Care by Fine-Tuning Clinical Embedding Models: Development and Evaluation Study

Background: Embedding models are critical components of Retrieval Augmented Generation (RAG) systems for retrieving and searching unstructured medical data. However, existing models are predominantly trained on publicly available English datasets, limiting their effectiveness in non-English health care settings. More importantly, these models lack training on real-world clinical documents, leading to inaccurate context retrieval when integrated into RAG systems for health care applications. This gap is particularly pronounced in specialized medical documentation containing domain-specific terminology, abbreviations, and nuanced clinical language. Objective: This retrospective study aimed to develop and validate embedding models specifically trained on real-world clinical documents from multiple medical specialties to improve medical information retrieval (IR) and RAG system performance in both German and English language contexts. Methods: We fine-tuned embedding models, so-called sentence transformers, using the multilingual-e5-large architecture as a foundation. Training data consisted of approximately 11 million question-answer pairs synthetically generated from 400,000 diverse clinical documents from a large German tertiary hospital, spanning 163,840 patients and 282,728 clinical cases between 2018 and 2023. The large language model generated medically relevant questions and corresponding answers for each document. The dataset was additionally pseudonymized and translated into English to aim for broader applicability. Models were evaluated in 2 distinct scenarios: IR using questions with multiple relevant passages, and RAG system performance in both cross-patient and patient-centered contexts. Results: In the IR evaluation, the fine-tuned miracle model achieved a mAP@100 of 0.27, outperforming the multilingual-e5-large baseline (0.14) and state-of-the-art models such as bge-m3 (0.11). In the RAG evaluation, the model demonstrated robust performance comparable with the baseline in the constrained patient-centered scenario (BERTScore F1 0.781 vs 0.778) and showed moderate improvements in the unconstrained cross-patient setting (BLEURT 0.56 vs 0.53). Notably, the model trained on pseudonymized data achieved comparable retrieval performance (mAP@100 0.25) and the highest scores for patient-centered contextual precision (0.93). Performance gains were robust in the German dataset, while the translated English model demonstrated promising results as a proof of concept for cross-lingual transfer. Conclusions: By leveraging a comprehensive real-world dataset spanning multiple medical specialties and using large language models for synthetic question generation, we successfully created and validated domain-specific embedding models. These models can improve medical IR in large-scale search spaces and perform competitively in constrained RAG applications. By publishing the models trained on pseudonymized data, other health care institutions can integrate or adapt these embedding models to their needs. This work establishes a reproducible framework for developing domain-specific clinical embedding models, with the potential to improve data retrieval in medical settings.
  •  

Why this battery company is pivoting to AI

Qichao Hu doesn’t mince words about how he sees the state of the battery industry. β€œAlmost every Western battery company has either died or is going to die. It’s kind of the reality,” he says.

Hu is the CEO of SES AI, a Massachusetts-based battery company. It once had aims of making huge amounts of advanced lithium metal batteries for major industries like electric vehiclesβ€”but now the company is placing its bets on AI materials discovery.

Hu sees the pivot as an essential one. β€œIt’s just not possible for a Western company to build a sustainable business,” he says. The company is still making some batteries, but only for smaller markets like drones rather than those that would require higher volumes, like EVs. The new focus is the company’s battery materials discovery platformβ€”which it can either license to other battery companies or use to develop materials to sell.Β 

Some leading US EV battery companies have folded in recent months, and others, like SES AI, are making dramatic changes in strategy. This shift in who’s building batteries and where they’re doing it could shape the future geopolitics of energy.Β 

The work that would eventually evolve into SES AI began at MIT, where Hu completed his graduate research. His battery work was aimed at applications in oil and gas exploration. The industry uses sensors that go deep underground, where temperatures can top 120 Β°C (about 250 Β°F). The team hoped to develop a battery that could withstand those high temperatures and last longer on a single charge.Β 

The chosen technology was a solid polymer lithium metal battery. These cells use lithium metal for their anode and a polymer for their electrolyte (the material that ions move through in a battery cell). Together, these components can increase the energy density of a cell significantly, relative to the lithium-ion batteries that are common in personal devices and EVs today. (Lithium-ion batteries generally use a graphite material for their anode and a liquid for the electrolyte.)

That solid-state battery technology became the foundation of Solid Energy, a startup Hu founded that spun out from MIT in 2012 and raised its first private investment in 2013.

The team eventually realized that underground oil exploration was a small market, so after several years of operation they began to focus on electric vehicles, which were starting to come into the mainstream. After the team tweaked the chemistry to work better at lower temperatures, the company built its first pilot facility in Massachusetts and eventually another facility in Shanghai.

By 2021, the battery industry was booming, Hu recalls, and EVs were the hottest industry to be in. There was a ton of interest in next-generation battery technology from major automakers at the time, and Solid Energy started developing technology with GM, Hyundai, and Honda.

Larger vehicles, like SUVs and trucks, seemed like a good fit for next-generation batteries, Hu says. Massive vehicles like the ones Americans like to drive would need lighter batteries so they could have a reasonable range without being prohibitively heavy.

The company also shifted its chemistry focus, and in 2022 it announced a battery with a silicon anode rather than a lithium metal one. That shift could help make the battery easier to manufacture.

Since then, growth in the EV market has slowed, at least in the US, partly because of major pullbacks in funding from the Trump administration. EV tax credits for drivers, a key piece of support pushing Americans toward electric options, ended in late 2025. With the market for large electric cars in trouble, Hu says, β€œnow we have to look at every market.”  

The AI materials discovery platform on which it’s pinning many of its hopes is called Molecular Universe. The company seeks not only to provide its software to other battery companies but also to identify new battery materials and either license them or sell them to those companies.

vials of electrolytes inside a machine at the synthesis foundry
COURTESY OF SES AI

The platform has already identified six new electrolyte materials, according to the company. Hu says one is an additive that could help improve the lifetime of batteries with silicon anodes.Β 

One of the challenges with silicon anodes is that they tend to swell a lot during use, which can cause physical damage and prevent efficient charging and discharging. To address the problem, the industry typically uses a material called fluoroethylene carbonate (FEC), which can help form an elastic film on the anode so the battery can still charge effectively. That additive can degrade at high temperatures, though, producing gases that can harm a battery’s lifetime. The SES platform identified a compound that works like FEC but doesn’t release those gases.

The company’s long history and deep battery knowledge could help make its platform a useful tool, Hu says. He sees the actual model as less crucial than SES’s domain expertise and data from years of making and testing batteries.Β 

β€œBy not actually making the physical battery, we’re actually able to scale and then generate revenue faster,” he says.Β 

But some experts are skeptical about the near-term prospects for AI materials discovery to revive the industry. β€œNew materials development, as much as we thought that was what people wanted (and, frankly, it should be what the cell makers want)β€”I don’t know that that seems to be the real linchpin of the battery industry’s progress,” says Kara Rodby, a technical principal at Volta Energy Technologies, a venture capital firm that focuses on the energy storage industry.

Investors are pulling back, and a slowdown in public support is making things difficult for some parts of the battery industry, she adds: β€œI don’t know that the ability to discover any new material is going to unlock anything new for the battery industry at this point in time.”

  •  

Robot-Assisted Therapy for Upper Limb Rehabilitation After Stroke: Umbrella Review

Background: Stroke is a leading cause of long-term upper limb disability, severely impacting patients’ independence and quality of life. Robot-assisted therapy (RAT) has emerged as a promising, high-intensity rehabilitation alternative. However, conclusions from existing systematic reviews on its efficacy are inconsistent and often lack a holistic framework, limiting their use for guiding personalized clinical decisions. Objective: This study aims to systematically synthesize recent evidence on RAT for upper limb rehabilitation after stroke. Guided by the International Classification of Functioning, Disability and Health framework, it moves beyond singular outcomes to provide a multidimensional evaluation across body function, activity, and participation levels. The review aims to provide stratified guidance for clinical decision-making based on patient- and intervention-specific characteristics, thereby supporting evidence-based practice and informing future research. Methods: This study included systematic reviews and meta-analyses published from January 1, 2019, to December 26, 2025, comparing RAT with conventional therapy for upper limb rehabilitation after stroke. Overall, 6 databases, including PubMed, Web of Science, and Embase, were searched. Two reviewers (XZ and LZ) independently performed study selection, data extraction, and quality assessment using the AMSTAR 2 tool. The synthesis integrated outcome measures and subgroup analyses derived from the included studies. Results: This umbrella review included 21 meta-analyses encompassing 535 randomized controlled trials and 27,598 patients across acute, subacute, and chronic stroke stages. According to AMSTAR 2, 17 reviews were high quality, 3 moderate, and 1 critically low. The synthesis demonstrated that RAT was superior in improving upper limb motor function, but no statistically significant advantages were observed in activities of daily living compared to conventional therapy. Subgroup analyses revealed that treatment effects were influenced by stroke stage, upper limb motor impairment level, and robot type. Conclusions: RAT is an effective intervention for improving upper limb motor function after stroke. However, its benefits are primarily observed at the level of body function, with limited evidence for long-term maintenance. The current evidence is constrained by significant outcome heterogeneity and methodological limitations inherent to umbrella reviews. Future research should validate these findings in broader clinical practice, focus on translating functional gains into sustained improvements in daily activities and participation, and include cost-effectiveness evaluations. Trial Registration: PROSPERO CRD42024497183; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024497183
  •  

Systems Biology and Multi-Omics in Asthma and COPD: A Systematic Review of Computational Approaches (2010-2024)

J Asthma Allergy. 2026 Mar 19;19:575312. doi: 10.2147/JAA.S575312. eCollection 2026.

ABSTRACT

Systems biology approaches have contributed to advancing our understanding of complex respiratory diseases including asthma and chronic obstructive pulmonary disease (COPD). This systematic review evaluates the application of systems biology methodologies in respiratory medicine, focusing on multi-omics data integration and computational techniques for biomarker discovery and mechanistic understanding. Following PRISMA 2020 guidelines, we conducted a comprehensive literature search across Web of Science and Scopus databases, identifying 117 peer-reviewed documents published from 2010 to 2024. The review methodology employed bibliometric analysis combined with qualitative synthesis of included studies. Results demonstrate steady growth in systems biology applications for asthma and COPD research, with publication rates increasing by approximately 0.5 articles per year (R2 = 0.73, p < 0.001). Bibliometric analysis identified five major research clusters: systems biology as a foundational methodological framework (Basic Theme), COPD-focused research as the most developed area (Motor Theme), gene expression analysis, disease classification approaches, and specialized lung disease investigations (Niche Theme). Multi-omics integration studies achieved 82-91% accuracy in disease classification tasks, with transcriptomics-based asthma endotyping validated in over 1500 patients across multiple cohorts. Network analysis approaches identified hub genes (IL-6, TNF-Ξ±, MMP9) replicated across three independent studies. Machine learning applications demonstrated 80-90% accuracy for diagnostic and prognostic tasks, though external validation remains limited, with only 15% of reviewed studies including independent validation cohorts. Significant challenges persist in data integration, computational reproducibility, and clinical translation. Most studies employed modest sample sizes (median n=89), and population diversity was limited, with 89% conducted in European-ancestry populations. This review provides a comprehensive assessment of systems biology progress in respiratory medicine, identifies methodological gaps, and highlights the need for standardized protocols, larger collaborative studies, and rigorous external validation to advance clinical implementation of systems biology findings in asthma and COPD management.

PMID:41878747 | PMC:PMC13007689 | DOI:10.2147/JAA.S575312

  •  

Systems Biology and Multi-Omics in Asthma and COPD: A Systematic Review of Computational Approaches (2010-2024)

J Asthma Allergy. 2026 Mar 19;19:575312. doi: 10.2147/JAA.S575312. eCollection 2026.

ABSTRACT

Systems biology approaches have contributed to advancing our understanding of complex respiratory diseases including asthma and chronic obstructive pulmonary disease (COPD). This systematic review evaluates the application of systems biology methodologies in respiratory medicine, focusing on multi-omics data integration and computational techniques for biomarker discovery and mechanistic understanding. Following PRISMA 2020 guidelines, we conducted a comprehensive literature search across Web of Science and Scopus databases, identifying 117 peer-reviewed documents published from 2010 to 2024. The review methodology employed bibliometric analysis combined with qualitative synthesis of included studies. Results demonstrate steady growth in systems biology applications for asthma and COPD research, with publication rates increasing by approximately 0.5 articles per year (R2 = 0.73, p < 0.001). Bibliometric analysis identified five major research clusters: systems biology as a foundational methodological framework (Basic Theme), COPD-focused research as the most developed area (Motor Theme), gene expression analysis, disease classification approaches, and specialized lung disease investigations (Niche Theme). Multi-omics integration studies achieved 82-91% accuracy in disease classification tasks, with transcriptomics-based asthma endotyping validated in over 1500 patients across multiple cohorts. Network analysis approaches identified hub genes (IL-6, TNF-Ξ±, MMP9) replicated across three independent studies. Machine learning applications demonstrated 80-90% accuracy for diagnostic and prognostic tasks, though external validation remains limited, with only 15% of reviewed studies including independent validation cohorts. Significant challenges persist in data integration, computational reproducibility, and clinical translation. Most studies employed modest sample sizes (median n=89), and population diversity was limited, with 89% conducted in European-ancestry populations. This review provides a comprehensive assessment of systems biology progress in respiratory medicine, identifies methodological gaps, and highlights the need for standardized protocols, larger collaborative studies, and rigorous external validation to advance clinical implementation of systems biology findings in asthma and COPD management.

PMID:41878747 | PMC:PMC13007689 | DOI:10.2147/JAA.S575312

  •  

The Efficiency Attenuation Phenomenon: A Computational Challenge to the Language of Thought Hypothesis

arXiv:2603.22312v1 Announce Type: new Abstract: This paper computationally investigates whether thought requires a language-like format, as posited by the Language of Thought (LoT) hypothesis. We introduce the ``AI Private Language'' thought experiment: if two artificial agents develop an efficient, inscrutable communication protocol via multi-agent reinforcement learning (MARL), and their performance declines when forced to use a human-comprehensible language, this Efficiency Attenuation Phenomenon (EAP) challenges the LoT. We formalize this in a cooperative navigation task under partial observability. Results show that agents with an emergent protocol achieve 50.5\% higher efficiency than those using a pre-defined, human-like symbolic protocol, confirming the EAP. This suggests optimal collaborative cognition in these systems is not mediated by symbolic structures but is naturally coupled with sub-symbolic computations. The work bridges philosophy, cognitive science, and AI, arguing for pluralism in cognitive architectures and highlighting implications for AI ethics.
  •  

Intelligence Inertia: Physical Principles and Applications

arXiv:2603.22347v1 Announce Type: new Abstract: While Landauer's principle establishes the fundamental thermodynamic floor for information erasure and Fisher Information provides a metric for local curvature in parameter space, these classical frameworks function effectively only as approximations within regimes of sparse rule-constraints. They fail to explain the super-linear, and often explosive, computational and energy costs incurred when maintaining symbolic interpretability during the reconfiguration of advanced intelligent systems. This paper introduces the property of intelligence inertia and its underlying physical principles as foundational characteristics for quantifying the computational weight of intelligence. We demonstrate that this phenomenon is not merely an empirical observation but originates from the fundamental non-commutativity between rules and states, a root cause we have formally organized into a rigorous mathematical framework. By analyzing the growing discrepancy between actual adaptation costs and static information-theoretic estimates, we derive a non-linear cost formula that mirrors the Lorentz factor, characterizing a relativistic J-shaped inflation curve -- a "computational wall" that static models are blind to. The validity of these physical principles is examined through a trilogy of decisive experiments: (1) a comparative adjudication of this J-curve inflation against classical Fisher Information models, (2) a geometric analysis of the "Zig-Zag" trajectory of neural architecture evolution, and (3) the implementation of an inertia-aware scheduler wrapper that optimizes the training of deep networks by respecting the agent's physical resistance to change. Our results suggest a unified physical description for the cost of structural adaptation, offering a first-principle explanation for the computational and interpretability-maintenance overhead in intelligent agents.
  •  

From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents

arXiv:2603.22386v1 Announce Type: new Abstract: Large language model (LLM)-based systems are becoming increasingly popular for solving tasks by constructing executable workflows that interleave LLM calls, information retrieval, tool use, code execution, memory updates, and verification. This survey reviews recent methods for designing and optimizing such workflows, which we treat as agentic computation graphs (ACGs). We organize the literature based on when workflow structure is determined, where structure refers to which components or agents are present, how they depend on each other, and how information flows between them. This lens distinguishes static methods, which fix a reusable workflow scaffold before deployment, from dynamic methods, which select, generate, or revise the workflow for a particular run before or during execution. We further organize prior work along three dimensions: when structure is determined, what part of the workflow is optimized, and which evaluation signals guide optimization (e.g., task metrics, verifier signals, preferences, or trace-derived feedback). We also distinguish reusable workflow templates, run-specific realized graphs, and execution traces, separating reusable design choices from the structures actually deployed in a given run and from realized runtime behavior. Finally, we outline a structure-aware evaluation perspective that complements downstream task metrics with graph-level properties, execution cost, robustness, and structural variation across inputs. Our goal is to provide a clear vocabulary, a unified framework for positioning new methods, a more comparable view of existing body of literature, and a more reproducible evaluation standard for future work in workflow optimizations for LLM agents.
  •  

Computational Arbitrage in AI Model Markets

arXiv:2603.22404v1 Announce Type: new Abstract: Consider a market of competing model providers selling query access to models with varying costs and capabilities. Customers submit problem instances and are willing to pay up to a budget for a verifiable solution. An arbitrageur efficiently allocates inference budget across providers to undercut the market, thus creating a competitive offering with no model-development risk. In this work, we initiate the study of arbitrage in AI model markets, empirically demonstrating the viability of arbitrage and illustrating its economic consequences. We conduct an in-depth case study of SWE-bench GitHub issue resolution using two representative models, GPT-5 mini and DeepSeek v3.2. In this verifiable domain, simple arbitrage strategies generate net profit margins of up to 40%. Robust arbitrage strategies that generalize across different domains remain profitable. Distillation further creates strong arbitrage opportunities, potentially at the expense of the teacher model's revenue. Multiple competing arbitrageurs drive down consumer prices, reducing the marginal revenue of model providers. At the same time, arbitrage reduces market segmentation and facilitates market entry for smaller model providers by enabling earlier revenue capture. Our results suggest that arbitrage can be a powerful force in AI model markets with implications for model development, distillation, and deployment.
  •  

Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context Length

arXiv:2603.22608v1 Announce Type: new Abstract: Users often rely on Large Language Models (LLMs) for processing multiple documents or performing analysis over a number of instances. For example, analysing the overall sentiment of a number of movie reviews requires an LLM to process the sentiment of each review individually in order to provide a final aggregated answer. While LLM performance on such individual tasks is generally high, there has been little research on how LLMs perform when dealing with multi-instance inputs. In this paper, we perform a comprehensive evaluation of the multi-instance processing (MIP) ability of LLMs for tasks in which they excel individually. The results show that all LLMs follow a pattern of slight performance degradation for small numbers of instances (approximately 20-100), followed by a performance collapse on larger instance counts. Crucially, our analysis shows that while context length is associated with this degradation, the number of instances has a stronger effect on the final results. This finding suggests that when optimising LLM performance for MIP, attention should be paid to both context length and, in particular, instance count.
  •  

Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning

arXiv:2603.22619v1 Announce Type: new Abstract: LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify such issues, yet fail to reflect this in standard generative responses. This reveals a fundamental know-act gap between discriminative recognition and generative behavior. Prior work largely characterizes this issue in narrow settings, such as math word problems or question answering, with limited focus on how to integrate these two modes. In this work, we present a comprehensive analysis using FaultyScience, a newly constructed large-scale, cross-disciplinary benchmark of faulty scientific questions. We show that the gap is pervasive and stems from token-level autoregression, which entangles task selection (validate vs. answer) with content generation, preventing discriminative knowledge from being utilized. To address this, we propose DeIllusionLLM, a task-level autoregressive framework that explicitly models this decision. Through self-distillation, the model unifies discriminative judgment and generative reasoning within a single backbone. Empirically, DeIllusionLLM substantially reduces answer-despite-error failures under natural prompting while maintaining general reasoning performance, demonstrating that self-distillation is an effective and scalable solution for bridging the discriminative-generative know-act gap
  •  

Graph-Aware Late Chunking for Retrieval-Augmented Generation in Biomedical Literature

arXiv:2603.22633v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems for biomedical literature are typically evaluated using ranking metrics like Mean Reciprocal Rank (MRR), which measure how well the system identifies the single most relevant chunk. We argue that for full-text scientific documents, this paradigm is incomplete: it rewards retrieval precision while ignoring retrieval breadth -- the ability to surface evidence from across a document's structural sections. We propose GraLC-RAG, a framework that unifies late chunking with graph-aware structural intelligence, introducing structure-aware chunk boundary detection, UMLS knowledge graph infusion, and graph-guided hybrid retrieval. We evaluate six strategies on 2,359 IMRaD-filtered PubMed Central articles using 2,033 cross-section questions and two metric families: standard ranking metrics (MRR, Recall@k) and structural coverage metrics (SecCov@k, CS Recall). Our results expose a sharp divergence: content-similarity methods achieve the highest MRR (0.517) but always retrieve from a single section, while structure-aware methods retrieve from up to 15.6x more sections. Generation experiments show that KG-infused retrieval narrows the answer-quality gap to delta-F1 = 0.009 while maintaining 4.6x section diversity. These findings demonstrate that standard metrics systematically undervalue structural retrieval and that closing the multi-section synthesis gap is a key open problem for biomedical RAG.
  •  

Benchmarking Multi-Agent LLM Architectures for Financial Document Processing: A Comparative Study of Orchestration Patterns, Cost-Accuracy Tradeoffs and Production Scaling Strategies

arXiv:2603.22651v1 Announce Type: new Abstract: The adoption of large language models (LLMs) for structured information extraction from financial documents has accelerated rapidly, yet production deployments face fundamental architectural decisions with limited empirical guidance. We present a systematic benchmark comparing four multi-agent orchestration architectures: sequential pipeline, parallel fan-out with merge, hierarchical supervisor-worker and reflexive self-correcting loop. These are evaluated across five frontier and open-weight LLMs on a corpus of 10,000 SEC filings (10-K, 10-Q and 8-K forms). Our evaluation spans 25 extraction field types covering governance structures, executive compensation and financial metrics, measured along five axes: field-level F1, document-level accuracy, end-to-end latency, cost per document and token efficiency. We find that reflexive architectures achieve the highest field-level F1 (0.943) but at 2.3x the cost of sequential baselines, while hierarchical architectures occupy the most favorable position on the cost-accuracy Pareto frontier (F1 0.921 at 1.4x cost). We further present ablation studies on semantic caching, model routing and adaptive retry strategies, demonstrating that hybrid configurations can recover 89\% of the reflexive architecture's accuracy gains at only 1.15x baseline cost. Our scaling analysis from 1K to 100K documents per day reveals non-obvious throughput-accuracy degradation curves that inform capacity planning. These findings provide actionable guidance for practitioners deploying multi-agent LLM systems in regulated financial environments.
  •  

Detecting outliers of pursuit eye movements: a preliminary analysis of autism spectrum disorder

arXiv:2603.22705v2 Announce Type: new Abstract: Background: Autism spectrum disorder (ASD) is characterized by significant clinical and biological heterogeneity. Conventional group-mean analyses of eye movements often mask individual atypicalities, potentially overlooking critical pathological signatures. This study aimed to identify idiosyncratic oculomotor patterns in ASD using an "outlier analysis" of smooth pursuit eye movement (SPEM). Methods: We recorded SPEM during a slow Lissajous pursuit task in 18 adults with ASD and 39 typically developed (TD) individuals. To quantify individual deviations, we derived an "outlier score" based on the Mahalanobis distance. This score was calculated from a feature vector, optimized via Principal Component Analysis (PCA), comprising the temporal lag ($\Delta$t) and the spatial deviation ($\Delta$s). An outlier was statistically defined as a score exceeding $\sqrt{10}$ (approximately 3.16$\sigma$) relative to the TD normative distribution. Results: While the TD group exhibited a low outlier rate of 5.1%, the ASD group demonstrated a significantly higher prevalence of 38.9% (7/18) (binomial P = 0.0034). Furthermore, the mean outlier score was significantly elevated in the ASD group (3.00 $\pm$ 2.62) compared to the TD group (1.52 $\pm$ 0.80; P = 0.002). Notably, these extreme deviations were captured even when conventional mean-based comparisons showed limited sensitivity. Conclusions: Our outlier analysis successfully visualized the high degree of idiosyncratic atypicality in ASD oculomotor control. By shifting the focus from group averages to individual deviations, this approach provides a sensitive metric for capturing the inherent heterogeneity of ASD, offering a potential baseline for identifying clinical subtypes.
  •  

Beyond Binary Correctness: Scaling Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

arXiv:2603.22744v1 Announce Type: new Abstract: Large language models excel on objectively verifiable tasks such as math and programming, where evaluation reduces to unit tests or a single correct answer. In contrast, real-world enterprise work is often subjective and context-dependent: success hinges on organizational goals, user intent, and the quality of intermediate artifacts produced across long, multi-tool workflows. We introduce LH-Bench, a three-pillar evaluation design that moves beyond binary correctness to score autonomous, long-horizon execution on subjective enterprise tasks. The pillars are: (i) expert-grounded rubrics that give LLM judges the domain context needed to score subjective work, (ii) curated ground-truth artifacts that enable stepwise reward signals (e.g., chapter-level annotation for content tasks), and (iii) pairwise human preference evaluation for convergent validation. We show that domain-authored rubrics provide substantially more reliable evaluation signals than LLM-authored rubrics (kappa = 0.60 vs. 0.46), and that human preference judgments confirm the same top-tier separation (p
  •  

AgriPestDatabase-v1.0: A Structured Insect Dataset for Training Agricultural Large Language Model

arXiv:2603.22777v1 Announce Type: new Abstract: Agricultural pest management increasingly relies on timely and accurate access to expert knowledge, yet high quality labeled data and continuous expert support remain limited, particularly for farmers operating in rural regions with unstable/no internet connectivity. At the same time, the rapid growth of AI and LLMs has created new opportunities to deliver practical decision support tools directly to end users in agriculture through compact and deployable systems. This work addresses (i) generating a structured insect information dataset, and (ii) adapting a lightweight LLM model ($\leq$ 7B) by fine tuning it for edge device uses in agricultural pest management. The textual data collection was done by reviewing and collecting information from available pest databases and published manuscripts on nine selected pest species. These structured reports were then reviewed and validated by a domain expert. From these reports, we constructed Q/A pairs to support model training and evaluation. A LoRA-based fine-tuning approach was applied to multiple lightweight LLMs and evaluated. Initial evaluation shows that Mistral 7B achieves an 88.9\% pass rate on the domain-specific Q/A task, substantially outperforming Qwen 2.5 7B (63.9\%), and LLaMA 3.1 8B (58.7\%). Notably, Mistral demonstrates higher semantic alignment (embedding similarity: 0.865) despite lower lexical overlap (BLEU: 0.097), indicating that semantic understanding and robust reasoning are more predictive of task success than surface-level conformity in specialized domains. By combining expert organized data, well-structured Q/A pairs, semantic quality control, and efficient model adaptation, this work contributes towards providing support for farmer facing agricultural decision support tools and demonstrates the feasibility of deploying compact, high-performing language models for practical field-level pest management guidance.
  •  

Reliable Classroom AI via Neuro-Symbolic Multimodal Reasoning

arXiv:2603.22793v1 Announce Type: new Abstract: Classroom AI is rapidly expanding from low-level perception toward higher-level judgments about engagement, confusion, collaboration, and instructional quality. Yet classrooms are among the hardest real-world settings for multimodal vision: they are multi-party, noisy, privacy-sensitive, pedagogically diverse, and often multilingual. In this paper, we argue that classroom AI should be treated as a critical domain, where raw predictive accuracy is insufficient unless predictions are accompanied by verifiable evidence, calibrated uncertainty, and explicit deployment guardrails. We introduce NSCR, a neuro-symbolic framework that decomposes classroom analytics into four layers: perceptual grounding, symbolic abstraction, executable reasoning, and governance. NSCR adapts recent ideas from symbolic fact extraction and verifiable code generation to multimodal educational settings, enabling classroom observations from video, audio, ASR, and contextual metadata to be converted into typed facts and then composed by executable rules, programs, and policy constraints. Beyond the system design, we contribute a benchmark and evaluation protocol organized around five tasks: classroom state inference, discourse-grounded event linking, temporal early warning, collaboration analysis, and multilingual classroom reasoning. We further specify reliability metrics centered on abstention, calibration, robustness, construct alignment, and human usefulness. The paper does not report new empirical results; its contribution is a concrete framework and evaluation agenda intended to support more interpretable, privacy-aware, and pedagogically grounded multimodal AI for classrooms.
  •  
❌