❌

Normal view

ConSensus: Multi-Agent Collaboration for Multimodal Sensing

arXiv:2601.06453v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However, accurately interpreting heterogeneous multimodal sensor data remains a fundamental challenge. We show that a single monolithic LLM often fails to reason coherently across modalities, leading to incomplete interpretations and prior-knowledge bias. We introduce ConSensus, a training-free multi-agent collaboration framework that decomposes multimodal sensing tasks into specialized, modality-aware agents. To aggregate agent-level interpretations, we propose a hybrid fusion mechanism that balances semantic aggregation, which enables cross-modal reasoning and contextual understanding, with statistical consensus, which provides robustness through agreement across modalities. While each approach has complementary failure modes, their combination enables reliable inference under sensor noise and missing data. We evaluate ConSensus on five diverse multimodal sensing benchmarks, demonstrating an average accuracy improvement of 7.1% over the single-agent baseline. Furthermore, ConSensus matches or exceeds the performance of iterative multi-agent debate methods while achieving a 12.7 times reduction in average fusion token cost through a single-round hybrid fusion protocol, yielding a robust and efficient solution for real-world multimodal sensing tasks.

A Large-Scale Study on the Development and Issues of Multi-Agent AI Systems

arXiv:2601.07136v1 Announce Type: cross Abstract: The rapid emergence of multi-agent AI systems (MAS), including LangChain, CrewAI, and AutoGen, has shaped how large language model (LLM) applications are developed and orchestrated. However, little is known about how these systems evolve and are maintained in practice. This paper presents the first large-scale empirical study of open-source MAS, analyzing over 42K unique commits and over 4.7K resolved issues across eight leading systems. Our analysis identifies three distinct development profiles: sustained, steady, and burst-driven. These profiles reflect substantial variation in ecosystem maturity. Perfective commits constitute 40.8% of all changes, suggesting that feature enhancement is prioritized over corrective maintenance (27.4%) and adaptive updates (24.3%). Data about issues shows that the most frequent concerns involve bugs (22%), infrastructure (14%), and agent coordination challenges (10%). Issue reporting also increased sharply across all frameworks starting in 2023. Median resolution times range from under one day to about two weeks, with distributions skewed toward fast responses but a minority of issues requiring extended attention. These results highlight both the momentum and the fragility of the current ecosystem, emphasizing the need for improved testing infrastructure, documentation quality, and maintenance practices to ensure long-term reliability and sustainability.

app.build: A Production Framework for Scaling Agentic Prompt-to-App Generation with Environment Scaffolding

arXiv:2509.03310v2 Announce Type: replace Abstract: We present app.build (https://github.com/neondatabase/appdotbuild-agent), an open-source framework that improves LLM-based application generation through systematic validation and structured environments. Our approach combines multi-layered validation pipelines, stack-specific orchestration, and model-agnostic architecture, implemented across three reference stacks. Through evaluation on 30 generation tasks, we demonstrate that comprehensive validation achieves 73.3% viability rate with 30% reaching perfect quality scores, while open-weights models achieve 80.8% of closed-model performance when provided structured environments. The open-source framework has been adopted by the community, with over 3,000 applications generated to date. This work demonstrates that scaling reliable AI agents requires scaling environments, not just models -- providing empirical insights and complete reference implementations for production-oriented agent systems.

Key Information Influencing Patient Decision-Making About AI in Health Care: Survey Experiment Study

Background: Artificial Intelligence (AI)-enabled devices are increasingly used in healthcare. However, there has been limited research on patients’ informational preferences, including which elements of AI device labeling enhance patient understanding, trust, and acceptance. Clear and effective patient-facing communication is essential to address patient concerns and support informed decision-making regarding AI-enabled care. Objective: Using simulated AI device labels in a cardiovascular context, we evaluated three aims. First, we identified key information elements that influence patient trust and acceptance of an AI device. Second, we examined how these effects varied based on patient characteristics. Third, we explored how patients evaluated informational content of AI labels and their perceived effectiveness of the AI labels in informing decision-making about the use of AI device, building trust in the device, and shaping their intention to use it in their healthcare. Methods: We recruited 340 US patients from ResearchMatch.org to participate in a web-based survey that contained two experiments. In the discrete choice experiment (DCE), participants indicated preferences in terms of trust and acceptance regarding 16 pairs of simulated AI device labels that varied across eight types of information needs identified in our previous qualitative work. In the single profile factorial experiment (SPFE), participants evaluated four randomly assigned label prototypes regarding the label’s legibility, comprehensibility, information overload, credibility, and perceived effectiveness in informing about the AI device, as well as participants’ trust in the AI device and intention to use the device in their healthcare. Data was analyzed using mixed effects binary or ordinal logistic regression. Results: The DCE showed that information about regulatory approval, high device performance, provider oversight, and AI’s value added to usual care significantly increased the likelihood of patient trust by 14.1-19.3% and acceptance by 13.3-17.9%. Subgroup analyses revealed variations based on patient characteristics such as familiarity with AI, health literacy, and recency of last medical checkup. The SPFE showed that patients reported good label comprehension, and that information about provider oversight, regulatory approval, device performance, and AI’s added value improved perceived credibility and effectiveness of the AI label (odds ratios [ORs] range 1.35-2.05), reduced doubts in the AI device (ORs range 0.61- 0.77), and increased trust and intention to use the AI device (ORs range 1.47-1.73). However, information about data privacy and safety management protocols are less influential. Conclusions: Patients value information about an AI device’s performance, provider oversight, regulatory status, and added value during decision-making. Providing transparent, easily understandable information about these aspects is critical to support patient determinations of trust and acceptance of AI-enabled healthcare. Information elements’ impact on patient trust and acceptance varies by patient characteristics, highlighting the need for a tailored approach to address the concerns of diverse patient groups about AI in healthcare.

FACTS Benchmark Suite Introduced to Evaluate Factual Accuracy of Large Language Models

12 January 2026 at 15:55

A new industry benchmark aimed at systematically evaluating the factual accuracy of LLMs has been released with the launch of the FACTS Benchmark Suite. Developed by the FACTS team in collaboration with Kaggle, the suite expands earlier work on factual grounding and introduces a broader, multi-dimensional framework for measuring how reliably language models produce factually correct responses.

By Robert Krzaczyński
❌