❌

Reading view

MedBuild AI: An Agent-Based Hybrid Intelligence Framework for Reshaping Agency in Healthcare Infrastructure Planning through Generative Design for Medical Architecture

arXiv:2511.11587v1 Announce Type: cross Abstract: Globally, disparities in healthcare infrastructure remain stark, leaving countless communities without access to even basic services. Traditional infrastructure planning is often slow and inaccessible, and although many architects are actively delivering humanitarian and aid-driven hospital projects worldwide, these vital efforts still fall far short of the sheer scale and urgency of demand. This paper introduces MedBuild AI, a hybrid-intelligence framework that integrates large language models (LLMs) with deterministic expert systems to rebalance the early design and conceptual planning stages. As a web-based platform, it enables any region with satellite internet access to obtain guidance on modular, low-tech, low-cost medical building designs. The system operates through three agents: the first gathers local health intelligence via conversational interaction; the second translates this input into an architectural functional program through rule-based computation; and the third generates layouts and 3D models. By embedding computational negotiation into the design process, MedBuild AI fosters a reciprocal, inclusive, and equitable approach to healthcare planning, empowering communities and redefining agency in global healthcare architecture.
  •  

How can we assess human-agent interactions? Case studies in software agent design

arXiv:2510.09801v2 Announce Type: replace Abstract: LLM-powered agents are both a promising new technology and a source of complexity, where choices about models, tools, and prompting can affect their usefulness. While numerous benchmarks measure agent accuracy across domains, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we propose PULSE, a framework for more efficient human-centric evaluation of agent designs, which comprises collecting user feedback, training an ML model to predict user satisfaction, and computing results by combining human satisfaction ratings with model-generated pseudo-labels. Second, we deploy the framework on a large-scale web platform built around the open-source software agent OpenHands, collecting in-the-wild usage data across over 15k users. We conduct case studies around how three agent design decisions -- choice of LLM backbone, planning strategy, and memory mechanisms -- impact developer satisfaction rates, yielding practical insights for software agent design. We also show how our framework can lead to more robust conclusions about agent design, reducing confidence intervals by 40% compared to a standard A/B test. Finally, we find substantial discrepancies between in-the-wild results and benchmark performance (e.g., the anti-correlation between results comparing claude-sonnet-4 and gpt-5), underscoring the limitations of benchmark-driven evaluation. Our findings provide guidance for evaluations of LLM agents with humans and identify opportunities for better agent designs.
  •  

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

arXiv:2510.24702v1 Announce Type: cross Abstract: Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this work, we argue that the bottleneck is not a lack of underlying data sources, but that a large variety of data is fragmented across heterogeneous formats, tools, and interfaces. To this end, we introduce the agent data protocol (ADP), a light-weight representation language that serves as an "interlingua" between agent datasets in diverse formats and unified agent training pipelines downstream. The design of ADP is expressive enough to capture a large variety of tasks, including API/tool use, browsing, coding, software engineering, and general agentic workflows, while remaining simple to parse and train on without engineering at a per-dataset level. In experiments, we unified a broad collection of 13 existing agent training datasets into ADP format, and converted the standardized ADP data into training-ready formats for multiple agent frameworks. We performed SFT on these data, and demonstrated an average performance gain of ~20% over corresponding base models, and delivers state-of-the-art or near-SOTA performance on standard coding, browsing, tool use, and research benchmarks, without domain-specific tuning. All code and data are released publicly, in the hope that ADP could help lower the barrier to standardized, scalable, and reproducible agent training.
  •  

Integration of multi-omics profiling reveals an epigenetic-based molecular classification of lung adenocarcinoma: implications for drug sensitivity and immunotherapy response prediction

Front Pharmacol. 2025 Feb 19;16:1540477. doi: 10.3389/fphar.2025.1540477. eCollection 2025.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD) remains a major cause of cancer-related mortality worldwide, with high heterogeneity and poor prognosis. Epigenetic dysregulation plays a crucial role in LUAD progression, yet its potential in molecular classification and therapeutic prediction remains largely unexplored.

METHODS: We performed an integrated multi-omics analysis of 432 LUAD patients from TCGA and 398 patients from GEO datasets. Using consensus clustering and random survival forest (RSF) algorithms, we established an epigenetic-based molecular classification system and constructed a prognostic model. The model's performance was validated in multiple independent cohorts, and its biological implications were investigated through comprehensive functional analyses.

RESULTS: We identified two distinct molecular subtypes (CS1 and CS2) with significant differences in epigenetic modification patterns, immune microenvironment, and clinical outcomes (P = 0.005). The RSF-based prognostic model demonstrated robust performance in both training (TCGA-LUAD) and validation (GSE72094) cohorts, with time-dependent AUC values ranging from 0.625 to 0.694. Low-risk patients exhibited enhanced immune cell infiltration, particularly CD8+ T cells and M1 macrophages, and showed better responses to immune checkpoint inhibitors. Drug sensitivity analysis revealed subtype-specific therapeutic vulnerabilities, with low-risk patients showing higher sensitivity to conventional chemotherapy and targeted therapy.

CONCLUSION: Our study establishes a novel epigenetic-based classification system and predictive model for LUAD, providing valuable insights into patient stratification and personalized treatment selection. The model's ability to predict immunotherapy response and drug sensitivity offers practical guidance for clinical decision-making, potentially improving patient outcomes through precision medicine approaches.

PMID:40046740 | PMC:PMC11879945 | DOI:10.3389/fphar.2025.1540477

  •  
❌