❌

Normal view

LLM Agents as Social Scientists: A Human-AI Collaborative Platform for Social Science Automation

arXiv:2604.01520v1 Announce Type: new Abstract: Traditional social science research often requires designing complex experiments across vast methodological spaces and depends on real human participants, making it labor-intensive, costly, and difficult to scale. Here we present S-Researcher, an LLM-agent-based platform that assists researchers in conducting social science research more efficiently and at greater scale by "siliconizing" both the research process and the participant pool. To build S-Researcher, we first develop YuLan-OneSim, a large-scale social simulation system designed around three core requirements: generality via auto-programming from natural language to executable scenarios, scalability via a distributed architecture supporting up to 100,000 concurrent agents, and reliability via feedback-driven LLM fine-tuning. Leveraging this system, S-Researcher supports researchers in designing social experiments, simulating human behavior with LLM agents, analyzing results, and generating reports, forming a complete human-AI collaborative research loop in which researchers retain oversight and intervention at every stage. We operationalize LLM simulation research paradigms into three canonical reasoning modes (induction, deduction, and abduction) and validate S-Researcher through systematic case studies: inductive reproduction of cultural dynamics consistent with Axelrod's theory, deductive testing of competing hypotheses on teacher attention validated against survey data, and abductive identification of a cooperation mechanism in public goods games confirmed by human experiments. S-Researcher establishes a new human--AI collaborative paradigm for social science, in which computational simulation augments human researchers to accelerate discovery across the full spectrum of social inquiry.

Predictive Value of Machine Learning for Poststroke Mortality Risk: Systematic Review and Meta-Analysis

Background: People with stroke face a high mortality risk, and an accurate prediction model is essential to the guidance of clinical decision-making in this population. Recently, with growing attention paid to machine learning (ML) in stroke care, some researchers have investigated the effectiveness of ML in predicting the mortality risk in stroke. However, systematic evidence is still lacking for its effectiveness. Objective: This systematic review aims to evaluate the value of ML in predicting the stroke mortality risk. The findings are expected to offer an evidence-based basis for developing and assessing clinical risk prediction tools. Methods: A search was made in Cochrane Library, PubMed, Embase, and Web of Science up to June 23, 2025, and studies that reported a complete performance of ML in predicting stroke mortality were included. Studies with only risk factors analyzed were excluded. The risk of bias of the included studies was assessed using PROBAST (Prediction model Risk of Bias Assessment Tool). Pooled risk ratios with 95% CIs and prediction intervals (PIs) were derived using the Hartung-Knapp-Sidik-Jonkman method under a random-effects model. Subgroup analyses were also conducted by model type, stroke type, patient source, and treatment background. Moreover, a metaregression was conducted on the C-index for out-of-hospital mortality at different time points to explore the influence of time factors on the model’s predictive performance. Results: Sixty-eight studies were included (23 predicting in-hospital mortality and 45 predicting out-of-hospital mortality), describing the development of 75 prediction models and 43 external validations. The follow-up period was 1 month to 15 years. For predicting in-hospital mortality, the external validation set had a pooled C-index of 0.727 (95% CI 0.677-0.781, 95% PI 0.521-1.000), with sensitivity and specificity of 0.64 (95% CI 0.57-0.70) and 0.74 (95% CI 0.70-0.77), respectively. For predicting out-of-hospital mortality, the pooled C-index was 0.847 (95% CI 0.808-0.887, 95% PI 0.750-0.956) in the external validation set, with sensitivity and specificity of 0.71 (95% CI 0.55-0.82) and 0.76 (95% CI 0.74-0.78), respectively. Comparatively, the overall pooled C-indexes were 0.788 (95% CI 0.766-0.810, 95% PI 0.621-0.999) and 0.812 (95% CI 0.798-0.826, 95% PI 0.693-0.952), respectively. The metaregression revealed a gradual decline in the predictive performance of the overall model and logistic regression model alone, whereas a random forest model maintained sustained performance. Age, National Institutes of Health Stroke Scale score, and stroke-related complications were the most frequently used variables for modeling. Conclusions: This is the first meta-analysis to demonstrate that ML-based prediction of stroke mortality is feasible. The performance of ML supports its role as an auxiliary tool for identifying high-risk populations, thereby optimizing clinical monitoring and resource allocation. However, due to substantial heterogeneity and a relatively high risk of bias in available studies, caution is warranted in real-world application. The effectiveness of ML may vary across settings, and external validation is recommended before broader implementation. Trial Registration: PROSPERO CRD420251086321; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251086321

Breast Cancer Screening Knowledge and Sentiments in Singaporean Women: Mixed Methods Study Using Topic Modeling, Sentiment Analysis, and Structured Questionnaire Data

Background: Mammography screening uptake in Singapore remains below 40% despite campaigns and subsidies. Natural language processing (NLP) can extract nuanced attitudes from free text that fixed response options miss, revealing latent factors influencing breast cancer (BC) screening behavior. Objective: This study characterized women’s attitudes toward mammography using mixed methods data, examined associations between BC awareness and screening willingness, and identified barriers and facilitators through NLP of free-text responses. Methods: We conducted a cross-sectional study within the multicenter cohort in Singapore (October 2021-December 2023). In total, 4169 women aged 35‐59 years (median 48, IQR 43‐54) were recruited via convenience sampling (3 hospitals and 2 polyclinics). Participants completed online structured questionnaires on demographics and screening history, then a BC education quiz with feedback. Participants answering >80% correctly were classified as “BC-aware.” Posteducation, participants reported screening willingness (motivated or neutral) with optional free-text explanations. Logistic regression models (adjusted for study site, age, ethnicity, marital status, housing, and education) examined the associations with willingness. For 3819 English-language respondents, biterm topic modeling identified themes and sentiment analysis quantified emotional tone. Statistical significance: =.05. Results: Overall, 79% (3287/4169) were BC-aware, and 94% (3908/4169) reported increased motivation posteducation. BC-aware women had higher screening motivation than BC-unaware women (adjusted odds ratio [aOR] 2.88, 95% CI 2.19‐3.80;

Spectral Surgery: Training-Free Refinement of LoRA via Gradient-Guided Singular Value Reweighting

arXiv:2603.03995v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) improves downstream performance by restricting task updates to a low-rank parameter subspace, yet how this limited capacity is allocated within a trained adapter remains unclear. Through a geometric and empirical study across multiple tasks and backbones, we find that trained LoRA updates often exhibit an inefficient spectrum: task effects concentrate in a small subset of singular directions, while many remaining components are neutral or detrimental, motivating post-hoc refinement within the learned subspace. We propose Spectral Surgery, a training-free refinement that decomposes a LoRA update with SVD, estimates per-component sensitivity using gradients on a small calibration set, and reweights singular values under a magnitude constraint while keeping the learned directions fixed. Across Llama-3.1-8B and Qwen3-8B on four benchmarks, Spectral Surgery yields consistent gains (up to +4.4 points on CommonsenseQA and +2.4 pass@1 on HumanEval) by adjusting only $\approx 1{,}000$ scalar coefficients. These results demonstrate that SVD-structured, low-cost parameter editing can serve as a practical route to improving trained LoRA adapters in a purely post-hoc manner.
❌