❌

Normal view

Design Guidelines for Online Health Forums: User-Centered Design Approach

Background: Online health forums are used widely, yet evidence of their effectiveness is inconsistent. Evidence-based forum design guidance grounded in theory and lived experience could improve the efficacy and outcomes of these forums for the many people using them worldwide. Objective: This study aimed to draw on the experience of online forum users and staff, and insights from existing research on technology design and self-determination theory, to generate a set of theoretically grounded guidelines for safe and well-being–supportive online forums. Methods: We conducted 54 semistructured interviews (36 forum users and 18 forum staff) and 4 design workshops with forum staff, combined with input from a multidisciplinary research team. Principles of qualitative framework analysis were used to adapt a preexisting framework for well-being–supportive technology to the context of online health forums. Results: The resulting design guidelines are framed around 4 overarching principles relating to the psychological needs for autonomy, competence, and relatedness as defined by self-determination theory, and the additional need for safety in online forums. Each principle is presented alongside pragmatic design heuristics and specific implementation strategies. Conclusions: User experiences of online health forums are mixed. We have drawn on self-determination theory to propose evidence-informed guidelines for service development and refinement, adaptable to specific user groups across diverse settings. International Registered Report Identifier (IRRID): RR2-https://doi.org/10.1136/bmjopen-2023-075142

Java News Roundup: New OpenJDK JEPs, CDI 5.0, Spring, Open Liberty, RefactorFirst, ADK for Kotlin

15 September 2026 at 04:15

This week's Java roundup for September 7th, 2026, features news highlighting: new JEPs for ahead-of-time compilation and structured concurrency; GA releases of Jakarta CDI 5.0 and ADK for Kotlin 1.0; the September 2026 edition of Open Liberty; point releases of TornadoVM and RefactorFirst; a maintenance release of Micronaut; and first releases candidates of Groovy 6.0 and Gradle 9.8.

By Michael Redlich

Evaluation of the Square Eyes Model as a Screening Tool for Identifying Digital Technologies in Wearable Camera Images Among Children: Laboratory Study

Background: Accurate measurements of children’s digital technology use are essential for understanding its potential implications on health and well-being. Wearable cameras can provide such measurements, but image coding is a high burden for researchers. Machine learning–based object-recognition models have the potential to reduce this burden by identifying images containing technology. Objective: This study aims to evaluate the performance of an object recognition model, the Square Eyes model, as a screening tool for identifying technologies in wearable camera images among children for further human review, as well as to examine the potential influence of face-blurring methods on the model’s performance. Methods: This study used data collected on 48 children (aged 3‐14 y) during an approximately 1-hour laboratory session. The children performed various technology-related tasks while wearing a camera. A total of 221,226 images were coded by humans and processed through the Square Eyes model. The performance of the Square Eyes model as a screening tool was evaluated by (1) assessing agreement between the model and human coding; (2) evaluating the N-back algorithm, an algorithm embedded in the model aimed to flag images requiring human review; and (3) examining the potential influence of facial-blurring on model performance. Results: Humans detected technology in 92,745 (41.9%) images, and the Square Eyes model detected technologies with an overall accuracy of 78.0%. When considering specific technologies, agreement between the model and human coders was the highest for (n=19,148, 54.3%) and (n=8492, 44.5%) and lowest for smaller devices such as (n=2600, 31.3%) and (n=3685, 25.1%). The model’s N-back algorithm effectively flagged images that required further human review, with only 7144 (3.2%) images that were not flagged for screening containing a human-coded technology. An explorative analysis indicated that using a square face-blurring with border could have reduced the model’s ability to accurately detect technologies. Conclusions: The Square Eyes model demonstrated overall satisfying accuracy in detecting technologies and successfully flagged images that required further review by humans. These findings suggest that the model could be used as an effective screening tool for reducing the burden of human coding. However, the model could be improved to more accurately detect smaller devices, and the form of facial blurring in images should be considered.

Digitally Adapting LGBTQ-Affirmative Cognitive Behavioral Therapy for Chinese Men Who Have Sex With Men Living With HIV: User-Centered Design Approach

Background: Chinese men who have sex with men living with HIV (MSMLWH) experience substantial psychological distress driven by minority stress and HIV-related challenges. However, culturally tailored digital mental health interventions that address HIV-specific maladaptive cognitive schemas and culturally specific psychosocial stressors remain scarce in China. Objective: This study aimed to systematically adapt an evidence-based cognitive behavioral therapy (CBT) intervention Effective Skills to Empower Effective Men (ESTEEM) into a WeChat (Tencent) Mini-Program–based intervention (iESTEEM) specifically for Chinese MSMLWH and to evaluate its preliminary feasibility and usability. Methods: We used a three-phase user-centered design approach guided by the Assessment, Decision, Adaptation, Production, Topical Experts, Integration, Training, and Testing (ADAPT-ITT) framework. The study proceeded in three phases: (1) a qualitative needs assessment using semistructured interviews with 20 MSMLWH (mean age 23.25, SD 3.08 years); (2) systematic intervention adaptation and platform development, including theater testing (n=5); and (3) a 2-week pilot study involving 10 MSMLWH and five counselors to evaluate feasibility, usability, and acceptability through focus groups and objective platform analytics. Results: Phase 1 identified 3 major themes of psychological distress: persistent health anxiety fueled by catastrophizing, intersectional stigma internalization, the disclosure dilemma, and intimacy barriers rooted in defectiveness and shame schemas. Participants also prioritized anonymity and bite-sized learning. Guided by these findings, iESTEEM was developed as a counselor-assisted, privacy-preserving WeChat Mini-Program incorporating HIV-specific scenarios, multimodal learning modules, and a back-end risk-alert system. During the 2-week pilot, participants logged into the platform 14.1 (SD 6.7) times per person and completed 134.3 (SD 103.1) minutes of learning activities; all participants accessed module 1, and 90% (9/10) accessed modules 2‐5. Anxiety scores decreased from 8.9 (SD 2.3) to 7.2 (SD 3.0), whereas depression scores remained stable. All participants expressed a willingness to continue using the program and to recommend it to peers. Participants and counselors endorsed its contextual relevance, privacy protections, and clinical utility. Conclusions: This study provides a theory- and evidence-informed model for culturally adapting digital mental health interventions for highly stigmatized populations. By integrating lesbian, gay, bisexual, transgender, and queer (LGBTQ)-affirmative CBT principles, HIV-specific adaptations, and a privacy-preserving, counselor-assisted WeChat Mini-Program, iESTEEM demonstrated promising preliminary feasibility, acceptability, and engagement among Chinese MSMLWH. These findings support the potential of culturally tailored digital interventions to expand access to psychological support for this stigmatized population in resource-constrained settings. Ongoing randomized controlled trials will further evaluate its efficacy, implementation outcomes, and mechanism of action. Trial Registration: Chinese Clinical Trial Registry ChiCTR2400080263; https://www.chictr.org.cn/showproj.html?proj=216926
  • ✇MIT Technology Review
  • Donated livers can be made biologically younger Jessica Hamzelou
    Once an organ is removed from a donor’s body, the clock starts ticking. Surgeons usually flush the organ with a preservative solution, bag it, and put it on ice—where it immediately starts to degrade. The team has a matter of hours to get it into a recipient’s body. There’s another option—one that has been growing in popularity in recent years, especially for donated organs that aren’t in the healthiest state. Some hospitals opt to put them on machines that pump them with nutrients and rem
     

Donated livers can be made biologically younger

15 September 2026 at 00:11

Once an organ is removed from a donor’s body, the clock starts ticking. Surgeons usually flush the organ with a preservative solution, bag it, and put it on ice—where it immediately starts to degrade. The team has a matter of hours to get it into a recipient’s body.

There’s another option—one that has been growing in popularity in recent years, especially for donated organs that aren’t in the healthiest state. Some hospitals opt to put them on machines that pump them with nutrients and remove waste products, usually for around six to 12 hours. It’s a bit like being back in a body.

This allows doctors to assess the organs, and some recent studies suggest that time spent on these perfusion machines helps them do better once they’re transplanted. Now, scientists have found that perfused organs seem to get younger, at least at a molecular level.

The research, shared with MIT Technology Review, provides molecular clues as to why organs from younger donors are known to have a higher success rate. It might also help explain why perfused organs are less likely to fail once they make it into a recipient. 

The researchers behind the study hope to find new ways to test the health of donated organs and potentially develop additional tools to repair organs that might otherwise be discarded. “If [we] can improve the utilization of organs beyond what the current systems can do, then that’s a win in my book,” says Jesse Poganik, who studies aging at Brigham and Women’s Hospital in Boston and coauthored the study.

Clocking organs

Poganik—along with colleagues including Heidi Yeh and Alban Longchamp, transplant surgeons at Mass General Brigham—used “aging clocks” to assess donated livers. These are scientific tools designed to measure biological age—a result that is meant to convey more about the health status of an organ (or person) than chronological age.

In an initial experiment, the team used a clock to look at the patterns of chemical marks on DNA in 37 samples taken from 19 donated livers. Such epigenetic patterns are known to change as we age. But when the team compared samples from livers kept on ice and those that were perfused, the team found a “striking” pattern in the latter.

“Machine-perfused livers, in spite of being older or having other disadvantageous characteristics, had a biological age that was lower than [non-perfused] livers that were chronologically younger,” says Yeh, who led the work.

To investigate further, Yeh and her colleagues analyzed another 208 samples from 103 donated livers. This time, they used different aging clocks—ones that essentially measure how genes are working. They studied samples biopsied from the livers after they had been stored for up to around six hours either in cold storage or on machine perfusion.

In most cases, they also assessed a second sample taken around an hour after the livers had been transplanted into a recipient. Once the organ’s blood supply is reestablished in the body, “you have a few other things to do,” says Longchamp. “Then you just do a quick biopsy before you close.”

According to the clocks, which were developed to measure age and risk of death, the machine-perfused livers were biologically younger, the team found. “Pumping them at 34 degrees with oxygen and nutrients actually reversed the biological age,” says Longchamp. The results have been been shared with colleagues at an industry conference, he says. 

“If you adjust out chronological age … to have a fair head-to-head comparison, the difference between the two is on the order of 30%,” says Poganik. “It’s logical to say that perfusion drives this effect.”

The biological ages of all the livers tended to increase as soon as they were put into a recipient’s body, probably as a result of stresses on the organs. But still, the effect endured—the perfused organs remained biologically younger. 

Nathanael Raschzok, a transplant surgeon at Charité Universitätsmedizin Berlin in Germany who was not involved in the research, says the work is impressive. But it’s not yet clear what these changes might mean for the recipients of these organs, he says. The organs in the study were donated by people in their 30s, 40s, and 50s. Raschzok wants to know the effect of perfusion on the liver of an 80-year-old. “Every so often, we use organs from 70-, 80-, 85-year-old donors,” he says.

A better understanding of why the organs appear to be getting biologically younger might lead to therapies that achieve the same effect with a drug that could potentially be used to treat a donated organ for a fraction of the price, he adds. That’s important because perfusion is expensive—Raschzok says it costs around €10,000 in Germany (a quarter of the budget for a transplant), while the cost in the US comes to around $80,000 to $100,000 per organ, says Yeh.

Molecular repair

Yeh and her colleagues weren’t able to study most of the livers before perfusion. That’s because donated organs are generally not considered to be under the purview of the hospital until they’ve been placed on perfusion machines, she says. (Organ procurement procedures vary, but for the team as Mass General Brigham, donated organs are put on perfusion devices at the donor’s hospital. “There’s this sort of nebulous period where it’s not clear who the organ belongs to,” says Yeh.)

Still, by looking at the genes and molecular pathways that seem to be altered in perfused organs, she and her colleagues can garner some clues. At a molecular level, the team saw changes in cell pathways linked to inflammation and the structure of tissues, for example. They also saw more activity in a pathway that allows cells to remove and recycle damaged cell parts, says Yeh.

Poganik hopes to develop some kind of test that would determine which organs, on the basis of their biological age, are suitable for transplantation. He and his colleagues are also experimenting with potential drug treatments that might push the biological age of an organ even lower.

In the meantime, any liver that is not from a “perfect, young, brain-dead donor” could probably benefit from perfusion, says Yeh. The devices are already transforming transplant surgery. Just a few years ago, she says, she and her colleagues would avoid using livers from people who’d suffered a circulatory death (when the heart stops beating and there’s a damaging lack of blood flow to organs) and were over 40. Today, they use livers from such donors over the age of 70. “Perfusion has completely changed the landscape of transplantation in the last three years,” she says.

Podcast: How Will We Train Developers If AI Does the Routine Work: A Conversation with Scott Hanselman

14 September 2026 at 19:00

In this podcast, Michael Stiefel spoke to Scott Hanselman about developing new software engineers when artificial intelligence agents are doing most of the work on which junior developers were trained. Hanselman suggests the software industry should adopt a preceptorship model similar to the nursing profession.

By Scott Hanselman

Unveiling the Gastric Microbiome: Novel Insights into Early Detection and Pathogenesis of Gastric Cancer

Probiotics Antimicrob Proteins. 2026 Sep 14. doi: 10.1007/s12602-026-11134-3. Online ahead of print.

ABSTRACT

Gastric cancer (GC) remains a major global health burden and is frequently diagnosed at advanced stages owing to the limited sensitivity, invasiveness, and restricted availability of current screening strategies. Increasing evidence indicates that the gastric microbiome including Helicobacter pylori and diverse non-H. pylori bacteria, fungi, and viruses actively contribute to gastric carcinogenesis by modulating mucosal immunity, chronic inflammation, epithelial barrier integrity, and metabolism. High-throughput sequencing has revealed reproducible dysbiosis signatures in GC, characterized by enrichment of taxa such as Lactobacillus, Streptococcus, and Fusobacterium, and depletion of beneficial commensals, including Bifidobacterium and short-chain fatty acid producing anaerobes, some of which show promise as diagnostic or prognostic biomarkers. This review summarizes current knowledge on bacterial, fungal, and viral inhabitants of gastric tumors, highlighting their mechanistic roles in tumor initiation and progression and their potential as microbial indicators of disease. It further evaluates noninvasive and minimally invasive early detection strategies based on fecal and salivary microbiota profiling, urinary extracellular vesicles, and metabolomic fingerprints, alongside multi-omics integration and machine-learning models that combine microbial and host features to improve risk stratification. Finally, the review discusses therapeutic avenues including microbiota modulation, immunotherapy microbiome interactions, and personalized medicine frameworks that incorporate microbial signatures into clinical decision-making, underscoring the need for standardized, multicenter studies to translate gastric microbiome insights into robust tools for early detection and targeted intervention in GC.

PMID:42734869 | DOI:10.1007/s12602-026-11134-3

Unveiling the Gastric Microbiome: Novel Insights into Early Detection and Pathogenesis of Gastric Cancer

Probiotics Antimicrob Proteins. 2026 Sep 14. doi: 10.1007/s12602-026-11134-3. Online ahead of print.

ABSTRACT

Gastric cancer (GC) remains a major global health burden and is frequently diagnosed at advanced stages owing to the limited sensitivity, invasiveness, and restricted availability of current screening strategies. Increasing evidence indicates that the gastric microbiome including Helicobacter pylori and diverse non-H. pylori bacteria, fungi, and viruses actively contribute to gastric carcinogenesis by modulating mucosal immunity, chronic inflammation, epithelial barrier integrity, and metabolism. High-throughput sequencing has revealed reproducible dysbiosis signatures in GC, characterized by enrichment of taxa such as Lactobacillus, Streptococcus, and Fusobacterium, and depletion of beneficial commensals, including Bifidobacterium and short-chain fatty acid producing anaerobes, some of which show promise as diagnostic or prognostic biomarkers. This review summarizes current knowledge on bacterial, fungal, and viral inhabitants of gastric tumors, highlighting their mechanistic roles in tumor initiation and progression and their potential as microbial indicators of disease. It further evaluates noninvasive and minimally invasive early detection strategies based on fecal and salivary microbiota profiling, urinary extracellular vesicles, and metabolomic fingerprints, alongside multi-omics integration and machine-learning models that combine microbial and host features to improve risk stratification. Finally, the review discusses therapeutic avenues including microbiota modulation, immunotherapy microbiome interactions, and personalized medicine frameworks that incorporate microbial signatures into clinical decision-making, underscoring the need for standardized, multicenter studies to translate gastric microbiome insights into robust tools for early detection and targeted intervention in GC.

PMID:42734869 | DOI:10.1007/s12602-026-11134-3

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

arXiv:2609.11977v1 Announce Type: new Abstract: Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost--performance Pareto frontier. Supporting evaluations in tool calling, coding, and instruction following further show that this specialization preserves broad agentic capability. We release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • The Computational Primitives of Adaptation Jonathan W. Page
    arXiv:2609.11989v1 Announce Type: new Abstract: Research on adaptive systems has traditionally focused on behavior (what organisms do) and mechanism (how their machinery works). This paper focuses on a third level, computation, which considers what adaptive systems must compute to survive and reproduce. It is proposed that adaptation has its own computational structure, comprising a small set of primitive operations common to all adaptive systems, regardless of their physical form. Six primitiv
     

The Computational Primitives of Adaptation

arXiv:2609.11989v1 Announce Type: new Abstract: Research on adaptive systems has traditionally focused on behavior (what organisms do) and mechanism (how their machinery works). This paper focuses on a third level, computation, which considers what adaptive systems must compute to survive and reproduce. It is proposed that adaptation has its own computational structure, comprising a small set of primitive operations common to all adaptive systems, regardless of their physical form. Six primitives, Arouse, Orient, Valence, Position, Boundary, and Attune, were selected using four criteria: necessity for existence, universality across independently evolved lineages, evolutionary conservation, and irreducibility. From this, two main implications follow. First, in biological systems, the primitives provide a substrate-neutral method for cross-species comparisons, reframing elaborative behaviors like attention, memory, and decision-making as combinations of these computations. Second, in artificial systems, the primitives offer a new way to view current challenges in machine intelligence, such as confabulation, prompt injection, distractibility, reward hacking, and catastrophic forgetting, suggesting that these issues may arise from a lack of these computations. Thus, for biological and artificial systems that persist, no specific physical substrate is necessary; rather, these primitives must be implemented if the systems are to be adaptive. What biology offers to inform machine intelligence, then, is not the brain's neural design but the computational functions it evolved to perform. The Computational Primitives Theory (CPT) presented here is a working hypothesis, intended for further refinement through discussion, empirical testing, and application.

DU-NO: A Parameter-Efficient Double U-Shaped Neural Operator for Phase-Resolving Wave Modeling

arXiv:2609.12115v1 Announce Type: new Abstract: Phase-resolving wave models such as FUNWAVE-TVD are the accuracy standard for nearshore dynamics, resolving the shoaling, refraction, and breaking of individual waves, but their cost rules them out for the ensembles, uncertainty quantification, and real-time warning that operational forecasting demands. Neural operators promise solver-level accuracy at a fraction of that cost, yet on wave-dominated fields the accurate ones are large: hybrid spectral-convolutional operators such as U-FNO (the strongest baseline in our study after DU-NO) buy their fidelity with tens of millions of parameters. We introduce DU-NO (Double U-shaped Neural Operator), a multiscale U-shaped spectral operator that attaches lightweight convolutional U-Net branches only at its two shallowest encoder and decoder levels. The placement follows a sampling argument: high-wavenumber content exists only on fine grids, so the local, full-band pathways go where that content lives, while the coarse, band-limited levels stay purely spectral. A depth-decaying mode schedule holds the model to 3.64M parameters, an order of magnitude below U-FNO. On our publicly released FUNWAVE-TVD benchmark, DU-NO attains the best autoregressive rollout error of six identically trained architectures, improving on U-FNO by 14.9% with 10.8x fewer parameters, and a frequency-band analysis shows the gain holds across all bands, including the high-wavenumber band where truncated-spectral operators collapse. Parameter-matched controls confirm the gain is architectural: rescaled to the same 3.6M budget, the best baseline still trails DU-NO by 28.6%. The advantage carries beyond nearshore waves: DU-NO matches the strongest baselines on 2D Navier-Stokes and wins clearly on PDEBench shallow-water rollouts. Code, trained models, and evaluation artifacts are available at https://anonymous.4open.science/r/duno-code-5A7B/.

GLARE: Generative Learning via Adversarial Reward Estimation For Social Dynamics Forecasting

arXiv:2609.12165v1 Announce Type: new Abstract: Meeting continuation requires tracking the agenda, speaker roles, participant intentions, and disagreement across long multi-party discussions. We introduce the Meeting Dynamic Forecasting Benchmark (MDFB), constructed from 2,207 real-world meetings and 24,794 future-facing queries. Given a transcript prefix and an active question, a model generates a plausible multi-turn continuation in one call. We evaluate utility---progress toward the question---and human-likeness---plausible conversational flow and role consistency---without requiring exact reproduction of the observed future. We further present GLARE, an adaptation of adversarial imitation learning to conditional language generation. A discriminator ranks the observed continuation above samples from the current actor, and its score supplies a KL-regularized policy reward; retraining on current-policy negatives allows the reward landscape to evolve with the actor. GLARE attains average human-evaluated win rates of 0.66 on utility and 0.70 on human-likeness, outperforming SFT and SPIN while remaining below the observed human continuation. We also demonstrate MDFB as a social reasoning arena for comparing general-purpose models, including closed-source systems, through reference-assisted judgments. Together, these studies illustrate the benchmark's use for both task-specific learning and output-based evaluation of meeting behavior.

WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation

arXiv:2609.12171v1 Announce Type: new Abstract: Enterprise settings provide a challenging environment for question-answering agents, which often rely on Retrieval-Augmented Generation, Deep Research (DR), and related techniques. Much of this challenge comes from the complexity of enterprise data: information is often spread across evolving and potentially conflict- ing emails, chat messages, documents, and other artifacts. Existing benchmarks typically have limited real-world complexity, short-form responses, and unnatural queries, so they often fail to capture the challenges of enterprise settings. In this work, we introduce an automated pipeline for generating synthetic datasets of emails reflecting realistic workplace scenarios, along with long- and short-form questions and gold answers grounded in the data. Our method simulates long-running enterprise projects spanning several months and involving up to 25 interacting employees across multiple roles. The data emphasizes ambiguity, distributed information, and naturally occurring queries. To validate the pipeline, we evaluate few standard agentic baselines on our datasets using the latest frontier models. We find that aggregate scores averaged over all queries remain below 80% for each dataset, indicating significant room for improvement. These findings suggest that more work remains to be done for enterprise deployment and underscore the importance of realistic, high-complexity evaluation data for developing stronger real-world enterprise DR systems.

Soft Symbol Grounding for Prototypical Concepts

arXiv:2609.12247v1 Announce Type: new Abstract: Neuro-symbolic models are usually trained with supervision only on final labels, leaving the intermediate concepts unobserved. Since many concept assignments are consistent with a given label, training can predict labels correctly while recovering the wrong concepts, a failure known as a reasoning shortcut. Prototypical networks reduce shortcuts by anchoring each concept to a few labeled examples, but existing methods still couple perception and reasoning through a hand-crafted, task-specific differentiable loss that must be redesigned for every task. We introduce \textbf{Soft-PNet}, which removes this loss: it reframes concept grounding as a Metropolis walk over a precomputed cache of feasible symbolic solutions, guided by a prototype distribution built from a single labeled anchor per concept, and trains against one KL objective between the prototype-weighted cache and the network's concept predictions. The objective is identical across tasks and remains applicable when the solution space cannot be enumerated. On \texttt{MNIST-EvenOdd}, Visual Sudoku, and \texttt{Kand-Logic} under scarce supervision, Soft-PNet matches loss-engineered prototypical networks at the concept and label levels and recovers concepts that soft-grounding baselines miss, with no loss engineering and lower training time.

GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs

arXiv:2609.12265v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly asked to reason over structured data such as graphs, yet how reliably they can carry out multi-step graph algorithms in language remains unclear. Existing evaluations tend to use simple tasks on small graphs, to score code generation rather than reasoning over the graph itself, or to fix a single input format. We introduce Graph Theory Bench (GT Bench), a benchmark covering 24 classical graph problems in 44 task-structure settings, with over 100,000 examples across four representations: natural language, structured language, adjacency list, and adjacency matrix. Evaluating eight LLMs on GT Bench shows that accuracy is strongly tied to the input representation, that the best representation shifts with graph density, size, and topology as well as with the model, and that this sensitivity persists, attenuated, in the strongest reasoning models. Building on these observations, we propose the Graph Theory Agent (GTA), which pairs a preference-trained representation selector with plan-and-decompose scaffolding around a frozen executor LLM. GTA lifts Phi-4 from 53.5% to 69.1% on the benchmark's easy split and from 33.0% to 41.5% on its hard split, outperforming eight prompting and agent baselines, and transfers without retraining to GraCoRe and NLGraph. Code for benchmark generation and evaluation: https://github.com/xzx34/GTA. The project homepage is available at https://xzx34.github.io/gta/.

Robust Prototypical Networks for Few-Shot Sensor Fault Diagnosis

arXiv:2609.12287v1 Announce Type: new Abstract: Industrial fault diagnosis often operates with only a handful of labeled fault examples, making few-shot learning attractive for sensor monitoring. Standard prototypical networks are simple and effective; however, their class prototypes may become unstable in the very-low-shot regime because each decision relies on a small support set. We propose \emph{Multi-Episode Prototypical Networks} (MEPN), which aggregate prototypes from multiple disjoint support episodes and use their mean as the final class representative, reducing prototype variance without changing the encoder architecture. We evaluate MEPN on the DeFACTO sensor dataset using five-way fault classification with synthetic bias, drift, spike, and noise faults injected into real industrial measurements. Over 100 independent runs, MEPN reaches \textbf{\SensorOneShotGcpn\%} in the per-episode one-shot setting ($K\!=\!1$ shot, aggregated over $N_{\text{agg}}\!=\!10$ support episodes), substantially above single-episode baselines. Under an equal 10-sample support budget, MEPN and ProtoNet at $K\!=\!10$ are statistically indistinguishable, confirming prototype accumulation as the mechanism rather than superior fixed-budget learning.

Hybrid Physics-AI Framework of Body Center of Mass Dynamics from Wrist-Worn Sensors

arXiv:2609.12304v1 Announce Type: new Abstract: Wrist-worn IMU has been widely used for daily-life health monitoring. Yet, it does not fully represent whole-body dynamics, for which the body center of mass (COM) is considered the physiological reference standard. Therefore, this work proposes a simplified kinematic model (KM), which is designed to map the wrist IMU to the COM acceleration. It is built upon several reductive assumptions that enable the solvability of the dynamic equations based on wrist IMU measurements alone. This work further proposes three types of hybrid AI modeling methods, namely human kinematic model-based neural network (HKM-NN) models, to leverage the power of both grey-box and black-box modeling. The HKM-NN methods include serial learning (ser-) and two approaches of simultaneous learning (sim1- and sim2-). The proposed models are trained and tested using our dataset, which includes wrist IMU measurements and ground-truth COM measurements from 10 healthy volunteers during six gait activities and sit-to-stand (SS) transitional movement. The results demonstrate the feasibility of estimating COM acceleration from wrist IMU measurements. Our KM model yields satisfactory results, with an error ranging from 6.7% to 12.5% for gait activities and 5.6% for the SS. In comparison with the KM model, our HKM-NN models significantly enhance the performance, achieving 5.3% to 9.3% errors for gait activities, and the best error of 3.9% for the SS. In addition, the HKM-NN models demonstrate distinct robustness characteristics under noisy test conditions, with sim1-/sim2- generally maintaining greater robustness under Gaussian perturbations, while the KM model exhibits comparatively strong robustness under salt-and-pepper noise. These findings highlight the importance of combining biomechanical structure with data-driven learning for wearable sensing applications operating under imperfect and noisy measurement conditions.

AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems

arXiv:2609.12320v1 Announce Type: new Abstract: Traditional large language models (LLMs) are scoped to individual user sessions, limiting their knowledge to a single conversation and preventing them from learning user preferences that evolve over time. Existing agentic memory systems address this limitation but generally operate at the individual-user level, restricting the public knowledge that could be shared across users to improve downstream responses. We introduce AIM (Agentic Interoperable Memory), a unified, privacy-aware memory framework that enables multi-agent, multi-user LLM systems to persistently manage private and shared memory. AIM dynamically classifies information as private, scoped to one user and inaccessible to others, or public, accessible to all users. It enforces index-level access controls so that private memories are retrievable only by their owner, protecting sensitive data while allowing beneficial shared knowledge to improve coordination and consistency. We also introduce MUMBench (Multi-User Memory Benchmark), a dataset of multi-user interactions containing private and shareable information across four domains. To our knowledge, MUMBench is the first public dataset designed to evaluate multiple memory operations, including retrieval, creation, update, and deletion, in a multi-user environment. Across three independent runs on MUMBench, AIM achieves 96.0% visibility classification accuracy, 58.8% strict operation accuracy, and 70.5% state-aware operation accuracy.

Toward Robust Personalized Alignment for LLMs: Mitigating Persona Drift in Multi-Turn Dialogue

arXiv:2609.12373v1 Announce Type: new Abstract: Persona drift remains a central challenge for personalized language models, as user profiles evolve over long interactions rather than remain permanently fixed. Models must therefore revise persistent persona states when preferences genuinely change, while avoiding updates driven by transient, ambiguous, or unresolved observations. We propose CORE, which separates turn-local evidence from persistent persona-state revision and selectively updates grounded user preferences through uncertainty-aware belief revision. We also introduce PERSIST, a held-out post-anchor benchmark for persona-state robustness under sequential interaction stress, covering ambiguity, conflict, and controlled social influence. Across ALOE, PersonaChat, and PERSIST, CORE improves personalized alignment and robustness, with complementary gains in normalized closed-slot state fidelity. Human evaluation and mechanistic controls further support explicit update control beyond stronger generation or persistent memory alone.
❌