❌

Reading view

Show-Harness: Just a VLM Agent Can Play Robots

arXiv:2609.10522v1 Announce Type: cross Abstract: Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.
  •  

A Human Audit of OpenAIs AI-Generated Mathematical Proofs

arXiv:2608.14673v3 Announce Type: replace Abstract: We assess 18 chapter-specific reviews of the ten mathematical results announced by OpenAI on 1 August 2026, alongside review standards, Lean formalizations, subsequent research, and mathematical references. The article audits this review record without claiming a complete reconstruction of all ten proofs. No confirmed substantive mathematical error in a principal result remains in the examined assessments, although review depth varies and some dependencies remain partly checked. Chapter 8 presents the strongest reservation: a specialist review requests major revision of compressed analytic arguments. In Chapter 6, an apparent polarity error was withdrawn after an overbar lost during PDF extraction was recovered from the typeset source. Subsequent research independently reuses the Chapter 3 proof mechanism and confirms that Connes's rigidity conjecture is false, without independently reproducing Chapter 4's stronger infinite-family result. Among the cited follow-ups, Chapter 7 receives the strongest direct theorem-level corroboration through a stronger hardness theorem. Related equality results in Chapter 8 do not verify the analytic inequality proof. Some follow-ups disclose material AI assistance. We argue that confidence should combine formal checking, human reconstruction, independent mathematical use, and a public record supporting correction of both proofs and reviews.
  •  

AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP

arXiv:2608.30107v2 Announce Type: replace-cross Abstract: Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and motivate more explicit geographic metadata for country-aware NLP evaluation.
  •  
  •  

Foaming photopolymers as a high-resolution biomimetic printing platform

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10968-9

Deep-foam photolithography uses light-controlled polymer foaming to create high-resolution, multifunctional microstructures with tunable optical, wetting and fluid-handling properties for advanced manufacturing applications.
  •  

Advancing conflict research and response through satellite-derived data

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11004-6

Integrating satellite-derived war-damage data with text-based fatality records through improvement, enrichment and fusion mitigates limitations inherent in each source, revealing complex violence dynamics beyond fatality-centric paradigms, as case studies from Ukraine and Myanmar illustrate.
  •  

Imaging cellular activity across all organs reveals body-wide circuits

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10979-6

An imaging system developed to record cellular activity throughout the whole body of zebrafish captures cellular organ dynamics and identifies multiple distributed circuits.
  •  

Proximity-guided graph learning reveals tumour-associated proximity antigens

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11003-7

A proximity-mapping atlas defines tumour-associated proximity antigens, revealing disease-associated membrane spatial communities, and identifies EGFR–CDCP1 as a co-target pair that enhances tumour killing by multispecific therapeutics.
  •  

Arrest of human spermatogenesis at the pachytene stage is frequently accompanied by disruption of the piRNA pathway

Cell Death Discovery, Published online: 03 September 2026; doi:10.1038/s41420-026-03327-0

Arrest of human spermatogenesis at the pachytene stage is frequently accompanied by disruption of the piRNA pathway
  •  

Transcriptomic profiling reveals complement activation in chronic pancreatitis adjacent to pancreatic ductal adenocarcinoma

Immunobiology. 2026 Aug 28;231(5):153239. doi: 10.1016/j.imbio.2026.153239. Online ahead of print.

ABSTRACT

Chronic pancreatitis (CP) and pancreatic ductal adenocarcinoma (PDAC) frequently coexist, yet distinguishing inflammatory changes secondary to malignancy from primary pancreatitis remains challenging. Here, we present an integrated multi-omics analysis of spatially distinct pancreatic tissue compartments obtained from a single patient undergoing pancreaticoduodenectomy for PDAC, combining histopathology, transcriptomics, immunofluorescence, and meta-transcriptomic microbial profiling. Morphological assessment identified tumor tissue, adjacent CP, and normal pancreas. Transcriptomic profiling of macro-dissected samples revealed differential gene expression and enrichment of the classical complement pathway in the CP compartment, which was qualitatively supported by immunofluorescent detection of C1q and C3 along the ductal epithelium. Meta-transcriptomic analysis detected a limited number of bacterial taxa and no viral RNA; these findings were interpreted conservatively given the constraints of low-biomass tissue profiling and the single-patient design. Although causal inference and generalizability are limited by the single-patient design, this study demonstrates the feasibility of integrating pathology with transcriptomic and microbial analyses to generate a hypothesis-generating, multi-omics framework for exploring inflammatory-malignant interactions in pancreatic disease.

PMID:42679433 | DOI:10.1016/j.imbio.2026.153239

  •  

Catching MRI outliers: unsupervised detection and localization of MRI artefacts and clinical anomalies using deep learning

arXiv:2605.24609v1 Announce Type: cross Abstract: Artificial intelligence is increasingly integrated into radiotherapy workflows, yet such pipelines remain vulnerable to out-of-distribution image data that may introduce unexpected behavior in clinical tasks. Deep learning-based anomaly detection for pelvic magnetic resonance imaging (MRI) remains largely unexplored, and transparent evaluation of its feasibility for full automation is limited. We developed and evaluated a fully automated, unsupervised anomaly-detection framework for pelvic and brain MRI. A two-stage framework was trained on reference images from public datasets: LUND-PROBE for pelvic MRI, and IXI, fastMRI, and fastMRI+ for brain MRI. In the first stage, MRI slices were compressed into discrete tokens; in the second, the distribution of normal tokens was modeled. Anomaly evidence was estimated by combining perceptual image differences with token-surprisal scores based on negative log-likelihood. Automated detection was evaluated on pelvic MRI with synthetic global and real clinical anomalies, and on brain MRI with clinically annotated fastMRI+ abnormalities. Sensitivity, specificity, area under the receiver operating characteristic curve (AUC), and false-positive behavior in held-out normal cases were assessed. The framework achieved robust detection across hidden evaluation cohorts, with AUCs of 0.97 (95% CI, 0.95-0.98) and 0.81 (95% CI, 0.74-0.87) for pelvic and brain MRI, respectively. Heatmap analysis showed strong spatial agreement between detected anomalies and ground-truth locations, supporting localization accuracy and interpretability. These results support the potential of unsupervised anomaly detection as an automated MRI quality-control layer for radiotherapy workflows, with transparent visualization of image regions likely to compromise downstream AI-based tasks.
  •  

Multi-Agent Specification-based Metamorphic Testing of FMU-Based Simulations

arXiv:2605.25101v1 Announce Type: cross Abstract: In many industrial domains, the Functional Mock-up Interface (FMI) is used to exchange simulation models as Functional Mock-up Units (FMUs) across different partners using various modelling tools. This opens up the possibilities for simulation-based verification and validation using FMUs for ensuring reliable system behaviour. However, deriving effective test oracles for these simulation models remains challenging due to the absence of explicit expected outputs. This limits the applicability of conventional testing approaches, which require access to the internal workings of the systems. Metamorphic testing (MT) addresses this limitation by leveraging metamorphic relations (MRs), but extracting such relations from specifications remains largely a manual and error-prone process. To address this challenge, we propose an LLM-powered multi-agent workflow for specification-based metamorphic testing of FMU-based simulation models. The approach takes functional and interface specifications as input and orchestrates multiple agents to extract requirements and derive MRs. These MRs are expressed using Given-When-Then patterns to structure input conditions (Given), transformations (When), and expected output behaviours (Then). These relations are then used to generate metamorphic test cases, execute simulations, and evaluate output consistency across multiple sessions. We evaluate the approach on a Lube Oil Cooling system FMU, demonstrating its ability to automatically generate meaningful MRs and corresponding test cases. Preliminary results indicate that the proposed workflow can effectively support the systematic verification and validation of dynamic simulation models by reducing manual effort and improving test generation.
  •  

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point

arXiv:2502.08047v5 Announce Type: replace Abstract: Recent progress in GUI agents has substantially improved visual grounding, yet robust planning remains challenging, particularly when the environment deviates from a canonical initial state. In real applications, users often invoke assistance mid-workflow, where software may be partially configured, steps may have been executed in different orders, or the interface may differ from its default setup. Such task-state variability is pervasive but insufficiently evaluated in existing GUI benchmarks. To address this gap, we introduce WorldGUI, a benchmark covering ten widely used desktop and web applications with tasks instantiated under diverse, systematically constructed initial states. These variations capture realistic human-computer interaction settings and enable diagnostic evaluation of an agent's ability to recover, adapt plans, and handle non-default contexts. We further present WorldGUI-Agent, a simple and model-agnostic framework that organizes planning and execution around three critique stages, improving reliability in dynamic environments. Experiments demonstrate that state-of-the-art GUI agents exhibit substantial performance degradation under non-default initial conditions, revealing limited robustness and fragile planning behaviors. Our benchmark and framework provide a foundation for developing more adaptable and reliable GUI agents. The code and data are available at https://github.com/showlab/WorldGUI.
  •  

Leveraging Spreading Activation for Improved Document Retrieval in Knowledge-Graph-Based RAG Systems

arXiv:2512.15922v3 Announce Type: replace Abstract: Despite initial successes and a variety of architectures, retrieval-augmented generation systems still struggle to reliably retrieve and connect the multi-step evidence required for complicated reasoning tasks. Most of the standard RAG frameworks regard all retrieved information as equally reliable, overlooking the varying credibility and interconnected nature of large textual corpora. GraphRAG approaches offer potential improvement to RAG systems by integrating knowledge graphs, which structure information into nodes and edges, capture entity relationships, and enable multi-step logical traversal. However, GraphRAG is not always an ideal solution, as it depends on high-quality graph representations of the corpus. Such representations usually rely on manually curated knowledge graphs, which are costly to construct and update, or on automated graph-construction pipelines that are often unreliable. Moreover, systems following this paradigm typically use large language models to guide graph traversal and evidence retrieval. In this paper, we propose a novel RAG framework that uses a spreading activation algorithm to retrieve information from a corpus of documents connected by an automatically constructed heterogeneous knowledge graph. This approach reduces reliance on semantic knowledge graphs, which are often incomplete due to information loss during information extraction, avoids LLM-guided graph traversal, and improves performance on multi-hop question answering. Experiments show that our method achieves better or comparable performance to several state-of-the-art RAG methods and can be integrated as a plug-and-play module with different iterative RAG pipelines. When combined with chain-of-thought iterative retrieval, it yields up to a 39% absolute improvement in answer correctness over naive RAG, while achieving these results with small open-weight language models.
  •  

SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned

arXiv:2602.07666v3 Announce Type: replace-cross Abstract: DARPA's AI Cyber Challenge (AIxCC, 2023--2025) is the largest competition to date for building fully autonomous cyber reasoning systems (CRSs) that leverage recent advances in AI -- particularly large language models (LLMs) -- to discover and remediate vulnerabilities in real-world open-source software. This paper presents the first systematic analysis of AIxCC. Drawing on design documents, source code, execution traces, and discussions with organizers and competing teams, we examine the competition's structure and key design decisions, characterize the architectural approaches of finalist CRSs, and analyze competition results beyond the final scoreboard. Our analysis reveals the factors that truly drove CRS performance, identifies genuine technical advances achieved by teams, and exposes limitations that remain open for future research. We conclude with lessons for organizing future competitions and broader insights toward deploying autonomous CRSs in practice.
  •  

A comparison of deep multiomics profiles across ethnicity, geography, and age

Multiomics profiling of healthy individuals reveals differences across molecular layers and key pathways related to immune, metabolic, and microbiome-linked processes across ethnicities, while geographic relocation reshapes these networks and influences aging trajectories.
  •  

Fronto-insular circuit mechanisms of accelerated intermittent theta burst stimulation

An optogenetic model of accelerated intermittent theta burst stimulation reveals cell type-specific plasticity mechanisms and a key role for a fronto-insular circuit in driving the antidepressant effects of this treatment in humans.
  •  
❌