❌

Normal view

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

arXiv:2609.13009v1 Announce Type: new Abstract: Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.

Investigation Into the Association Between Neurotransmitters, Immune Features, and Lung Adenocarcinoma: Identifying GABA-Related Features Using Machine Learning Methods

Stem Cells Int. 2026 May 21;2026:3060138. doi: 10.1155/sci/3060138. eCollection 2026.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD), a predominant subtype of non-small cell lung cancer (NSCLC), is associated with a high mortality rate. Currently, there are no reliable or sensitive biomarkers or prognostic methodologies available for its early detection or diagnosis. Gamma-aminobutyric acid (GABA), a pivotal inhibitory neurotransmitter within the central nervous system (CNS), primarily exerts its effects through interactions with GABA receptors (GABARs). Recent studies have increasingly highlighted GABA's significant role in mediating the initiation and progression of various tumors beyond the CNS. Nonetheless, research investigating the role of GABA in LUAD is limited, and the specific molecular and cellular mechanisms underlying its interactions remain to be fully elucidated.

METHODS: We developed an innovative machine learning framework designed to screen GABA-related genes (GABARgenes) at both single-cell and large transcriptomic levels. This framework encompasses 10 algorithms and 101 combinatorial pairing patterns, which facilitate the construction of consistent GABA-related features (GABARFs). The framework's performance was assessed using both a training set and an external validation set. To provide a quantitative prognostic tool for clinical application, we established a nomogram that incorporates GABARF. Additionally, we conducted multiomics analyses, including genomics, single-cell transcriptomics, and comprehensive transcriptomics, to derive and consolidate more extensive prognostic features. We also evaluated the response of GABARF-defined risk subgroups to immunotherapy and identified potential personalized therapeutic agents for specific risk categories.

RESULTS: Among the 124 GABARgenes analyzed, 38 demonstrated a significant correlation with overall survival (OS) in patients. Our machine learning-derived GABARF exhibited exceptional performance in predicting prognosis and clinical outcomes, showing promise in forecasting the onset and progression of LUAD. Multivariate analysis confirmed that GABARF serves as an independent prognostic factor for OS in LUAD. Furthermore, distinct GABARF risk subgroups exhibited significant differences in biological function, mutation status, and tumor immune infiltration. Notably, there were significant variations in the immunophenoscore (IPS) across the risk subgroups. GABARF risk stratification aligns with stemness properties of tumor cells, indicating that high-risk patients may harbor tumors with enhanced stemness features that contribute to their poor prognosis and reduced immunotherapy response. Sensitivity analyses of conventional LUAD therapies indicated that patients in the low-risk group may derive greater benefit from immune checkpoint inhibitors (ICIs), while those in the high-risk group may exhibit heightened sensitivity to first-line chemotherapy agents. Furthermore, LDHA overexpression was found to promote proliferation and migration, while inhibiting apoptosis. In addition, overexpression of LDHA can upregulate the expression of stemness markers CD133, SOX2, and OCT4 in LUAD cells, enhancing the malignant phenotype of tumor cells.

CONCLUSION: This study presents a novel machine learning-based model for GABARF, which shows promise as a potential tool to aid in prognostic prediction, targeted prevention, and individualized treatment planning in LUAD. Initial investigations into the interaction mechanisms of GABARF at the molecular, cellular, and tumor immune microenvironment (TIME) levels in LUAD have commenced. The GABARF model not only serves as a prognostic indicator but may also reflect the stemness status of LUAD tumors, offering insights into personalized treatment strategies that account for both neural-immune-stemness interactions.

PMID:42181964 | PMC:PMC13191820 | DOI:10.1155/sci/3060138

Investigation Into the Association Between Neurotransmitters, Immune Features, and Lung Adenocarcinoma: Identifying GABA-Related Features Using Machine Learning Methods

Stem Cells Int. 2026 May 21;2026:3060138. doi: 10.1155/sci/3060138. eCollection 2026.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD), a predominant subtype of non-small cell lung cancer (NSCLC), is associated with a high mortality rate. Currently, there are no reliable or sensitive biomarkers or prognostic methodologies available for its early detection or diagnosis. Gamma-aminobutyric acid (GABA), a pivotal inhibitory neurotransmitter within the central nervous system (CNS), primarily exerts its effects through interactions with GABA receptors (GABARs). Recent studies have increasingly highlighted GABA's significant role in mediating the initiation and progression of various tumors beyond the CNS. Nonetheless, research investigating the role of GABA in LUAD is limited, and the specific molecular and cellular mechanisms underlying its interactions remain to be fully elucidated.

METHODS: We developed an innovative machine learning framework designed to screen GABA-related genes (GABARgenes) at both single-cell and large transcriptomic levels. This framework encompasses 10 algorithms and 101 combinatorial pairing patterns, which facilitate the construction of consistent GABA-related features (GABARFs). The framework's performance was assessed using both a training set and an external validation set. To provide a quantitative prognostic tool for clinical application, we established a nomogram that incorporates GABARF. Additionally, we conducted multiomics analyses, including genomics, single-cell transcriptomics, and comprehensive transcriptomics, to derive and consolidate more extensive prognostic features. We also evaluated the response of GABARF-defined risk subgroups to immunotherapy and identified potential personalized therapeutic agents for specific risk categories.

RESULTS: Among the 124 GABARgenes analyzed, 38 demonstrated a significant correlation with overall survival (OS) in patients. Our machine learning-derived GABARF exhibited exceptional performance in predicting prognosis and clinical outcomes, showing promise in forecasting the onset and progression of LUAD. Multivariate analysis confirmed that GABARF serves as an independent prognostic factor for OS in LUAD. Furthermore, distinct GABARF risk subgroups exhibited significant differences in biological function, mutation status, and tumor immune infiltration. Notably, there were significant variations in the immunophenoscore (IPS) across the risk subgroups. GABARF risk stratification aligns with stemness properties of tumor cells, indicating that high-risk patients may harbor tumors with enhanced stemness features that contribute to their poor prognosis and reduced immunotherapy response. Sensitivity analyses of conventional LUAD therapies indicated that patients in the low-risk group may derive greater benefit from immune checkpoint inhibitors (ICIs), while those in the high-risk group may exhibit heightened sensitivity to first-line chemotherapy agents. Furthermore, LDHA overexpression was found to promote proliferation and migration, while inhibiting apoptosis. In addition, overexpression of LDHA can upregulate the expression of stemness markers CD133, SOX2, and OCT4 in LUAD cells, enhancing the malignant phenotype of tumor cells.

CONCLUSION: This study presents a novel machine learning-based model for GABARF, which shows promise as a potential tool to aid in prognostic prediction, targeted prevention, and individualized treatment planning in LUAD. Initial investigations into the interaction mechanisms of GABARF at the molecular, cellular, and tumor immune microenvironment (TIME) levels in LUAD have commenced. The GABARF model not only serves as a prognostic indicator but may also reflect the stemness status of LUAD tumors, offering insights into personalized treatment strategies that account for both neural-immune-stemness interactions.

PMID:42181964 | PMC:PMC13191820 | DOI:10.1155/sci/3060138

GPA: Learning GUI Process Automation from Demonstrations

arXiv:2604.01676v2 Announce Type: replace-cross Abstract: GUI Process Automation (GPA) is a lightweight but general vision-based Robotic Process Automation (RPA), which enables fast and stable process replay with only a single demo. Addressing the fragility of traditional RPA and the non-deterministic risks of current vision language model-based GUI agents, GPA introduces three core benefits: (1) Robustness via Sequential Monte Carlo-based localization to handle rescaling and detection uncertainty; (2) Deterministic and Reliability safeguarded by readiness calibration; and (3) Privacy through fast, fully local execution. This approach delivers the adaptability, robustness, and security required for enterprise workflows. It can also be used as an MCP/CLI tool by other agents with coding capabilities so that the agent only reasons and orchestrates while GPA handles the GUI execution. We conducted a pilot experiment to compare GPA with Gemini 3 Pro (with CUA tools) and found that GPA achieves higher success rate with 10 times faster execution speed in finishing long-horizon GUI tasks.

GPA: Learning GUI Process Automation from Demonstrations

arXiv:2604.01676v1 Announce Type: cross Abstract: GUI Process Automation (GPA) is a lightweight but general vision-based Robotic Process Automation (RPA), which enables fast and stable process replay with only a single demo. Addressing the fragility of traditional RPA and the non-deterministic risks of current vision language model-based GUI agents, GPA introduces three core benefits: (1) Robustness via Sequential Monte Carlo-based localization to handle rescaling and detection uncertainty; (2) Deterministic and Reliability safeguarded by readiness calibration; and (3) Privacy through fast, fully local execution. This approach delivers the adaptability, robustness, and security required for enterprise workflows. It can also be used as an MCP/CLI tool by other agents with coding capabilities so that the agent only reasons and orchestrates while GPA handles the GUI execution. We conducted a pilot experiment to compare GPA with Gemini 3 Pro (with CUA tools) and found that GPA achieves higher success rate with 10 times faster execution speed in finishing long-horizon GUI tasks.

Low-dose intestinal irradiation enhances the efficacy and prognosis of PD-1 blockade in metastatic non-small cell lung cancer

Clin Cancer Res. 2026 Mar 18. doi: 10.1158/1078-0432.CCR-25-4153. Online ahead of print.

ABSTRACT

PURPOSE: Intestinal low-dose irradiation (ILDR) may enhance immunotherapy efficacy by modulating the gut microbiota and metabolism; however, its role in metastatic non-small cell lung cancer (mNSCLC), particularly in the first-line setting, remains unclear.

EXPERIMENTAL DESIGN: This multicenter retrospective and prospective study included mNSCLC patients receiving first- and second-line programmed cell death protein 1 (PD-1) inhibitors along with abdominopelvic radiotherapy between 2018 and 2025. Patients were stratified by the mean intestinal radiation dose into <1 Gy, 1-3 Gy, and >3 Gy groups and treatment outcomes were compared. The blood and fecal samples were subjected to multi-omics profiling.

RESULTS: g>309 patients were included in the retrospective analysis. Optimal efficacy was observed with a small intestinal mean radiation dose (SIMRD) of 1-3 Gy, showing longer progression-free survival (PFS, 10.2 months) and overall survival (OS, 22.8 months) (P < 0.01), which was consistent across subgroups. Compared with 1-3 Gy, SIMRD >3 Gy (Hazard ratio [HR] = 4.87, P < 0.001) and <1 Gy (HR = 1.85, P < 0.001) independently predicted worse OS. Prospective results confirmed the best disease control rate (P = 0.041) and PFS (P = 0.046) with SIMRD of 1-3 Gy. Responders were enriched in Bacillota, Clostridia, and indole derivatives, particularly indole-3-carboxylic acid. Moreover, the 1-3 Gy group exhibited increased circulating macrophage inflammatory protein-3α and reduced circulating α4β7+ regulatory T cells.

CONCLUSIONS: ILDR influences the efficacy of PD-1 blockade in patients with mNSCLC, particularly when SIMRD is maintained within the 1-3 Gy range, likely through modulation of the gut microbiota-metabolite-immune axis.

PMID:41849236 | DOI:10.1158/1078-0432.CCR-25-4153

❌