❌

Reading view

Large Language Model–Generated Patient Instructions for Prescriptions in Primary Health Care: Preclinical Algorithm Validation

Background: The application of generative artificial intelligence to simplify medication use instructions has the potential to enhance people’s health by improving treatment adherence. Objective: We evaluated the performance of large language models (LLMs) in generating medication usage instructions to complement prescriptions in primary health care. Methods: This randomized, blinded experimental preclinical study used prescription-inducing scenarios, assigned to 62 health care professionals, to validate instructions generated by LLMs during electronic prescriptions. The instructions were generated by ChatGPT-4.0 (OpenAI), Llama3.1-8B (Meta), and Llama3.1-8B-RAG (Meta) using retrieval-augmented generation based on patient information leaflets. Performance metrics assessed adequacy, completeness, clarity, language simplification, usefulness, and errors in the generated instructions, with scores to analyze overall and individual metrics. Results: The 3 models yielded high overall scores for producing qualified instructions (ChatGPT-4.0: median 88.4, IQR 22.8; Llama3.1-8B: median 66.5, IQR 50.9; Llama3.1-8B-RAG: median 79.9, IQR 34.4; Kruskal-Wallis test P=.003). Llama3.1-8B-RAG received evaluations with similar overall scores to ChatGPT-4.0 (post hoc test, P=.05) and similar to Llama3.1-8B (post hoc test, P=.44). ChatGPT-4.0 outperformed Llama3.1-8B (Bonferroni test, P<.001). Regarding specific domains, Llama3.1-8B-RAG received scores equivalent to those of ChatGPT-4.0 for adequacy (mean 6.24, SD 2.3 vs mean 6.82, SD 2.1; post hoc test, P=.54); completeness (mean 5.94, SD 2.2 vs 6.55, SD 1.9; post hoc test P=.38), clarity (mean 5.77, SD 2.4 vs mean 6.68, SD 1.9; post hoc test P=.09), and usefulness (mean 5.42, SD 2.4 vs mean 5.96, SD 2.2; post hoc test P=.63). ChatGPT-4.0 received higher scores in the language simplification criterion than Llama3.1-8B-RAG (mean 7.05, SD 1.5 vs mean 5.44, SD 2.6; post hoc test P<.001). Interrater variability in assigning scores ranged from 4.2% (n=3) to 85.8% (n=6) among primary health care professionals. Instructions leading to incorrect use of the medication had similar frequency among the models(ChatGPT-4.0: n=15, 22.7%; Llama3.1-8B: n=19, 22.8%; Llama3.1-8B-RAG: n=19, 22.8%; chi-square test P=.71). The frequencies of hallucination were similar (ChatGPT-4.0: n=7, 10.6%; Llama3.1-8B: n=9, 13.6%; Llama3.1-8B-RAG: n=6, 9.1%; chi-square test P=.67). Conclusions: The open-source LLM enhanced with external information presented similar performance to the closed-source model, except for ChatGPT4.0, which was superior in language simplification of messages. LLM generation demonstrated potential for instructing patients on medication use. Nonetheless, the introduction of this innovation into the electronic prescribing workflow demands prescriber validation for human oversight of the technology and requires a strategy for LLM performance governance.
  •  

STAT+: Five biotech news updates to stay on top of today

Want to stay on top of the science and politics driving biotech today? Sign up to get our biotech newsletter in your inbox.

Hi! Hope you had a nice extended weekend.

Today: Eli Lilly’s gene-editing data seems promising for high cholesterol, an AI drug discovery CEO dispelled some AI drug discovery hype, and the new interim FDA chief is so far well received.

Continue to STAT+ to read the full story…

© Adobe

  •  

A reality check on the AI jobs hysteria

Haven’t you heard? White-collar jobs are going away, decimated by AI. Waves of layoffs in the tech sector (most recently at Coinbase and Meta and Cisco) are said to presage what will soon come for all of us knowledge workers. But before you quit your job as a software developer or financial analyst—or tech journalist—and look to join the plumbers’ union, it’s worth considering today’s economic research on whether artificial intelligence has actually begun to devour white-collar work.

The short answer is: No.

Despite the warning by some of an imminent jobs apocalypse that will destroy much of if not most such work, or the rumblings about a “permanent underclass,” there’s scant evidence that AI has yet had any large-scale impact on the US labor market. 

Analysis of the data gathered for the US Bureau of Labor Statistics (BLS) shows that the unemployment rate for the jobs potentially most affected by AI is actually lower than that for occupations less exposed to the technology. And, critically in the mind of economists, there are no signs that large numbers of people are shifting from jobs threatened by AI to supposedly safer ones, such as those involving mostly manual labor.

While the current labor statistics don’t preclude a sudden job upheaval in the coming years, they do throw doubt on the inevitability of the doomsday scenarios and the pace at which they’d unfold. Everyone in the AI community, it seems, is predicting that the technology will soon wipe out jobs, and everyone, it also seems, knows some young wannabe workers who can’t find one. Perhaps we haven’t seen any major disruption in the labor market statistics yet, people often say, but just wait. 

But maybe we should pay attention to what the data is showing us. And right now, the numbers paint a picture of a relatively stable labor market in which AI disruptions remain largely speculative.

“It could be disruptive, but the data is telling us right now that disruption is not yet here, and we have time to plan.”

“All of the available evidence to date suggests that AI’s impact on current labor market conditions is likely small right now,” says Erika McEntarfer, a labor economist who headed the BLS until President Trump fired her last fall after a jobs report that displeased the administration. (Not surprisingly, BLS reports of sluggish job growth have continued since her dismissal.)

McEntarfer, who is now a fellow at the Stanford Institute for Economic Policy Research, says the relatively small impact that AI is having so far on today’s labor market “surprises many people, but it shouldn’t. What we know from history is that it takes time for innovations to work their way through changes in industries and changes in occupations. AI is unlikely to transform labor markets until it first transforms businesses.”

McEntarfer points to US Census data showing that only one in five companies are using AI in any business function. “The data are a great reality check on the fear that AI will be enormously disruptive,” she says. “It could be. It likely will be disruptive, but the data is telling us right now that disruption is not yet here, and that we have time to plan.”

Things ain’t great—but the question is why

The US job market, to be sure, sucks for many, especially younger would-be workers. Unemployment rates for recent college graduates stand at around 5.6%, well above the level for all workers. It’s a rate not seen since the pandemic and the years immediately after the 2008 recession. Even more troubling is that hiring rates have been particularly dismal during the post-covid economy, a trend that hits hard at young people trying to enter the workforce. If you’re a recent college graduate and looking for a tech job, no one, it can seem, is hiring.

There are signs that AI is contributing to the pain for the 22-to-25-year-olds seeking jobs in software development and other occupations that are feeling a big impact from AI. But these professions represent just a sliver of the overall labor market. What’s more, it’s uncertain how much blame AI should get for the job woes. Similarly unknown is whether the loss of entry-level jobs in AI-exposed occupations is a harbinger of what’s coming for others or simply an isolated symptom of what economists refer to as a “low-fire, low-hire” labor market caused by a variety of macroeconomic forces.

Insights into these uncertainties will tell us much about our working fates in the transition to an AI economy. There are no shortage of confident assertions and predictions about what is about to happen; while some people forecast the end of work, others say economic history teaches us that technology advances always lead to more and better jobs eventually. 

The honest answer is that no one knows for sure what AI will bring and whether this time will be different. To help figure it out, we need better and far more comprehensive data.

The statistics gleaned from the federal government’s monthly survey of 60,000 households for the BLS provide a broad overview of the changes to the labor market, while academics and even some AI companies have begun trying to gain a more granular view of specific jobs that are being affected. But the existing data-gathering tools don’t adequately explain how AI is affecting the huge and diverse US labor market.

There’s a long list of questions that we don’t have the data to fully answer. How is AI being used in the workplace? Does the increased use of AI mean the technology will replace workers, or will it make them more productive and valuable? Which occupations and skills are most affected? Who is in most peril from the changes? As David Deming, a professor of economics at Harvard University, puts it: “We’re sort of flying blind.”

To gather more insight into some of these questions, Deming and his colleagues have been surveying several thousand people every three months since 2024, asking them basic questions: Do you use generative AI, and how often? Does it save you time at work? Tracking the answers over time gives the economists important clues (it’s used by a little over 40% of workers but adoption varies by sectors) and allows them to estimate productivity gains (they’ve found some, but nothing economy-shaking). It has also helps document how quickly AI has been adopted in the workplace and how it compares with earlier technologies such as the PC and the internet (the pace has been faster but roughly in the same ballpark).

It’s far from a complete picture of how AI is changing work. But it provides some intriguing results; for example, a fair number of workers in manufacturing and other industrial sectors have tried AI. Deming’s results show that while businesses in general might be relatively slow to formally adopt the technology, lots of their employees are using it.

Getting a picture of these early adopters and how they’re using AI provides a “crystal ball for the future of the labor market,” Deming says. “It gives you important clues about how it’s going to be used tomorrow, and who’s going to be affected, and who’s going to be harmed and how do we need to get ready for it. It’s a diagnostic of what’s coming down the road.”

But what it doesn’t tell you is the fate of various jobs.

The young are most vulnerable

Analysis of how AI will affect jobs typically begins with identifying so-called exposure of various occupations to the technology. This approach is based on the idea that any given job is a collection of tasks. By evaluating which tasks can be performed by, say, the latest large language model, researchers gauge an occupation’s overall exposure. A small army of economists have created a slew of such studies, meticulously ranking hundreds of jobs and scrambling to update the results as the capabilities of generative AI keep exploding. 

The results have often triggered a panic, with graphics showing the growing vulnerability of different jobs to AI.

But by themselves the exposure results are not a true predictor of which jobs will be lost to AI. That depends on the kinds of tasks done by the technology, the extent to which the AI is adopted, various business calculations about the value of workers, and even the costs of deploying AI. But the exposure findings are a valuable starting point. 

In a working paper called “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence,” researchers at the Stanford Digital Economy Lab looked at 950 jobs, placing the occupations into five categories from least exposed to most. Then they used a vast data set from ADP, the world’s largest payroll provider, to look at employment growth in each of the categories. Their exclusive access to the ADP data set, which is far larger than the one available through the BLS, allows the researchers to better spot impacts by demographic. When they examined what was happening to different age groups, says Erik Brynjolfsson, the director of the lab who led the effort, “it was extremely striking.”

They spotted the drop in head count for 22-to-25-year-olds in the most exposed occupations, such as software development and customer service, beginning in late 2022, when ChatGPT was first publicly released. Other researchers reported evidence that the decline in these jobs began well before ChatGPT and questioned whether the labor market could react so quickly to the introduction of AI technology. 

But while the Stanford researchers acknowledge that other factors in addition to AI probably contributed to the early declines, they say that after controlling for those factors, they saw convincing evidence of a significant effect from AI after 2024 and growing in 2025 to a 16% decline in entry-level jobs in AI-exposed occupations. In contrast, head count grew for older workers in the same occupations, as did the number of jobs in the less exposed occupations.

Digging deeper into the data, the researchers found another important clue, though one that wasn’t totally unexpected. The impact on head counts depended on how AI was being used. It was specifically the jobs where tasks could be automated (that is, AI could do them “with minimal human involvement”) that accounted for the decrease in employment—jobs for people like software developers. In jobs where AI was mainly used but to augment human work, head counts grew faster than the average for entry-level workers.

That’s consistent with one explanation for the woes of many young would-be workers. It could be, according to the Stanford paper, that entry-level jobs depend more on the types of knowledge that people acquire through education but that can readily be mimicked by AI; the authors call this codified knowledge. It might be particularly easy to automate such tasks as entry-level coding. In contrast, older workers have more so-called tacit knowledge, the type based on their experience. That type of wisdom is harder for AI to replace.

Despite the findings about AI’s impact on young workers, Bharat Chandar, an economist at Stanford and one of the authors (along with Brynjolfsson and Ruyu Chen), stresses that it’s still early when it comes to understanding how the technology will affect jobs in the future. It could be that the job loss will spread to older workers and to less AI-exposed occupations, he says. But Chandar says it is also possible that firms and workers will adjust to shifting labor demands, and the effects will level off or even disappear.

To track how it plays out, the Stanford Digital Economy Lab is about to launch a regularly updated project providing data on how AI is transforming the economy.

The Stanford research and other work has put a particular spotlight on coding, a task at which AI is getting extremely adept. 

A recent paper by economists at the Federal Reserve Board found, not surprisingly, that annual employment growth for coders has slowed significantly—by about 3%—since the introduction of ChatGPT. But here’s a critical detail: Overall employment for coders continues to grow. Employment in coding jobs is still rising, they noted, just more slowly than before 2022. 

In short, coding jobs are not going away, at least not anytime soon. But it’s an occupation that is clearly being transformed by AI.

One of the somewhat surprising wrinkles uncovered by recent research is that wages in sectors highly exposed to AI have risen relatively fast since the introduction of ChatGPT. One explanation is that employers are still willing to pay for the kinds of knowledge and experience that are, at least for now, hard to replace with AI. If true, this suggests not the end of work in AI-exposed jobs but, more specifically, the demise of the typical career model in which young graduates are hired to do software tasks that can be automated and are slowly trained to gain that valuable tacit experience. The earn-while-you-learn model might finally be broken—at least for some occupations.

The simple truth could be that coding skills are no longer a guarantee of a job. That may help to explain the drop-off of computer science majors at schools around the country. Future canaries in the cubicles are sniffing out the dangers of looking for a job when their skills can be matched by AI.

But a closer look at the data shows that students are not necessarily turning away from AI-related careers. Rather, they appear to be tailoring their skills to the changes they see underway as AI becomes increasingly important for various disciplines. Interest is rising in AI-adjacent fields like data science and cybersecurity. One fast-growing major: artificial intelligence itself (a recent addition to many college offerings).

Is this time different?

Anxiety over the potential of AI to replace workers is nothing new. I wrote “How Technology Is Destroying Jobs” in 2013, describing how a slew of new digital technologies, including AI, were beginning to threaten white-collar work. I wasn’t alone. It was a popular theme at a time when the labor market was sluggish and jobs were scarce. 

In one of his last days in office in late 2016, President Obama issued a report written by his top economic and science advisors warning that AI was threatening workers. Among the findings was that automated vehicles—especially driverless trucks—could eliminate 2.2 million to 3.1 million existing US jobs.  Around the same time, one of the pioneers of AI, Geoffrey Hinton, said that “people should stop training radiologists” because it was “completely obvious” the occupation was soon to be replaced by AI.

None of these predictions came true, of course (nor did so-called technological unemployment occur during several earlier tech-related job panics). The forecasts were often wrong about the pace of the technological advances—we’re still waiting for fleets of driverless trucks on the highways—and failed to understand the complex portfolio of tasks that make up many jobs. AI has indeed become a tool for screening radiology images, but there are more radiologists than ever. It turns out that human radiologists perform a multitude of valuable tasks, including interpreting results and interacting with patients, that can’t be accomplished with AI (yet).

Perhaps this time is different, and we can put aside the lessons of economic history. Certainly, AI has gained unimaginable powers to do humanlike tasks. Perhaps it will devour jobs in ways that we’ve never seen before. And perhaps that will happen abruptly, without a warning buried in the labor statistics. But the previous bouts of AI job anxiety still hold a prescient lesson: Our real focus needs to be less on the dystopian fears and more on the very real transitions in the workplace that will likely affect millions of people.

“Even if there is not mass or even increased unemployment, the transition could still be very difficult,” says Jed Kolko, senior fellow at the Peterson Institute for International Economics and former undersecretary of commerce in the Biden administration. “And what does a difficult transition period mean? It means people losing jobs, or people’s jobs being redefined in ways that make those jobs pay worse or be less meaningful. And some people whose jobs are threatened may not be able to adapt.”

The more we understand this transition, the better prepared we’ll be to deal with it.  And for that we’ll need better and more complete data.

For McEntarfer, the former commissioner of the BLS, the real question is the speed of any disruption. “If it happens at the normal pace of technological change, labor markets will have time to adapt. If there is a sudden and severe disruption, then that will be a big challenge for policymakers,” she says. “That’s really the most important question facing us right now: how rapid this transformation is going to be.” And, she adds, “we’ll know by watching the data.”

Two decades ago, the country was caught flat-footed by the so-called China shock as free-trade policies led to an influx of imports and the devastation of manufacturing jobs in many parts of the country. It took years for researchers to understand the data showing how the trade policies, generally welcomed by economists, were destroying communities. Today the threat of an economic transformation brought on by AI is far larger and points to potentially far more damage for huge groups of workers.

To head off another devastating labor transition, we will need well-timed government and business policies, especially programs to train and reskill workers. If McEntarfer and other labor economists are correct, we probably have time to design deliberate and effective strategies to manage the transition. But first we need to better understand what is going on—and how fast.

It’s hard to find an economist who is more enthusiastic about AI’s future than Stanford’s Brynjolfsson, who believes that we’re likely on the brink of a huge boost that will transform the economy. “Perhaps the best productivity growth of my lifetime is coming up,” he says.

But Brynjolfsson also warns that a lack of data is severely limiting our visibility into the economic and societal impacts that are coming. At a time when hundreds of billions are being spent on rolling out the technology, he says, “we’re not investing even 1% of that on understanding the transition.”

  •  

BoxLitE: A Faithful Knowledge Base Embedding Based on Convex Optimization

arXiv:2605.23937v1 Announce Type: new Abstract: Knowledge base (KB) embeddings aim at combining the capability of classical knowledge graph embeddings to generalize the information present in facts, the ABox, with conceptual knowledge represented in an ontology language, the TBox. Several authors have recently explored the idea of mapping concepts to convex regions in a vector space. This is useful to represent hierarchies, typically present in TBoxes, since more general concepts can be mapped to larger regions, containing those regions associated with more specific concepts. However, the power of convexity is rarely leveraged during the actual learning tasks. Here, we introduce BoxLitE, a KB embedding model for DL-Lite$^{\mathcal{H}}$ that allows for convex optimization. We show that for any satisfiable DL-Lite$^{\mathcal{H}}$ KB, there is a BoxLitE embedding that is a weakly faithful model. As a proof of concept, we show how to formulate the KB embedding task as a convex optimization problem and how to obtain embeddings with such desirable faithfulness properties.
  •  

Inference Time Context Sparsity: Illusion or Opportunity?

arXiv:2605.24168v1 Announce Type: new Abstract: Sparsity has long been a central theme in LLM efficiency, but its role in context processing remains unresolved. As LLM workloads shift toward longer contexts and agentic interactions, the compute and memory bottlenecks of attention become increasingly critical, raising the question of whether these constraints are fundamental. Our position is that these constraints are artificial and unnecessary, and that the future of LLM inference lies in extreme but principled sparsity along the context dimension. This position is supported by several strands of empirical and theoretical evidence. First, we find the insistence on dense attention unreasonable, since in a long context a query effectively projects O(N) attention information into a hidden space of dimension d
  •  

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

arXiv:2605.24219v2 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination benchmarks still evaluate only the final output, missing failures that originate in intermediate Thought-Action-Observation steps. We present Trajel, a dataset and evaluation framework for auditing trajectory-level hallucinations in multi-agent industrial workflows. Trajel introduces a five-type hallucination taxonomy (factual, referential, logical, procedural, and scope-based) over expert-annotated agent traces from AssetOpsBench. We benchmark supervised detection models at the subtask, trajectory, and long-context levels. Our results show that the most common failure modes are missed by existing benchmarks, that nearly half of hallucinated trajectories involve multiple types at once, and that automated detectors with high binary accuracy still misclassify the subtlest types. Trajectory-aware detection significantly outperforms standard post-hoc verification, making taxonomy-grounded evaluation necessary for safer agentic deployment.
  •  

Lattice theory and algebraic models for deep convolutional learning based on mathematical morphology

arXiv:2605.24608v1 Announce Type: new Abstract: We develop a rigorous algebraic framework for deep convolutional architectures, CNNs, ResNets, and encoder--decoder networks such as UNet, grounded in lattice theory and mathematical morphology. The central tool is the Matheron--Maragos--Banon--Barrera (MMBB) universal representation theory for translation-invariant operators, which we apply systematically to every layer of a standard deep network. The principal finding is that the standard CNN pipeline (linear convolution~$+$ ReLU~$+$ flat max-pooling) is a cross-lattice operator: the convolution is an erosion in the Fourier inf-semilattice while ReLU is a lattice-join closing and max-pooling is a dilation in the pointwise max-plus lattice, and their composition is a morphological opening in neither. A second finding is that the upper adjoint of ReLU in the pointwise lattice is a global (non-local) operator, the identity on globally non-negative functions and $-\infty$ otherwise, so no local morphological erosion can form an adjunction pair with ReLU. These two results together provide the precise algebraic reason why depth in standard CNNs introduces genuine representational power: the composed layer is not idempotent. Three layer designs that are genuine idempotent openings are identified and fully characterised: the pure max-plus morphological layer (pointwise lattice), the spectral Wiener layer (Fourier lattice), and the self-dual morphological layer. We establish a complete fixed-point and convergence theory. The framework also unifies max-pooling, strided convolution, and the Laplacian pyramid under the Goutsias--Heijmans adjoint pyramid theory, and gives the Activation--Pooling Dilation (APD) factorisation with its correct adjoint.
  •  

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can come from architectural and engine-level design. We present MDIA, a Multi-agent Diagnostic Intelligence Agent implemented as a 7-node specialty-routed clinical reasoning graph, on the full HealthBench Professional benchmark (n = 525), on a non-fine-tuned LLM. MDIA achieves 0.6272 under OpenAI's GPT-5.4-2026-03-05, which is +3.72 pp above the performance of OpenAI's ChatGPT for Clinicians. The experimental work shows that performance lift is attributable to system architecture: specialty routing, multi-turn context preservation, drug-state safety gating, site-filtered search, length-aware synthesis, and engine-level reliability. These findings support the view that agentic clinical benchmark performance is shaped both by the underlying foundation model and the orchestration architecture. Nevertheless, we also noticed notable differences when using other models as a grader; in particular, when using Gemini 2.5 Pro, MDIA scored 0.6585, which suggests that the choice of grader is a source of variability. Robust evaluation of LLMs would therefore require assessment across several independent grader models.
  •  

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

arXiv:2605.25272v1 Announce Type: new Abstract: While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when rankings reflect genuine capability differences versus evaluation artifacts. We introduce a framework for measuring the latent landscape in AI benchmark ecosystems. Applying Confirmatory Factor Analysis (CFA) and Generalizability Theory to 4,000+ models from the Open LLM Leaderboard, we decompose sources of ranking variance and establish: (1) structures assumed in current reporting practice underestimate the strength of relationships between benchmarks; (2) evidence of local dependence among leaderboard items, undermining uses of benchmarks as measurement instruments under current scoring systems; (3) contributor metadata explains more rank-relevant variance ($\approx9\%$) than architecture or deployment categories in this context; (4) a manifest-score "scaling law" slope has low reliability ($R_{\beta}=0.53$); by contrast, the latent general-factor size slope is highly stable across ecosystem controls ($R_g=0.97$). We are able to provide unique insights into benchmark dynamics, such as which benchmarks are a function of LLM size and which can be oppositely impacted by post-training practices. We provide actionable diagnostics to determine how benchmark rankings can be trusted and how benchmark design can be improved.
  •  

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models

arXiv:2605.25394v1 Announce Type: new Abstract: Large language models often generate confident but incorrect answers rather than abstaining when uncertain. This problem is particularly acute for small language models (SLMs), where computational constraints and autonomous operation amplify the need for reliable uncertainty detection. We propose _Second Guess_, a lightweight, parameter-free prompting technique for abstention in multiple-choice question answering (MCQA) that is well-suited for SLMs. Our key empirical insight is that models which truly know an answer will select it consistently, while uncertain models exhibit unstable behavior when an ``I don't know'' option is added. Evaluated on four open models (2B-8B parameters) and four benchmarks, Second Guess achieves the highest composite risk improvement of 10.81\%. Notably, it maintains an 8\% composite risk improvement on fine-tuned models where entropy-based methods degrade, and improves most for lower-performing models. All code and results required to reproduce this work is available in https://github.com/Mystic-Slice/second-guess
  •  

Credit Assignment with Resets in Language Model Reasoning

arXiv:2605.25507v2 Announce Type: new Abstract: Contemporary reinforcement learning with verifiable reward methods post-train language models on multi-step reasoning by assigning a single outcome reward uniformly across all tokens in a trajectory. Such uniform assignment ignores which steps contributed to success or failure. Improving credit assignment can address this limitation by enabling targeted refinement of faulty reasoning steps, rather than updating entire trajectories uniformly. Resets are one such simple mechanism, enabling more precise credit assignment by returning to an intermediate state and resampling counterfactual continuations, so that outcome differences can be attributed to decisions made at that point. We propose two such methods: Random-Reset Policy Optimization (RRPO), where reset states are drawn randomly from reasoning steps, and Self-Reset Policy Optimization (SRPO), where the model self-localizes the erroneous step in an incorrect trajectory and resets there. We analyze these methods within the Conservative Policy Iteration (CPI) framework. Extending CPI with a credit-assignment oracle that targets improvable states yields provable improvements over random resets. Across models and reasoning benchmarks, SRPO consistently outperforms standard GRPO and RRPO by sampling multiple suffix continuations at a self-localized reset and learning from their rewards, using only the model itself with no external supervision.
  •  

FLOATBench: A Dataset and Benchmark for Floating Offshore Wind Turbine Tower Fatigue

arXiv:2605.25717v1 Announce Type: new Abstract: Most of the world's offshore wind resource lies in waters too deep for fixed-bottom foundations, making floating offshore wind turbines (FOWTs) essential for deep-water deployment. As the industry scales toward $22$ MW class designs, tower fatigue becomes increasingly critical because larger structures amplify the coupled aero-hydro-servo-elastic loads induced by continuous wind and wave excitation. Accurate fatigue-damage prediction is therefore central to certification, design optimization, and cost reduction. Yet the field lacks a shared surrogate benchmark: studies report different simulations, splits, and metrics, making methods difficult to compare. We present FLOATBench, a public tabular benchmark with $582{,}120$ per-section fatigue-damage labels across three $22$ MW FOWT tower geometries, derived from $19{,}404$ high-fidelity OpenFAST simulations across the three towers ($6{,}468$ per tower: $1{,}078$ aligned wind/wave operating points $\times$ six turbulence seeds), labeled at $30$ cross-sections per tower. FLOATBench includes a regime-aware alpha-shape partition of the joint wind/wave operating envelope, stratifying test points into in-train, interpolation, and extrapolation regimes. It is paired with a reproducible evaluation harness covering three protocol levels: random validation (E1), within-tower regime-aware evaluation (E2), and cross-tower transfer (E3). The regime-aware protocol reveals rank shifts between global and extrapolation performance that random-split leaderboards cannot detect. To the authors' knowledge, FLOATBench is the first FOWT fatigue benchmark for tabular surrogate modeling, and offers an evaluation protocol that generalizes to engineering surrogates defined over physical operating envelopes. Dataset and code available at: https://github.com/Joao97ribeiro/FLOATBench.
  •  

Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework

arXiv:2605.23921v1 Announce Type: cross Abstract: This study seeks to determine the authority signals used by Anthropic's Claude AI in its presentation of sources when answering consumer health questions. While there exists a great deal of discourse around the quality of health citations that LLMs produce, there is limited information on the integrity of the sources the citations originate from, and to what extent the sources are, from what health professionals would consider, credible sources. This descriptive cross-sectional study used data from HealthSearchQA, which contains 3,172 consumer health questions curated by Google Research. After exclusions, a final dataset of 3,075 questions yielding 10,038 citations was analyzed. The Authority Signals Framework (Jacques et al., 2026) was applied to examine 10 authority signals across four domains for a disproportionate stratified sample of 542 sources. Established institutional sources accounted for 97.8% of all citations (n = 9,818). Medical Institutions were the most frequently cited organization type (36.5%), followed by Government Resources (31.6%) and Professional Associations (28.4%). Commercial Health Information comprised 2.2% (n = 220). The top 10 organizations accounted for 57.8% of all citations, with Mayo Clinic alone representing 24.7%. Among commercial sources in the focused sample, 86.4% displayed medical review statements, 82.5% used schema markup, and 71.8% had comprehensive content, while traditional institutional sources appeared in Claude's citations with or without these same markers. As Anthropic positions Claude for HIPAA-ready healthcare applications, these findings establish a baseline for Claude's citation behavior and demonstrate the utility of the Authority Signals Framework as a tool for ongoing, cross-platform evaluation of AI-mediated health information.
  •  

Multi-market value-stacking: Battery control for combined imbalance participation and non-uniform FCR bidding

arXiv:2605.23964v1 Announce Type: cross Abstract: The growing share of Renewable Energy Sources (RES) in modern power systems increases both grid imbalances and frequency deviations, reinforcing the need for ancillary services such as Frequency Containment Reserve (FCR) and passive balancing. Battery Energy Storage Systems (BESS) are well-suited for these services, but prior research typically relies on uniform FCR bids that remain constant throughout the control period. Such static bids fail to fully exploit BESS flexibility, as they do not balance the trade-off between reserving energy for FCR delivery and using it for imbalance arbitrage, limiting the achievable value in value-stacking settings. To address this limitation, we propose a two-stage control framework for the European context that introduces non-uniform FCR bids. In the first stage, we derive a time-varying bid sequence using data-driven Monte Carlo (MC) optimization. In the second stage, a Deep Reinforcement Learning (DRL) agent leverages the residual flexibility for real-time imbalance trading while proactively managing the State of Energy (SoE) to ensure compliance with FCR requirements. The framework is presented as a proof of concept, highlighting the potential benefits of time-varying bidding strategies. By incorporating daily cycle budgets and time-varying reserve commitments, our approach achieves a 7.56% profit increase compared to uniform baselines. These results show that non-uniform bidding can unlock additional value by more effectively aligning reserve obligations with rapidly changing imbalance opportunities.
  •  

Harnessing AtomisticSkills for Agentic Atomistic Research

arXiv:2605.24002v1 Announce Type: cross Abstract: Computational materials science and chemistry span vast knowledge domains and fractured software ecosystems. Although large language models (LLMs) have demonstrated research capabilities, scaling monolithic agents to manage the rigor and complexity of atomistic research remains a challenge. Here, we introduce AtomisticSkills, an open-source harness framework that empowers general-purpose AI coding agents to conduct atomistic research across materials science, chemistry, and drug discovery. By hierarchically decomposing scientific workflows into agent skills and tools, AtomisticSkills provides agents with modular, extensible, and plug-and-play research capabilities. The framework integrates more than 100 human-curated multidisciplinary skills, including database access, thermodynamics and kinetics modeling, and diverse simulation engines employing machine learning interatomic potentials (MLIPs) and density functional theory (DFT). We validate its functional coverage against scientific literature and demonstrate robust orchestration capabilities across diverse scientific campaigns: generative design of Li-ion solid-state electrolytes, high-throughput screening of metal-organic frameworks for CO2 capture, autonomous MLIP benchmarking and fine-tuning, multi-stage structure-based virtual screening for drug design, multimodal X-ray diffraction pattern analysis, and screening of Fe-oxide catalysts for oxygen evolution reaction. AtomisticSkills provides a critical agent infrastructure towards building fully autonomous AI scientists.
  •  

Verified SHAP: Provable Bounds for Exact Shapley Values of Neural Networks

arXiv:2605.24084v1 Announce Type: cross Abstract: Shapley additive explanations (SHAP) are widely recognised as computationally intractable for neural networks, since they induce an exponential search space over the input features. In this work, we take a first step towards scaling exact SHAP computation to larger search spaces by introducing an algorithm that leverages recent advances in neural network verification to compute arbitrarily tight exact lower and upper bounds on SHAP values for neural networks, ultimately recovering the exact SHAP values. We demonstrate that our approach scales to orders of magnitude larger search spaces than state-of-the-art exact methods. This provides an important first step towards exact SHAP computation and establishes a principled cornerstone for evaluating statistical approximation methods on larger search spaces.
  •  

MASt3R-Nav: WayPixel Navigation in Relative 3D Maps

arXiv:2605.24111v1 Announce Type: cross Abstract: Visual navigation ability is strongly tied to its underlying representation of the world. Unlike classical 3D maps that require globally-consistent geometry, image- or object-relative topological graphs almost entirely do away with geometric understanding. But, this comes at the cost of navigation capability, often limiting it to merely teach-and-repeat. In this work, we propose a novel map representation in the form of pixel-relative connectivity, which is geometrically accurate but does not require global geometric consistency. Inspired by recent progress in 3D grounded image matching, we construct a map from an image sequence through inter-image connectivity based on pixel correspondences in the relative 3D coordinate systems of individual image pairs. We then use this pixel-level graph to perform global path planning by approximating and sparsifying intra-image pixel connectivity. Through this, we derive a ''WayPixel Costmap'' representation and train a controller conditioned on it to predict a trajectory rollout. We show that this dense pixel-level costmap based on relative geometry is a more accurate conditioning variable for control prediction than its image- and object-level counterparts. This enables a highly capable navigation system, as validated on four types of navigation tasks in the simulator and through real world demonstrations.
  •  

PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection

arXiv:2605.24171v1 Announce Type: cross Abstract: Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacterized. We present PromptAudit, a controlled evaluation framework that isolates prompt effects by fixing the dataset, decoding, and parsing while varying only the prompting strategy. Using five prompting strategies across five open-weight models on 1,000 CVEs (6,074 code samples spanning 16 programming languages), we evaluate accuracy, recall, abstention, coverage, and effective F1. We find that standard chain-of-thought prompting achieves the strongest overall operational performance, while few-shot prompting provides model-dependent benefits that are most pronounced for prompt-sensitive models. In contrast, adaptive chain-of-thought frequently suppresses recall and self-consistency induces excessive abstention, sharply reducing effective performance. These results show that vulnerability detection behavior is jointly determined by the model and the prompt, and that prompt sensitivity is a first-class system property that must be explicitly characterized in evaluation and deployment.
  •  

GIBLy: Improving 3D Semantic Segmentation through an Architecture-Agnostic Lightweight Geometric Inductive Bias Layer

arXiv:2605.24243v1 Announce Type: cross Abstract: In 3D scene understanding, deep learning models rely on large models and extensive training to capture basic geometric structures that are present in the 3D data. However, existing methods lack explicit mechanisms to incorporate geometric information, such as learnable primitive shapes, often necessitating large models and more training data which in turn increases cost and can limit generalization. We introduce GIBLy, a lightweight geometric inductive bias layer that integrates learnable geometric priors into 3D segmentation pipelines. GIBLy enhances existing architectures -- whether MLP-based, convolution-based, or transformer-based -- by providing features aligned with simple geometric shapes (and thus human-interpretable) that improve segmentation performance with minimal computational overhead. We validate our approach across multiple 3D semantic segmentation benchmarks, demonstrating consistent performance gains, including up to +11.5% mIoU on TS40K with PTV3, while adding only 58K extra parameters. Our results highlight the benefit of explicitly encoding geometric structure to support accurate and efficient 3D scene understanding, with a lightweight add-on layer
  •  

Concept Drift Adaptation Using Self-Supervised and Reinforcement Learning In Android Malware Detection

arXiv:2605.24294v1 Announce Type: cross Abstract: Android malware detectors often degrade after deployment because of concept drift, while full retraining at each maintenance step is costly. We propose a chronological adaptive maintenance framework that models deployment-time maintenance as a sequential decision problem. The framework learns a stable latent representation through self-supervised learning during initialization, freezes the encoder, measures latent drift in the fixed representation space, and performs lightweight downstream adaptation using a trainable adapter and classification head. A proximal policy optimization controller selects low-cost maintenance actions based on the detector state, including current utility, retention on a fixed memory set, latent drift indicators, and update cost. We evaluate the framework under a causal deployment-style protocol on emulator and real Android malware datasets with static and dynamic features. Results show that the RL controller provides a strong cost-aware adaptation strategy, consistently remaining among the top-performing policies while achieving a favorable balance between temporal performance, memory retention, and maintenance cost under non-stationary deployment conditions.
  •  
❌