Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
arXiv:2511.02805v1 Announce Type: cross Abstract: Typical search agents concatenate the entire interaction history into the LLM context, preserving information integrity but producing long, noisy contexts, resulting in high computation and memory costs. In contrast, using only the current turn avoids this overhead but discards essential information. This trade-off limits the scalability of search agents. To address this challenge, we propose MemSearcher, an agent workflow that iteratively maint
-
cs.AI, q-bio.NC updates on arXiv.org
-
Tongyi DeepResearch Technical Report
arXiv:2510.24701v1 Announce Type: cross Abstract: We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data syn
Tongyi DeepResearch Technical Report
-
MIT Technology Review
-
A Chinese firm has just launched a constantly changing set of AI benchmarks
When testing an AI model, it’s hard to tell if it is reasoning or just regurgitating answers from its training data. Xbench, a new benchmark developed by the Chinese venture capital firm HSG, or HongShan Capital Group, might help to sidestep that issue. That’s thanks to the way it evaluates models not only on the ability to pass arbitrary tests, like most other benchmarks, but also on the ability to execute real-world tasks, which is more unusual. It will be updated on a regular basis to try to
A Chinese firm has just launched a constantly changing set of AI benchmarks
When testing an AI model, it’s hard to tell if it is reasoning or just regurgitating answers from its training data. Xbench, a new benchmark developed by the Chinese venture capital firm HSG, or HongShan Capital Group, might help to sidestep that issue. That’s thanks to the way it evaluates models not only on the ability to pass arbitrary tests, like most other benchmarks, but also on the ability to execute real-world tasks, which is more unusual. It will be updated on a regular basis to try to keep it evergreen.
This week the company is making part of its question set open-source and letting anyone use for free. The team has also released a leaderboard comparing how mainstream AI models stack up when tested on Xbench. (ChatGPT o3 ranked first across all categories, though ByteDance’s Doubao, Gemini 2.5 Pro, and Grok all still did pretty well, as did Claude Sonnet.)
Development of the benchmark at HongShan began in 2022, following ChatGPT’s breakout success, as an internal tool for assessing which models are worth investing in. Since then, led by partner Gong Yuan, the team has steadily expanded the system, bringing in outside researchers and professionals to help refine it. As the project grew more sophisticated, they decided to release it to the public.
Xbench approached the problem with two different systems. One is similar to traditional benchmarking: an academic test that gauges a model’s aptitude on various subjects. The other is more like a technical interview round for a job, assessing how much real-world economic value a model might deliver.
Xbench’s methods for assessing raw intelligence currently include two components: Xbench-ScienceQA and Xbench-DeepResearch. ScienceQA isn’t a radical departure from existing postgraduate-level STEM benchmarks like GPQA and SuperGPQA. It includes questions spanning fields from biochemistry to orbital mechanics, drafted by graduate students and double-checked by professors. Scoring rewards not only the right answer but also the reasoning chain that leads to it.
DeepResearch, by contrast, focuses on a model’s ability to navigate the Chinese-language web. Ten subject-matter experts created 100 questions in music, history, finance, and literature—questions that can’t just be googled but require significant research to answer. Scoring favors breadth of sources, factual consistency, and a model’s willingness to admit when there isn’t enough data. A question in the publicized collection is “How many Chinese cities in the three northwestern provinces border a foreign country?” (It’s 12, and only 33% of models tested got it right, if you are wondering.)
On the company’s website, the researchers said they want to add more dimensions to the test—for example, aspects like how creative a model is in its problem solving, how collaborative it is when working with other models, and how reliable it is.
The team has committed to updating the test questions once a quarter and to maintain a half-public, half-private data set.
To assess models’ real-world readiness, the team worked with experts to develop tasks modeled on actual workflows, initially in recruitment and marketing. For example, one task asks a model to source five qualified battery engineer candidates and justify each pick. Another asks it to match advertisers with appropriate short-video creators from a pool of over 800 influencers.
The website also teases upcoming categories, including finance, legal, accounting, and design. The question sets for these categories have not yet been open-sourced.
ChatGPT-o3 again ranks first in both of the current professional categories. For recruiting, Perplexity Search and Claude 3.5 Sonnet take second and third place, respectively. For marketing, Claude, Grok, and Gemini all perform well.
“It is really difficult for benchmarks to include things that are so hard to quantify,” says Zihan Zheng, the lead researcher on a new benchmark called LiveCodeBench Pro and a student at NYU. “But Xbench represents a promising start.”
-
Nature - Issue - nature.com science feeds
-
NBS1 lactylation is required for efficient DNA repair and chemotherapy resistance
Nature, Published online: 03 July 2024; doi:10.1038/s41586-024-07620-9Lactylation of NBS1 by TIP60 promotes homologous recombination-driven DNA repair and resistance to chemotherapy in cancer cells and links altered cancer cell metabolism to increase genome stability.
NBS1 lactylation is required for efficient DNA repair and chemotherapy resistance
Nature, Published online: 03 July 2024; doi:10.1038/s41586-024-07620-9
Lactylation of NBS1 by TIP60 promotes homologous recombination-driven DNA repair and resistance to chemotherapy in cancer cells and links altered cancer cell metabolism to increase genome stability.-
Nature - Issue - nature.com science feeds
-
Author Correction: A genomic mutational constraint map using variation in 76,156 human genomes
Nature, Published online: 15 January 2024; doi:10.1038/s41586-024-07050-7Author Correction: A genomic mutational constraint map using variation in 76,156 human genomes
Author Correction: A genomic mutational constraint map using variation in 76,156 human genomes
Nature, Published online: 15 January 2024; doi:10.1038/s41586-024-07050-7
Author Correction: A genomic mutational constraint map using variation in 76,156 human genomes-
Cell
-
Highly multiplexed bioactivity screening reveals human and microbiota metabolome-GPCRome interactions
PRESTO-Salsa is a multiplexed GPCR screening platform that enables the interrogation of GPCRome-wide bioactivities across diverse sample types. This platform was used to explore microbiota-induced GPCR activation and revealed the engagement of an immune cell-associated GPCR by a periodontal pathogen protease.
Highly multiplexed bioactivity screening reveals human and microbiota metabolome-GPCRome interactions
-
Most Recent Articles: Clinical Epigenetics
-
Predicting disease-free survival in colorectal cancer by circulating tumor DNA methylation markers
Recurrence represents a well-known poor prognostic factor for colorectal cancer (CRC) patients. This study aimed to establish an effective prognostic prediction model based on noninvasive circulating tumor DNA...
Predicting disease-free survival in colorectal cancer by circulating tumor DNA methylation markers
-
Cell Death Discovery nature.com science feeds
-
Cancer-associated fibroblasts promote the stemness and progression of renal cell carcinoma via exosomal miR-181d-5p
Cell Death Discovery, Published online: 01 November 2022; doi:10.1038/s41420-022-01219-7Cancer-associated fibroblasts promote the stemness and progression of renal cell carcinoma via exosomal miR-181d-5p
Cancer-associated fibroblasts promote the stemness and progression of renal cell carcinoma via exosomal miR-181d-5p
Cell Death Discovery, Published online: 01 November 2022; doi:10.1038/s41420-022-01219-7
Cancer-associated fibroblasts promote the stemness and progression of renal cell carcinoma via exosomal miR-181d-5p