❌

Normal view

  • ✇MIT Technology Review
  • A Chinese firm has just launched a constantly changing set of AI benchmarks Caiwei Chen
    When testing an AI model, it’s hard to tell if it is reasoning or just regurgitating answers from its training data. Xbench, a new benchmark developed by the Chinese venture capital firm HSG, or HongShan Capital Group, might help to sidestep that issue. That’s thanks to the way it evaluates models not only on the ability to pass arbitrary tests, like most other benchmarks, but also on the ability to execute real-world tasks, which is more unusual. It will be updated on a regular basis to try to
     

A Chinese firm has just launched a constantly changing set of AI benchmarks

23 June 2025 at 23:46

When testing an AI model, it’s hard to tell if it is reasoning or just regurgitating answers from its training data. Xbench, a new benchmark developed by the Chinese venture capital firm HSG, or HongShan Capital Group, might help to sidestep that issue. That’s thanks to the way it evaluates models not only on the ability to pass arbitrary tests, like most other benchmarks, but also on the ability to execute real-world tasks, which is more unusual. It will be updated on a regular basis to try to keep it evergreen. 

This week the company is making part of its question set open-source and letting anyone use for free. The team has also released a leaderboard comparing how mainstream AI models stack up when tested on Xbench. (ChatGPT o3 ranked first across all categories, though ByteDance’s Doubao, Gemini 2.5 Pro, and Grok all still did pretty well, as did Claude Sonnet.) 

Development of the benchmark at HongShan began in 2022, following ChatGPT’s breakout success, as an internal tool for assessing which models are worth investing in. Since then, led by partner Gong Yuan, the team has steadily expanded the system, bringing in outside researchers and professionals to help refine it. As the project grew more sophisticated, they decided to release it to the public.

Xbench approached the problem with two different systems. One is similar to traditional benchmarking: an academic test that gauges a model’s aptitude on various subjects. The other is more like a technical interview round for a job, assessing how much real-world economic value a model might deliver.

Xbench’s methods for assessing raw intelligence currently include two components: Xbench-ScienceQA and Xbench-DeepResearch. ScienceQA isn’t a radical departure from existing postgraduate-level STEM benchmarks like GPQA and SuperGPQA. It includes questions spanning fields from biochemistry to orbital mechanics, drafted by graduate students and double-checked by professors. Scoring rewards not only the right answer but also the reasoning chain that leads to it.

DeepResearch, by contrast, focuses on a model’s ability to navigate the Chinese-language web. Ten subject-matter experts created 100 questions in music, history, finance, and literature—questions that can’t just be googled but require significant research to answer. Scoring favors breadth of sources, factual consistency, and a model’s willingness to admit when there isn’t enough data. A question in the publicized collection is “How many Chinese cities in the three northwestern provinces border a foreign country?” (It’s 12, and only 33% of models tested got it right, if you are wondering.)

On the company’s website, the researchers said they want to add more dimensions to the test—for example, aspects like how creative a model is in its problem solving, how collaborative it is when working with other models, and how reliable it is.

The team has committed to updating the test questions once a quarter and to maintain a half-public, half-private data set.

To assess models’ real-world readiness, the team worked with experts to develop tasks modeled on actual workflows, initially in recruitment and marketing. For example, one task asks a model to source five qualified battery engineer candidates and justify each pick. Another asks it to match advertisers with appropriate short-video creators from a pool of over 800 influencers.

The website also teases upcoming categories, including finance, legal, accounting, and design. The question sets for these categories have not yet been open-sourced.

ChatGPT-o3 again ranks first in both of the current professional categories. For recruiting, Perplexity Search and Claude 3.5 Sonnet take second and third place, respectively. For marketing, Claude, Grok, and Gemini all perform well.

“It is really difficult for benchmarks to include things that are so hard to quantify,” says Zihan Zheng, the lead researcher on a new benchmark called LiveCodeBench Pro and a student at NYU. “But Xbench represents a promising start.”

  • ✇MIT Technology Review
  • Scaling integrated digital health MIT Technology Review Insights
    Around the world, countries are facing the challenges of aging populations, growing rates of chronic disease, and workforce shortages, leading to a growing burden on health care systems. From diagnosis to treatment, AI and other digital solutions can enhance the efficiency and effectiveness of health care, easing the burden on straining systems. According to the World Health Organization (WHO), spending an additional $0.24 per patient per year on digital health interventions could save more than
     

Scaling integrated digital health

Around the world, countries are facing the challenges of aging populations, growing rates of chronic disease, and workforce shortages, leading to a growing burden on health care systems. From diagnosis to treatment, AI and other digital solutions can enhance the efficiency and effectiveness of health care, easing the burden on straining systems. According to the World Health Organization (WHO), spending an additional $0.24 per patient per year on digital health interventions could save more than two million lives from non-communicable diseases over the next decade.

To work most effectively, digital solutions need to be scaled and embedded in an ecosystem that ensures a high degree of interoperability, data security, and governance. If not, the proliferation of point solutions— where specialized software or tools focus on just one specific area or function—could lead to silos and digital canyons, complicating rather than easing the workloads of health care professionals, and potentially impacting patient treatment. Importantly, technologies that enhance workforce productivity should keep humans in the loop, aiming to augment their capabilities, rather than replace them. 

Through a survey of 300 health care executives and a program of interviews with industry experts, startup leaders, and academic researchers, this report explores the best practices for success when implementing integrated digital solutions into health care, and how these can support decision-makers in a range of settings, including laboratories and hospitals. 


Key findings include: 


Health care is primed for digital adoption. The global pandemic underscored the benefits of value-based care and accelerated the adoption of digital and AI-powered technologies in health care. Overwhelmingly, 96% of the survey respondents say they are “ready and resourced” to use digital health, while one in four say they are “very ready.” However, 91% of executives agree interoperability is a challenge, with a majority (59%) saying it will be “tough” to solve. Two in five leaders say balancing security with usability is the biggest challenge for digital health. With the adoption of cloud solutions, organizations can enjoy the benefits of modernized IT infrastructure: 36% of the survey respondents believe scalability is the main benefit, followed by improved security (28%). 

Digital health care can help health care institutions transform patient outcomes—if built on the right foundations. Solutions like AI-powered diagnostics, telemedicine, and remote monitoring can offer measurable impact across the patient journey, from improving early disease detection to reducing hospital readmission rates. However, these technologies can only support fully connected health care when scaled up and embedded in ecosystems with robust data governance, interoperability, and security. 

Health care data has immense potential—but fragmentation and poor interoperability hinder impact. Health care systems generate vast quantities of data, yet much of it remains siloed or unusable due to inconsistent formats and incompatible IT systems, limiting scalability. 

Digital tools must augment, not overload, the workforce. With global health care workforce shortages worsening, digital solutions like clinical decision support tools, patient prediction, and remote monitoring can be seen as essential aids rather than threats to the workforce. Successful deployment depends on usability, clinician engagement, and training. 

Regulatory evolution, open data policies, and economic sustainability are key to scaling digital health. Even the best digital tools struggle to scale without reimbursement frameworks, regulatory support, and viable business models. Open data ecosystems are needed to unleash the clinical and economic value of innovation. Regulatory and reimbursement innovation is also critical to transitioning from pilot projects to high-impact, system-wide adoption.

Download the full report.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff.

This content was researched, designed, and written entirely by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Want to know where VCs are investing next? Be in the room at TechCrunch Disrupt 2025

23 June 2025 at 22:30
Early-stage founders, listen up! You will want a front row seat at the Builders Stage on October 27 at 1:00 p.m. PT. This session at TechCrunch Disrupt 2025 brings together Nina Achadjian, partner, Index Ventures; Jerry Chen, general partner, Greylock; and Viviana Faga, general partner, Felicis, all of whom will share their 2026 investment priorities […]
❌