❌

Reading view

Martin Frederik, Snowflake: Data quality is key to AI-driven growth

As companies race to implement AI, many are finding that project success hinges directly on the quality of their data. This dependency is causing many ambitious initiatives to stall, never making it beyond the experimental proof-of-concept stage.

So, what’s the secret to turning these experiments into real revenue generators? AI News caught up with Martin Frederik, regional leader for the Netherlands, Belgium, and Luxembourg at data cloud giant Snowflake, to find out.

“There’s no AI strategy without a data strategy,” Frederik says simply. “AI apps, agents, and models are only as effective as the data they’re built on, and without unified, well-governed data infrastructure, even the most advanced models can fall short.”

Improving data quality is key to AI project success

It’s a familiar story for many organisations: a promising proof-of-concept impresses the team but never translates into a tool that makes the company money. According to Frederik, this often happens because leaders treat the technology as the end goal.

Headshot of Martin Frederik, regional leader for the Netherlands, Belgium, and Luxembourg at AI data cloud giant Snowflake.

“AI is not the destination – it’s the vehicle to achieving your business goals,” Frederik advises.

When projects get stuck, it’s usually down to a few common culprits: the project isn’t truly aligned with what the business needs, teams aren’t talking to each other, or the data is a mess. It’s easy to get disheartened by statistics suggesting that 80% of AI projects don’t reach production, but Frederik offers a different perspective. This isn’t necessarily a failure, he suggests, but “part of the maturation process”.

For those who get the foundation right, the payoff is very real. A recent Snowflake study found that 92% of companies are already seeing a return on their AI investments. In fact, for every £1 spent, they’re getting back £1.41 in cost savings and new revenue. The key, Frederik repeats, is having a “secure, governed and centralised platform” for your data from the very beginning.

It’s not just about tech, it’s about people

Even with the best technology, an AI strategy can fall flat if the company culture isn’t ready for it. One of the biggest challenges is getting data into the hands of everyone who needs it, not just a select few data scientists. To make AI work at scale, you have to build strong foundations in your “people, processes, and technology.”

This means breaking down the walls between departments and making quality data and AI tools accessible to everyone.

“With the right governance, AI becomes a shared resource rather than a siloed tool,” Frederik explains. When everyone works from a single source of truth, teams can stop arguing about whose numbers are correct and start making faster and smarter decisions together.

The next leap: AI that reasons for itself

The true breakthrough we’re seeing now is the emergence of AI agents that can understand and reason over all kinds of data at once regardless of structure quality; from the neat rows and columns in a spreadsheet, to the unstructured information in documents, videos, and emails. Considering that this unstructured data makes up 80-90% of a typical company’s data, this is a huge step forward.

New tools are enabling staff, no matter their technical skill level, to simply ask complex questions in plain English and get answers directly from the data.

Frederik explains that this is a move towards what he calls “goal-directed autonomy”. Until now, AI has been a helpful assistant you had to constantly direct. “You ask a question, you get an answer; you ask for code, you get a snippet,” he notes.

The next generation of AI is different. You can give an agent a complex goal, and it will figure out the necessary steps on its own, from writing code to pulling in information from other apps to deliver a complete answer. This will automate the most time-consuming parts of a data scientist’s job, like “tedious data cleaning” and “repetitive model tuning.”

The result? It frees up your brightest minds to focus on what really matters. This elevates your people “from practitioner to strategist” and allows them to drive real value for the business. That can only be a good thing.

Snowflake is a key sponsor of this year’s AI & Big Data Expo Europe and will have a range of speakers sharing their deep insights during the event. Swing by Snowflake’s booth at stand number 50 to hear more from the company about making enterprise AI easy, efficient, and trusted.

See also: Public trust deficit is a major hurdle for AI growth

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Martin Frederik, Snowflake: Data quality is key to AI-driven growth appeared first on AI News.

  •  

Fine-Tuning Methods for Large Language Models in Clinical Medicine by Supervised Fine-Tuning and Direct Preference Optimization: Comparative Evaluation

Background: Large language model (LLM) fine tuning is the process of adjusting out-of-the-box model weights using a dataset of interest. Fine tuning can be a powerful technique to improve model performance in fields like medicine, where data access is restricted and LLMs may have poor out-of-the-box performance. Objective: In this study we investigated the benefits of fine tuning with supervised fine tuning (SFT) and direct preference optimization (DPO) across a range of LLM applications for medicine Methods: We use Llama3 7B and Mistral 7B v2 to compare the performance of SFT and DPO across four datasets for common natural language tasks in medicine. The tasks evaluated were simple classification, clinical reasoning, summarization, and clinical triage. Results: Clinical Reasoning accuracy increased 8% and 7% with DPO over SFT for Llama3 (p value 0.003) and Mistral2 (p value 0.004) respectively. Summarization quality, graded on a five point Likert scale, increased 0.13 and 0.10 for Llama3 and Mistral2 (p values
  •  

Comparative Evaluation of a Medical Large Language Model in Answering Real-World Radiation Oncology Questions: Multicenter Observational Study

Background: Large language models (LLMs) hold promise for supporting clinical tasks, particularly in data-driven and technical disciplines such as radiation oncology. While prior evaluation studies have focused on examination-style settings for evaluating LLMs, their performance in real-life clinical scenarios remains unclear. In the future, LLMs might be used as general AI assistants to answer questions arising in clinical practice. It is unclear how well a modern LLM, locally executed within the infrastructure of a hospital, would answer such questions compared with clinical experts. Objective: This study aimed to assess the performance of a locally deployed, state-of-the-art medical LLM in answering real-world clinical questions in radiation oncology compared with clinical experts. The aim was to evaluate the overall quality of answers, as well as the potential harmfulness of the answers if used for clinical decision-making. Methods: Physicians from 10 departments of European hospitals collected questions arising in the clinical practice of radiation oncology. Fifty of these questions were answered by 3 senior radiation oncology experts with at least 10 years of work experience, as well as the LLM Llama3-OpenBioLLM-70B (Ankit Pal and Malaikannan Sankarasubbu). In a blinded review, physicians rated the overall answer quality on a 5-point Likert scale (quality), assessed whether an answer might be potentially harmful if used for clinical decision-making (harmfulness), and determined if responses were from an expert or the LLM (recognizability). Comparisons between clinical experts and LLMs were then made for quality, harmfulness, and recognizability. Results: There were no significant differences between the quality of the answers between LLM and clinical experts (mean scores of 3.38 vs 3.63; median 4.00, IQR 3.00-4.00 vs median 3.67, IQR 3.33-4.00; P=.26; Wilcoxon signed rank test). The answers were deemed potentially harmful in 13% of cases for the clinical experts compared with 16% of cases for the LLM (P=.63; Fisher exact test). Physicians correctly identified whether an answer was given by a clinical expert or an LLM in 78% and 72% of cases, respectively. Conclusions: A state-of-the-art medical LLM can answer real-life questions from the clinical practice of radiation oncology similarly well as clinical experts regarding overall quality and potential harmfulness. Such LLMs can already be deployed within the local hospital environment at an affordable cost. While LLMs may not yet be ready for clinical implementation as general AI assistants, the technology continues to improve at a rapid pace. Evaluation studies based on real-life situations are important to better understand the weaknesses and limitations of LLMs in clinical practice. Such studies are also crucial to define when the technology is ready for clinical implementation. Furthermore, education for health care professionals on generative AI is needed to ensure responsible clinical implementation of this transforming technology.
  •  
❌