← sunnyray.com
The Sunny Ray Show · Episode Page

Why 80% of AI Work Is Hidden Data Prep — The Game-Changer No One Talks About

David · ML Twist (data preparation for AI) · 17:52
watch on youtube ↗

What we talked about.

In this episode of The Sunny Ray Show, host Sunny Ray talks with David, whose company ML Twist helps businesses prepare data for AI systems. David traces his path from a chemical engineering start to computer engineering after acing a Fortran class, then into adtech through an acquisition by DoubleClick, which was later bought by Google. He shares stories from his time managing accounts at Google, working on early mobile and streaming ad innovation, and later roles at MarketShare, New Star, TransUnion, and Oracle, where he saw firsthand how much unseen work goes into preparing data for data scientists. David explains that AI models almost always break at the data layer, not the code, and uses a racecar and fuel analogy to describe how ML Twist turns raw, messy data into AI ready fuel. He closes by explaining why companies bringing AI into their core products need to take data preparation seriously, and teases a possible part two conversation with Sunny.

ML Twist's David explains why the hidden, unglamorous work of prepping data is the real bottleneck for AI success.

The questions, and the answers.

Can you give a quick overview of what ML Twist does?

We help companies get their data ready for AI. People assume it is just get data, train models, tune the model, deploy. But that first step, getting the data ready, can actually be twenty different steps if you want really high quality data. Our software automates that step.

What first pulled you toward computers and tech growing up?

I was definitely the kid who liked computers. My dad had gotten an Apple 2e, which I remember fondly, and I played a lot of games on it before it evolved through the 360 and other systems. In college I started as a chemical engineering major, but the first A I ever got was in a Fortran programming class, which pushed me toward computer engineering instead.

Do you think being a self-proclaimed nerd, into things like Magic the Gathering and Dungeons and Dragons, actually shaped how you approach problem solving and data today?

It's tough to say categorically that it had an influence, since some people in my computer engineering program came from that background and others did not. What it did give me was a lot of alone time, and because computers fascinated me so much, that time went into things like installing new motherboards and really understanding the nuts and bolts of how computers work.

How did you go from engineering into the business and data side of adtech through the DoubleClick acquisition?

I graduated during the dotcom bust when it was hard to get a programming job, so I went to France, worked odd jobs including bartending, and landed a role helping open the US market for a company building an early linguistics model. After stints in New York and a move to the UK, I joined an early adtech company as an engineer, and shortly after we were acquired by DoubleClick, which was later bought by Google.

What was it like being inside Google after the DoubleClick acquisition?

Titles were fluid and kept changing through the acquisition, but essentially I was an account manager working with companies using DoubleClick to pass ad data back and forth. Adtech was one of the earliest innovators in mobile ads, streaming ads, and even ads in gaming, and working with those accounts gave me my first real taste of data for dashboards and eventually predictive models for ad performance.

When did you realize the technology had caught up to the AI hype and machine learning was ready for prime time?

The real light bulb moment was a couple years ago when one of our investors told me I should change the company name to AI Twist instead of ML Twist. Ironically, at conferences today many speakers are still careful to say ML, since a lot of what gets called AI is really machine learning like image recognition or NLP. GPT gave it a huge credibility bump, but that comment was when I felt AI had truly been embraced and was now expected.

You've said data scientists spend up to 80 percent of their time preparing data rather than building models. When did that fully hit you?

It was gradual, built up across roles at MarketShare, New Star, TransUnion, and Oracle, where I felt like a glorified data delivery boy for data scientists. Even after delivering data, it still had to go through more processing before it fit the models being built. It is nobody's fault, data is just fluid and constantly needs restating, which is part of what makes data science fascinating.

You've described ML Twist as the refinery that turns oil into high octane fuel for AI. Can you break that down simply?

Think of companies building the race cars and companies building the fuel. The car makers usually know what fuel they need but are not fuel specialists themselves. Scale AI is a good example, since Meta put fourteen billion dollars into a forty nine percent stake at a twenty eight billion dollar valuation rather than build that data prep capability alone. If someone has your gas covered, you can focus your time and talent on winning the race.

AI data preparationmachine learningadtechDoubleClick and GoogleML Twistdata science careersAI hype vs reality

David

ML Twist (data preparation for AI)

Building something daring? Sunny talks to founders like this every day. Fifteen minutes to see if your story belongs on the stage.

Claim your pre-interview
Built with help from AI. We use AI tools to research, draft, and assemble pages like this one. A human reviews everything, but if something looks off, tell us and we will fix it fast.