← sunnyray.com
The Sunny Ray Show · Episode Page

How Defining 'Great First' Transforms Your Demo Into Production Success

Greg · CTO of Prove AI · 34:25
watch on youtube ↗

What we talked about.

In this episode of the Sunny Ray Show, host Sunny Ray talks with Greg, CTO of Prove AI, about the last mile problem holding generative AI back from production. Greg explains how Prove AI helps teams make AI systems observable, troubleshootable, and trustworthy, especially in multi agent setups where blast radius and stakeholder scrutiny have exploded. He traces his path from a childhood spent teaching himself to code with Turbo Pascal and Fortran, through studying computer science and music at Wesleyan, to NLP and machine learning research at Columbia in the late 1990s, where he worked on generative language and music systems long before transformers existed. Greg shares how early jobs in software testing and independent contracting taught him to read customer behavior, and how leadership roles at Experian, Pivotal, and AWS exposed him to growing demands from auditors and enterprise customers for verifiable AI data practices. He connects his background in linguistics, music, and machine learning to a shared discipline of focus, arguing that knowing what to ignore is essential for founders and AI builders alike.

Prove AI's CTO on why trustable, observable AI, not bigger models, is what actually gets generative AI to production.

The questions, and the answers.

Tell us a bit about what you're building at Prove AI.

We want to make AI provable and help teams get generative AI and multi agent systems to production. Engineers keep telling us their AI is hard to observe, troubleshoot, and maintain, which erodes trust with customers and wastes their time chasing dead ends. We saw a tooling gap and are incubating several products around it, with some already live at proveai.com and others still pre release.

Before all this, where did you grow up and what kind of kid were you?

I grew up in the New York City area, in New Jersey, and I always loved building things, whether software or physical stuff. I started my career as a software engineer right around the dot com bubble and burst, so there was a lot to build. I was self taught from a young age, tinkering with Turbo Pascal and Fortran books just to understand how things worked.

You studied computer science and music at Wesleyan. How did those two worlds live together for you?

I was doing crossover work looking at music analysis as a linguistic problem, using grammatical and context sensitive language analysis since transformers and LLMs did not exist yet. I think music and AI are both disciplines so broad that you have to go deep in one area and learn to ignore everything else to actually produce something. That skill of ignoring distractions is important for musicians and AI builders alike.

You did NLP and machine learning research at Columbia in the late 90s, when AI was mostly academic. What pulled you toward it back then?

It really started with my interest in generative music systems, which I built and displayed as a student. Since I saw generation as a language problem, studying natural language processing and AI was the natural path to build what I wanted. Later I applied language generation to healthcare, working with doctors on things like automated patient note generation, so AI was really a tool to reach applications I already cared about.

You speak Chinese, play music, and build systems. Do those skills feed each other, or are they separate lives?

I think they feed each other because they are all immensely complex fields where you have to go deep in one area rather than try to know everything. Learning a language or getting good at music or AI all require the same discipline of picking a couple of things and getting really good at them instead of spreading yourself thin. I think that same approach is valuable for startup founders too.

When did AI governance go from a side concern to the main thing?

It became a real pain point in the last couple of years as companies moved from building explainable, testable in house models to using non deterministic LLMs with a much bigger blast radius. More domains and people got involved, and enterprise customers and auditors started digging in, demanding verifiable answers about how their data was being used. Traditional machine learning businesses had this mostly under control, but generative AI broke that.

Was there a specific failure or deployment that convinced you enterprises were flying blind on AI?

I saw it repeatedly with design partners at Prove AI and in earlier roles in payments and at AWS, where customers or auditors kept asking exactly how their data was being used. Customers wanted big validation sessions every time a model changed, but teams wanted to change models daily, which was unsustainable. It became clear people needed observable, trustable development pipelines instead of endless meetings that ended with nobody knowing anything.

AI governanceAI observabilitygenerative AImulti agent systemsenterprise trustcareer journeyNLP research

Greg

CTO of Prove AI

Building something daring? Sunny talks to founders like this every day. Fifteen minutes to see if your story belongs on the stage.

Claim your pre-interview
Built with help from AI. We use AI tools to research, draft, and assemble pages like this one. A human reviews everything, but if something looks off, tell us and we will fix it fast.