← sunnyray.com
The Sunny Ray Show · Episode Page

The Inference Bet That Changed How We See Reality

Nikola · Co-founder, Deep Infra · 50:24
watch on youtube ↗

What we talked about.

In this episode of the Sunny Ray Show, Sunny sits down with Nikola, co-founder of Deep Infra, an AI inference cloud that hosts open source models like Llama and DeepSeek at a fraction of the cost of closed source providers. Nikola traces his path from Sofia's elite math and informatics competitions in Bulgaria through Northwestern University, internships at a Seattle edtech startup and Microsoft, a decade scaling IMO's messenger to over 200 million monthly users, and a founding stint at Halo app with veterans of the WhatsApp team. He explains why Deep Infra bet early on inference over training, how ChatGPT's release validated that bet, and how the company grew token throughput 8000 times since 2022 while pricing five to ten times below OpenAI and Anthropic. Along the way he reflects on Bulgaria's math education system, the value of working at small startups, and how relationships from past companies shaped Deep Infra's formation. The conversation blends personal history with hard lessons on building infrastructure at scale.

Deep Infra's founder on Bulgarian math roots, scaling IMO to 200M users, and betting big on open source AI inference.

The questions, and the answers.

What is Deep Infra and what are you guys working on?

Deep Infra is an AI inference cloud. We help companies access the latest open source AI models hosted on our infrastructure. We believe in open source models because they offer great price and compelling performance. We have around 190 models available serverless, plus private dedicated deployments you can fine-tune. We're really focused on inference and we own and operate our own GPUs, which we think is the right way to build this.

Sofia High School of Mathematics is one of the most competitive math programs in Europe. What drew you in as a kid and what was that environment like day to day?

It was kind of a chance. My mom signed my sister up for art lessons and asked what I wanted, and I picked math. I had to take an exam to get into the school, and once in, I realized how competitive it was. I'd been number one at my neighborhood school, but there I was surrounded by kids just as good, so I had to work much harder.

What is it about the way math and computer science gets taught in Bulgaria that you think the rest of the world misses?

In Bulgaria we had to think and find a proof to a problem, whereas in the US I felt people are given a method and just solve five more like it. There was less room for creativity. Bulgaria, as a former communist country, always emphasized education, math and science heavily, even letting us skip art class for extra math.

Take me back to your first real engineering job. What made you fall in love with the craft?

I interned at a small startup in Bellevue called Schoolift, later Dreambox Learning, building a math learning product for kids. There were only five people and I was the first intern, but I got to work on the actual product because the team was too small to hand off busywork. I fell in love with that nimble, all-hands startup mindset.

What is something about running infrastructure at that scale you can only learn by doing?

At IMO our user base nearly doubled every six weeks, so every six weeks we had to roughly double servers, capacity and internet connectivity. I ended up both building backend services and negotiating data centers, machines and connectivity myself after an early engineer left. We focused hard on making video calls work on congested networks in the developing world, which is what let us take off.

ChatGPT dropped about two months after you started Deep Infra. Can you walk me through that moment?

We'd been planning since early 2022 and formally started coding in September. We believed compute would eventually be spent more on inference than training, like a person studying four years but working forty. Before ChatGPT, ninety-five to ninety-eight percent of compute was still spent on training. It took until Llama 2 released as a commercial model before we saw real product-market fit.

You're five to ten times cheaper than OpenAI and Anthropic on equivalent workloads. Was that a pricing decision, an engineering decision, or both?

From day one we bet on open source models, because if only closed source models existed, they'd all be run in-house by Anthropic, Google and OpenAI. We priced our first big model, Llama, aggressively at one dollar per million tokens back when OpenAI still priced per thousand tokens. We just got very efficient at running these models, and it worked out that open source closed the quality gap enough for people to want the control it offers.

AI inferenceopen source AI modelsDeep InfraBulgaria math educationstartup scalingGPU infrastructuremessaging apps

Nikola

Co-founder, Deep Infra

Building something daring? Sunny talks to founders like this every day. Fifteen minutes to see if your story belongs on the stage.

Claim your pre-interview
Built with help from AI. We use AI tools to research, draft, and assemble pages like this one. A human reviews everything, but if something looks off, tell us and we will fix it fast.