In this episode of the Sunny Ray Show, host Sunny Ray reconnects with Yan, founder of Sepion.ai, for an unplanned deep dive into what Yan calls a wisdom layer for enterprise AI. Yan explains that while building a product intelligence platform for product and customer experience teams, his team accidentally solved a much bigger two trillion dollar problem: enterprise data is not ready for AI, and existing rag and knowledge graph architectures still force agents to consume information in inefficient chunks. Sepion's engine instead cleans, organizes, and synthesizes all of a company's scattered data automatically, then hands agents already distilled answers rather than raw chunks, cutting token usage from millions to tens of thousands in benchmark tests. The conversation ranges into AI development curves, with Yan predicting full AGI around late 2027 or early 2028 based on comparing AI progress to human cognitive development, rising enterprise token costs ahead of anticipated Anthropic and OpenAI IPOs, and who might eventually want to acquire a company built around this kind of foundational efficiency layer.
Yan reveals how his startup accidentally cut AI agent token usage by over 99 percent while building for product teams.
What are you building and why should people care?
We set out to build a collective neural network, starting with a product intelligence platform for product teams since I spent 25 years in product design. To make our platform fully autonomous and promptless, we built proprietary tech that ended up solving a much bigger enterprise problem around getting data ready for AI, which is roughly a two trillion dollar problem industry wide.
So you weren't even trying to solve the token efficiency problem. The savings were a side effect of building for product teams?
Exactly. Our users just plug their data in, they don't organize or clean anything, our core engine does all of that. It cleans, structures, triangulates, validates, resolves conflicts, and builds automatic ontologies into what I call a wisdom layer. Once agents plug into that layer, in our last benchmark we hit 99.7 percent token efficiency compared to traditional approaches.
So the agents don't even need chunking or rag, the engine just knows what to pull? Is this basically killing the pre-processing industry?
Not fully, those architectures are still valid and we can even plug into them. But instead of organizing data into chunks an agent has to go search through, we essentially build a librarian for your data. Our system has already read everything and gives agents synthesized answers directly. In our benchmark, hybrid rag used over eight million tokens on a task where we used only 24,000.
You call it wisdom, not knowledge or context. What's the actual difference?
Context is a summary of something you read, and knowledge is the deep expertise you build over time, like my 25 years as a mechanical engineer would give me. Wisdom is different because I'm not just handing over data or pointing to where it lives, I'm giving the already synthesized, condensed answer. There's no fluff, no need to reason through it, it's straight to the point.
You said Frontier Labs will be selling time in five years. Is the real race no longer intelligence but who compresses the distance between question and answer?
Right, I think an AI year equals about 10 to 12 human years. Development felt like huge leaps from late 2022 through 2024, but since early 2025 it's incremental gains, similar to how human brains stop major development around age 30. I predict full AGI around late 2027 or early 2028, and after that the race becomes about speed and specificity, not raw intelligence.
If your engine works as you say, the data prep industry evaporates. Who fights you hardest?
I think it would be a frontier lab, probably Anthropic or OpenAI, trying to acquire us rather than fight us. Enterprises are trying to solve three things right now: keeping their data private instead of handing it fully to AI providers, making that data clean and ready for AI, and making agents efficient enough that they don't cost a fortune to run.
Does the wisdom layer become AI's memory, and can AGI even work without something like it?
That's a great question, I honestly haven't thought about it from that angle before. What I do know is our engine does three things for enterprise: protects their data, makes it AI ready, and reduces the cost of running agents. If a lab like Anthropic or OpenAI acquired us, maybe they'd bake this in as a unified layer their models sit on top of.
Building something daring? Sunny talks to founders like this every day. Fifteen minutes to see if your story belongs on the stage.
Claim your pre-interview