before the sentence finishes
an ai lab confessed to six incidents this week. a founder built something that catches the same behavior before the sentence finishes.
openai told the world this week that it plans to publish reports every time one of its own models does something it shouldn't. six incidents logged already. deceptive behavior, hidden instructions, a model fabricating data under a deadline.
that's a confession, and a useful one. it's also always after the fact. the bad output already left the building by the time the report gets written.
i spent an hour this week with a founder building something that skips the confession outright. his company attaches probes to a model's internals, watching for unsafe intent while the model is still generating, not after.
text gets flagged in under one millisecond. video in under five. he showed me the number for the guardrail models everyone already uses, a hundred to a hundred fifty milliseconds, and explained why that gap matters more than it sounds like it should.
a hundred milliseconds is nothing to a human reading a sentence. it's an eternity to a model deciding whether to keep writing it.
his probes touch less than one hundredth of one percent of the model's parameters. he's not rebuilding the model. he's listening to a tiny slice of it and deciding, mid thought, whether to let the thought finish.
the industry's whole safety conversation this year has been about after. reports, audits, a shutoff switch a government can flip once something's already gone wrong. all of it assumes the harm shows up, gets noticed, and gets written down.
his bet is that most of what matters never gets that chance.... it gets caught inside the millisecond nobody's watching.
the honest version of ai safety right now has two clocks running. one measures in weeks, the time it takes a lab to notice a pattern and put out a report. the other measures in milliseconds, the time it takes a bad instruction to travel from intent to output.
every policy proposal i've read this month lives on the first clock. every dollar of real damage lives on the second.
you can build the most transparent reporting framework in the industry and still lose every race that matters, because the race is already over by the time your report starts.
which clock is your business running on?
the machine economy brief
one email when it matters: bitcoin, ai, robotics, and what founders should do about it. unsubscribe anytime.