Ev Williams

speaker
383 appearances 2 recordings 2 series first heard Jun 2026 last heard 18 Jun

Ev Williams’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
2 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Jun 2026 with 2.

Appearances

newest first · ▶ plays the moment
where you could say, oh, well, that's like a negative for Anthropic in the labs and IPO prospects and everything like that.
The massive positive from this is this idea that with the right amount of test time compute, basically the right amount of inference thrown at a given query into a model or a harness with an open source model, you can replicate these results.
That would be a sign of a test time computer.
So with like, Noam Brown tweeted about this recently where he said, the way we think about these benchmark cards, you know, these scorecards where it's like, this model has larger numbers than the last model, aka good.
He's like, it's all wrong to think about because the models actually perform very, very differently if you just continue to throw more compute at it at test time, at inference time.
And so the really, really good thing
for I think all the frontier labs and just the inference market in general, companies like ours, like Fireworks that run inference platforms for their enterprise customers on open models, all of these companies that have to do that are in the token path is that it's very clear that if you just throw more test time compute at any frontier level model, you continue to get results.
they haven't found where the wall is, where you stop getting better results, the more test time compute you throw at it.
And so, there's been all of these like step function increases in how much compute AI models use.
We went from like very simple auto-aggressive, next token completion to these agentic models that do chain of thought and like agents use an order of magnitude more inference.
And now you have this whole thing around, well, if you wanna have mythos like performance,
Just spend a lot more compute, you know, do a lot more inference, spend a lot more tokens.
And so I think that the other side of this is that even though we've had such an insane growth in the amount of tokens process for the industry over the last three years, we actually might see another kink in the curve as more test time compute creates more and more of these unbelievable results for capabilities and things that you can do with models, whether or not they're, you know, Fable or Mythos or whatever OpenAI comes out with, but open source models as well.
Zooming out, I think the most important thing in all of this is that we step back from Anthropic and realize that the models are at a place now, whether they're from Anthropic or elsewhere, where they can autonomously find and chain multiple vulnerabilities together and orchestrate and attack autonomously.
If they do get good enough where they can take down parts of digital infrastructure that run the US economy or Western economies, then there is some concern about a nation state adversary like a North Korea or China or Russia having those capabilities.
So it's a real conversation that we're going to have to have as Western democracy very, very soon.
And we probably should have already had it.
And Anthropic right now is the poster boy for it because of the relationship that they have with the administration and the things that they've said.
But this is true for AI now.
This is an AI discussion.
Showing 141–160 of 383 · page 8 of 20 ← Previous Next →