TechCrunch Host
speaker
5,755 appearances
109 recordings
1 series
first heard Mar 2026
last heard 17 Aug
TechCrunch Host’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 109 in all, peaking in Jul 2026 with 23.
Appearances
And companies have little incentive to make those investments until something goes wrong.
matched
Bitterman said, I think the companies are not willing to extend the resources that are required to accomplish sufficient guardrails, and probably will not until they're
matched
Forced to.
matched
But there is another issue at hand.
matched
If they lock a model down too tight during testing, researchers might fail to discover capabilities before the model is released.
matched
This is just as dangerous, possibly more so than giving it too much freedom.
matched
And then the evaluation itself risks becoming the problem.
matched
So can safety evaluations be regulated?
matched
Well, the Trump administration is currently weighing a voluntary predeployment cybersecurity evaluation regime under which the government will get to assess the security risks of new powerful models thirty days before they are released publicly.
matched
The policy, the product of a Trump executive order which has been finalized behind closed doors, would not address safety evaluation incidents because they occur farther upstream of deployment.
matched
Yoon said the lesson that we've been learning in the last few months is that self-regulatory apparatus is just not enough anymore.
matched
There are competitive pressures that are incentivizing a race to the bottom on safety standards, and that is a perfect place for regulatory intervention.
matched
He went on to say, what we would need to cover this is some kind of controls on what's happening inside the labs while the models are being developed, both at the training stage and at the testing stage.
matched
The challenge is only likely to grow as the models do.
matched
A source familiar with irregular evaluations told TechCrunch that more capable models require more complex evaluations, often conducted quickly and at greater scale,
matched
Which opens the door for more mistakes.
matched
AISI, which intentionally gives some models internet access, told TechCrunch it's reviewing the balance between realistic testing and managing the risks those tests create.
matched
OpenAI said it's reviewing how it conducts third-party testing, as well as requirements around isolation, monitoring, and when evaluations should be stopped.
matched
Meta said it's still investigating the incident and plans to publish a retrospective once it has all the facts.
matched
In the end,
matched
Showing 321–340 of 5,755 · page 17 of 288
← Previous
Next →