Steve Gibson

speaker
17,616 appearances 14 recordings 1 series first heard May 2026 last heard 2 Sep

Steve Gibson’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
5 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 14 in all, peaking in Jul 2026 with 5.

Appearances

newest first · ▶ plays the moment
Some of that knowledge is dual use, meaning it can be used for good or bad.
For example, knowledge of cybersecurity can help patch critical security vulnerabilities, or it can be used to exploit them.
Knowledge of virology can help a researcher create a vaccine, but it can also help a malicious actor design a deadly pathogen.
Ideally, we would be able to balance three separate goals.
First, limiting access to dual use capabilities in as surgical a way as possible.
Second, allowing trusted users to access those same capabilities for beneficial purposes.
And third, doing all this without affecting
the model's performance on any other task.
Current safeguards are imperfect, they wrote.
We train models to refuse harmful requests and use classifiers to screen inputs and outputs for dangerous content.
These layers of protection guard against dangerous outputs, but they don't change the knowledge stored in the underlying model.
Despite our safeguards, a sufficiently determined attacker may still try to jailbreak the model, working past its defenses to access the dual-use knowledge.
A more robust protection against misuse would be to control what the model knows.
We've explored this before, they wrote.
In earlier work, we filtered information about chemical, biological, radiological, and nuclear weapons out of pre-training data.
And later showed that dual use knowledge can be confined to a remote to a removable slice of a model's weights.
But filtering is a blunt instrument, it produces one model with one fixed set of capabilities.
Using filtering, if you want a model version that can discuss it.
Advanced virology for deployment in a vetted biosecurity lab, say, and another version that cannot discuss that because it doesn't have the knowledge, they say you have to train two separate models.
Especially in the case of frontier models, which are large and very expensive to train, the cost to the developer would be prohibitive.
Showing 2161–2180 of 17,616 · page 109 of 881 ← Previous Next →