Steve Gibson

speaker
17,616 appearances 14 recordings 1 series first heard May 2026 last heard 2 Sep

Steve Gibson’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
5 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 14 in all, peaking in Jul 2026 with 5.

Appearances

newest first · ▶ plays the moment
Specifically for each model, we find that a single direction
Such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions.
While adding this direction elicits refusal on even harmless instructions.
They said, leveraging this insight, we propose a novel white box jailbreak method that surgically disables refusal with minimal effect on other capabilities.
Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal mediating.
direction.
Our findings underscore the brittleness, and this is the key, the brittleness of current safety fine-tuning methods.
In other words, just instructing the model not to answer the question, well, that works.
But if the weights are open, it turns out to be trivial to remove those instructions even after the fact and not having known what the instructions were.
I'll I'll explain a little more.
It's amazing.
So they they finish saying more broadly.
Our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior.
Okay, so what this group discovered was that
Any late-term model behavior imprinting.
can later be removed from such models.
And the way this is done, as I it's as I said, it's wonderful and wild.
They compare the models activations on harmful versus harmless prompts.
Compute the mean difference.
then project the weight matrices orthogonal to that direction.
Showing 2061–2080 of 17,616 · page 104 of 881 ← Previous Next →