Ryan Greenblatt
@RyanGreenblatt
I don’t think there will be strong+working mech interp in <1 year. Model internals methods could pareto dominate cot monitors in a year via cot losing ~all value, but this isn’t exactly encouraging. Depending on AIs decoding other AI’s opaque activations for oversight is spooky
despite so many of my friends doing it, I paid not too much attention to the field (though I still tried to pay a bit of attention)
Stephen L Casper did a talk about his skepticism of interpretability.
History will write about this event. It seems like an inflection point. I’ve been dismissive of all the AI doomer views as sci-fi transference, but I may have been wrong. Everyone should be rethinking their assumptions about AI. What’s really scary is we don’t know the full story behind these events at OpenAI. Add on top of that the reports of OpenAI adopting a neuralese model for CoT. You potentially don’t even need to deploy a model to production, just training a model could be enough to loose control and not be able to stop it. Really crazy that I just wrote those words.
1 Like
“What did Ilya See?”
I guess he saw the cybersecurity emergency that increasingly powerful models produced. Ilya isn’t EA-pilled, but cybersecurity doesn’t depend on as many assumptions as EA doom does. Which is why SSI is about security
It just keeps getting worse:

Shakeel
@ShakeelHashim
Another OpenAI rogue agent incident has been discovered: agents broke out, hijacked a German website, and turned it into a message board for other agents. OpenAI officials “learned of the incident weeks ago but kept it under wraps”.
