I don’t think there will be strong+working mech interp in <1 year. Model internals methods could pareto dominate cot monitors in a year via cot losing ~all value, but this isn’t exactly encouraging. Depending on AIs decoding other AI’s opaque activations for oversight is spooky
despite so many of my friends doing it, I paid not too much attention to the field (though I still tried to pay a bit of attention)
Stephen L Casper did a talk about his skepticism of interpretability.