Researchers are working on mechanistic interpretability of transformer language models, which can help understand their behavior and identify potential issues. They have developed a method to reverse-engineer transformers by breaking them down into human-interpretable pieces, such as attention heads and MLP layers.