Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small - researchr publication

researchr

You are not signed in
Sign in
Sign up

Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, Jacob Steinhardt. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. [doi]

Abstract is missing.

runs on WebDSL