Probing Classifiers are Unreliable for Concept Removal and Detection

Abhinav Kumar, Chenhao Tan, Amit Sharma 0007. Probing Classifiers are Unreliable for Concept Removal and Detection. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022. 2022. [doi]

Authors

Abhinav Kumar

This author has not been identified. Look up 'Abhinav Kumar' in Google

Chenhao Tan

This author has not been identified. Look up 'Chenhao Tan' in Google

Amit Sharma 0007

This author has not been identified. Look up 'Amit Sharma 0007' in Google