Safe and effective post-fine-tuning alignment in large language models

Minrui Jiang, Yuning Yang, Xiurui Xie, Pei Ke, Guisong Liu. Safe and effective post-fine-tuning alignment in large language models. Knowl.-Based Syst., 330:114523, 2025. [doi]

Authors

Minrui Jiang

This author has not been identified. Look up 'Minrui Jiang' in Google

Yuning Yang

This author has not been identified. Look up 'Yuning Yang' in Google

Xiurui Xie

This author has not been identified. Look up 'Xiurui Xie' in Google

Pei Ke

This author has not been identified. Look up 'Pei Ke' in Google

Guisong Liu

This author has not been identified. Look up 'Guisong Liu' in Google