What makes and breaks safety fine-tuning? a mechanistic study

Jain, S; Lubana, ES; Oksuz, K; Joy, T; Sanyal, A; Torr, P; Dokania, PK

Conference item

What makes and breaks safety fine-tuning? a mechanistic study

Abstract:: Safety fine-tuning helps align Large Language Models (LLMs) with human preferences for their safe deployment. To better understand the underlying factors that make models safe via safety fine-tuning, we design a synthetic data generation framework that captures salient aspects of an unsafe input by modeling the interaction between the task the model is asked to perform (e.g., "design") versus the specific concepts the task is asked to be performed upon (e.g., a "cycle" vs. a "bomb"). Using this, we investigate three well-known safety fine-tuning methods---supervised safety fine-tuning, direct preference optimization, and unlearning---and provide significant evidence demonstrating that these methods minimally transform MLP weights to specifically align unsafe inputs into its weights' null space. This yields a clustering of inputs based on whether the model deems them safe or not. Correspondingly, when an adversarial input (e.g., a jailbreak) is provided, its activations are closer to safer samples, leading to the model processing such an input as if it were safe. We validate our findings, wherever possible, on real-world models---specifically, Llama-2 7B and Llama-3 8B.

Publication status:: Published

Peer review status:: Peer reviewed

Actions

Email

Email this record

Send the bibliographic details of this record to your email address.

Your Email
Please enter the email address that the record information will be sent to.

-
Your message (optional)
Please add any additional information to be included within the email.
Cite

Cite this record

APA Style

Jain, S., Lubana, E. S., Oksuz, K., Joy, T., Sanyal, A., Torr, P., & Dokania, P. K. (2024). What makes and breaks safety fine-tuning? a mechanistic study. Proceedings of the Mechanistic Interpretability Workshop 2024 Hosted by the 13th International Conference on Machine Learning (ICML 2024).

MLA Style

Jain, S., et al. “What Makes and Breaks Safety Fine-Tuning? a Mechanistic Study.” Proceedings of the Mechanistic Interpretability Workshop 2024 Hosted by the 13th International Conference on Machine Learning (ICML 2024), OpenReview, 2024.

Chicago Style

Jain, S, ES Lubana, K Oksuz, T Joy, A Sanyal, P Torr, and PK Dokania. 2024. “What Makes and Breaks Safety Fine-Tuning? a Mechanistic Study.” In Proceedings of the Mechanistic Interpretability Workshop 2024 Hosted by the 13th International Conference on Machine Learning (ICML 2024). OpenReview.
Share
Print

Access Document

Files:: Jain_et_al_2024_What_makes_and.pdf

(Preview, Version of record, pdf, 6.5MB, Terms of use)

Publication website:: https://openreview.net/forum?id=BS2CbUkJpy

Authors

+ Jain, S More by this author

Role:: Author

+ Lubana, ES More by this author

Role:: Author

+ Oksuz, K More by this author

Role:: Author

+ Joy, T More by this author

Role:: Author

+ Sanyal, A More by this author

Role:: Author

More authors...

+ Engineering and Physical Sciences Research Council More from this funder

Funder identifier:: https://ror.org/0439y7842
Grant:: EP/W002981/1

Publisher:: OpenReview
Host title:: Proceedings of the Mechanistic Interpretability Workshop 2024 hosted by the 13th International Conference on Machine Learning (ICML 2024)
Publication date:: 2024-07-31
Acceptance date:: 2024-05-02
Event title:: Mechanistic Interpretability Workshop 2024 hosted by the 13th International Conference on Machine Learning (ICML 2024)
Event location:: Vienna, Austria
Event website:: https://icml2024mi.pages.dev/
Event start date:: 2024-07-27
Event end date:: 2024-07-27

Language:: English
Keywords:: AI safety

mechanistic interpretability

Safety fine tuning
Pubs id:: 2036881
Local pid:: pubs:2036881
Deposit date:: 2024-10-07

Terms of use

Copyright holder:: Jain et al.
Notes:: This paper was presented at the Mechanistic Interpretability Workshop 2024 hosted by the 13th International Conference on Machine Learning (ICML 2024), 27th July 2024, Vienna, Austria.

Licence:: CC Attribution (CC BY)

Views and Downloads

About views and downloads

If you are the owner of this record, you can report an update to it here: Report update to this record

Conference item

What makes and breaks safety fine-tuning? a mechanistic study

Actions

Access Document

Authors

Terms of use

Views and Downloads

Altmetrics

Dimensions

Conference item

What makes and breaks safety fine-tuning? a mechanistic study

Actions

Access Document

Authors

Funding

Bibliographic Details

Item Description

Terms of use

Metrics

Views and Downloads

Altmetrics

Dimensions