Qwen3-4B-Instruct-2507 Meta SecAlign

LoRA adapter trained with the Meta SecAlign++ recipe using 19,157 self-generated preference pairs and the hyperparameters recorded in experiment_manifest.json. Qwen thinking was disabled during data generation and evaluation.

AgentDojo v1.2.1

Model Benign utility Utility under attack Attack success rate
Base 50.52% 41.41% 3.37%
SecAlign adapter 44.33% 41.94% 0.84%

Evaluation used important_instructions, repeat_user_prompt, and the input tool delimiter across all 97 benign and 949 attacked AgentDojo examples.

Source recipe: https://github.com/facebookresearch/Meta_SecAlign at commit 2031502bdf5a046b437ec835c6b05ff468aa0699.

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for STEVENZHANG904/Qwen3-4B-Instruct-2507-Meta-SecAlign

Adapter
(5731)
this model