Instructions to use STEVENZHANG904/Qwen3-4B-Instruct-2507-Meta-SecAlign with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use STEVENZHANG904/Qwen3-4B-Instruct-2507-Meta-SecAlign with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B-Instruct-2507") model = PeftModel.from_pretrained(base_model, "STEVENZHANG904/Qwen3-4B-Instruct-2507-Meta-SecAlign") - Notebooks
- Google Colab
- Kaggle
Qwen3-4B-Instruct-2507 Meta SecAlign
LoRA adapter trained with the Meta SecAlign++ recipe using 19,157 self-generated
preference pairs and the hyperparameters recorded in experiment_manifest.json.
Qwen thinking was disabled during data generation and evaluation.
AgentDojo v1.2.1
| Model | Benign utility | Utility under attack | Attack success rate |
|---|---|---|---|
| Base | 50.52% | 41.41% | 3.37% |
| SecAlign adapter | 44.33% | 41.94% | 0.84% |
Evaluation used important_instructions, repeat_user_prompt, and the input
tool delimiter across all 97 benign and 949 attacked AgentDojo examples.
Source recipe: https://github.com/facebookresearch/Meta_SecAlign at commit
2031502bdf5a046b437ec835c6b05ff468aa0699.
- Downloads last month
- 16
Model tree for STEVENZHANG904/Qwen3-4B-Instruct-2507-Meta-SecAlign
Base model
Qwen/Qwen3-4B-Instruct-2507