Bielik-7B Polish Law Fine-tune (QLoRA)
Fine-tuned version of speakleash/Bielik-7B-Instruct-v0.1 on Polish legal Q&A data using QLoRA (4-bit) with FlashAttention 2.
Model Details
| Property |
Value |
| Base model |
speakleash/Bielik-7B-Instruct-v0.1 |
| Architecture |
Mistral-7B |
| Fine-tuning method |
QLoRA (4-bit NF4) |
| Adapter type |
LoRA |
| Language |
Polish |
| Domain |
Legal / Polish law |
Training Hardware
| Component |
Spec |
| GPU |
NVIDIA RTX 3060 12 GB |
| VRAM usage |
~10 GB (with FlashAttention 2) / ~10–11 GB (SDPA) |
| CUDA |
12.1 |
Training Parameters
LoRA Configuration
| Parameter |
Value |
r (rank) |
16 |
lora_alpha |
32 |
lora_dropout |
0.05 |
bias |
none |
target_modules |
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
task_type |
CAUSAL_LM |
Quantization (BitsAndBytes)
| Parameter |
Value |
| Quantization |
4-bit (load_in_4bit=True) |
| Quant type |
NF4 |
| Compute dtype |
bfloat16 |
| Double quantization |
Yes (saves ~0.4 GB VRAM) |
SFT / Training Config
| Parameter |
Value |
| Epochs |
3 |
| Per-device batch size |
2 |
| Gradient accumulation steps |
4 (effective batch size: 8) |
| Learning rate |
2e-4 |
| LR scheduler |
cosine |
| Warmup ratio |
0.05 |
| Max sequence length |
2048 |
| Precision |
bf16 |
| TF32 |
Yes (Ampere GPU benefit) |
| Optimizer |
paged_adamw_8bit |
| Gradient checkpointing |
Yes (use_reentrant=False) |
| Group by length |
Yes |
| Dataloader workers |
4 |
| Seed |
42 |
| Attention implementation |
FlashAttention 2 (flash_attention_2) |
Software Environment
| Library |
Version |
| PyTorch |
2.5.1+cu121 |
| Transformers |
4.47.0 |
| PEFT |
0.14.0 |
| TRL |
0.13.0 |
| BitsAndBytes |
0.45.0 |
| Datasets |
3.2.0 |
| Accelerate |
1.2.1 |
Dataset
Custom Polish legal Q&A dataset. Each sample is a single text field formatted as an instruction/response pair in Polish.