SLM Service Multinicho

A 4B parameter Small Language Model fine-tuned for human-quality customer service in Portuguese. Built with LoRA/QLoRA on Qwen3-4B-Instruct using Unsloth, optimized for consultative conversations, needs discovery, and empathetic support.

Base Model

Qwen3-4B-Instruct-2507 — a lightweight instruction-following model with strong multilingual performance.

Training Dataset

~10,000 synthetic examples generated by GPT-4o-mini following a structured prompt covering:

  • Consultative conversations (5,000)
  • Objection handling (2,000)
  • Memory & context continuity (2,000)
  • Bad vs ideal responses (1,000)

Categories include indecisive, busy, angry, curious, price inquiries, scheduling, complaints, and more. All examples are in Brazilian Portuguese.

Training Method

Fine-tuned with LoRA / QLoRA via Unsloth:

Hyperparameter Value
LoRA rank 16
LoRA alpha 16
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Sequence length 2048
Batch size 2 (gradient accumulation 4)
Learning rate 2e-4
Epochs 3
Optimizer AdamW 8-bit
Scheduler Cosine
Precision BF16 / FP16

4-bit QLoRA (default) or full LoRA — both supported via --qlora / --lora flags.

Results are merged to 16-bit for inference and GGUF export.

Intended Use

  • Customer service chatbots in Portuguese
  • Consultative sales conversations
  • First-contact qualification
  • Follow-ups and appointment scheduling
  • Complaint handling with empathy

Not intended for: general-purpose chat, factual question answering, code generation, or languages other than Portuguese.

Quantization

The uploaded GGUF file is f16 (16-bit float), preserving full fine-tuned quality.

  • File: slm-customer-service-f16.gguf

How to Use with Ollama

# Download the GGUF and create a Modelfile:
FROM ./slm-customer-service-f16.gguf
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER num_predict 200

# Import and run:
ollama create slm-service -f Modelfile
ollama run slm-service

How to Use with llama.cpp

./main -m slm-customer-service-f16.gguf \
  -p "<|im_start|>user\nQuanto custa o seguro?<|im_end|>\n<|im_start|>assistant\n" \
  --temp 0.7 \
  --top-p 0.9 \
  -n 200

Chat Format

This model uses the ChatML format:

<|im_start|>system
You are a helpful customer service assistant.<|im_end|>
<|im_start|>user
Olá, gostaria de saber mais sobre o seguro.<|im_end|>
<|im_start|>assistant
Claro! Vou ficar feliz em ajudar. Você já tem uma ideia do tipo de cobertura que procura?<|im_end|>

Links

Downloads last month
16
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for buildsource/slm-service

Adapter
(5826)
this model