Khushi-Business-AI-4bit (7B)

A 4-bit quantized, business-focused conversational AI built on Mistral 7B using Unsloth. Optimized for low VRAM (runs on 6GB+ VRAM) and for Indian business use-cases.

Built with ❤️ by Kartik Sharma | UX4567 17+ models | 1200+ downloads | Low-VRAM Specialist

Why this model?

  • Made with Unsloth: 2x faster fine-tuning & quantization
  • 7B params in 4-bit (BF16 base): ~4GB RAM me chal jayega
  • Business Tuned: Customer support, sales chat, lead handling
  • Bilingual: Hindi + English (Hinglish) samajhta hai
  • Two formats: safetensors + GGUF for Ollama / LM Studio

How to use (Transformers + Unsloth)

from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name = "UX4567/Khushi-Business-AI-4bit",
    max_seq_length = 2048,
    load_in_4bit = True,
)

prompt = "Ek customer bol raha hai: Mujhe refund chahiye, kya reply du?"
inputs = tokenizer([prompt], return_tensors = "pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens = 200)
print(tokenizer.batch_decode(outputs)[0])
Downloads last month
271
Safetensors
Model size
7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for UX4567/Khushi-Business-AI-4bit

Quantized
(111)
this model