AI & ML interests
One-click deployment of Open-source LLMs, on managed and dedicated GPUs.
Recent Activity
Organizations
Deploying Qwen3.8-2.4T-A95B with vLLM: Verified GPU Pods, Quants, and Serving Recipes
Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 314 tok/s Decode Generation [Benchmark]