Models run by the garage-inference stacks: LLM serving, speech-to-text and voice cloning sharing one GPU. Book and code on GitHub.
WΞNDΞL
bitdeep
AI & ML interests
Inference engineering: self-hosted LLMs, speech recognition, voice cloning and GPU lifecycle
Recent Activity
updated a collection 2 days ago
Garage Inference: self-hosted AI on one 24 GB GPU updated a collection 2 days ago
Garage Inference: self-hosted AI on one 24 GB GPU updated a collection 2 days ago
Garage Inference: self-hosted AI on one 24 GB GPUOrganizations
None yet