lyogavin/airllm — 23,706 Stars
AirLLM 70B inference with single 4GB GPU. Contribute to lyogavin/airllm development by creating an account on GitHub.
Watch Episode
About This Repo#
lyogavin/airllm — 23,706 ⭐
AirLLM 70B inference with single 4GB GPU. Contribute to lyogavin/airllm development by creating an account on GitHub.
Narration#
AirLLM. Run 70 billion parameter models on a 4 gigabyte GPU. Twenty-three thousand stars. It compresses and shards massive language models so they fit on hardware that should be impossible. LLaMA, Mistral, DeepSeek — pick your model, run it on a laptop. No cloud, no API keys, no limits. More tools like this — follow @Fork_Cast.