colibrì: Run 744B MoE AI Models on Consumer Hardware
A 744B MoE AI model that runs on consumer hardware with zero dependencies.
★ 4.5 / 5 iKeep scoreScreenshots
Quick facts
Overview
colibrì is a 744B Mixture of Experts (MoE) model designed to run on consumer hardware, leveraging pure C with zero dependencies. It efficiently tiers 19,456 experts across VRAM, RAM, and disk, enabling frontier-class AI performance without requiring datacenter resources.
Getting started with colibrì — tiny engine, immense model
- Clone the repositoryStart by cloning the colibrì repository from its official source to your local machine.

- Compile the sourceCompile the single C file implementation, ensuring your system meets the 25 GB memory requirement.
- Configure resourcesSet up the tiering of 19,456 experts across VRAM, RAM, and disk as per your hardware capabilities.
- Run the modelExecute the compiled binary to start running the 744B MoE model on your consumer hardware.
- Monitor performanceUse the live routing telemetry and per-expert heat tracking features to monitor model performance.
Pricing
| Plan | Price | Note |
|---|---|---|
| Open-source | Free | open-source |
💡 Indicative pricing — check the official rates at justvugg.github.io · Updated 7/21/2026
Key features
- 744B MoE model runs on 25 GB machines
- Pure C implementation with zero dependencies
- 19,456 experts tiered across VRAM, RAM, and disk
- Live routing telemetry and per-expert heat tracking
- Single C file for easy readability and modification
- Token-exact validation against reference transformers
Pros
- Runs frontier-class models on consumer hardware
- Transparent and open for study and improvement
- Efficient use of hardware resources
Cons
- Requires technical expertise to modify and optimize
- Performance may vary based on hardware configuration
Use cases
- AI researchers studying large language models
- Developers experimenting with MoE architectures
- Enthusiasts wanting to run frontier models locally
The verdict
colibrì is a groundbreaking tool for AI researchers and developers, offering unprecedented access to frontier-class models on consumer hardware. Its transparency and efficiency make it a must-try for those in the field.
✓ Who should use it
AI researchers and developers who need to run and study large MoE models without datacenter resources.
✕ Who should skip it
Casual users or those without technical expertise in AI and machine learning.
Categories
Community
- GitHub Issues ↗Official bug reports and feature discussions
Great for
Featured in
Frequently asked questions
Is colibrì — tiny engine, immense model free?
colibrì — tiny engine, immense model is free to use.
What is colibrì — tiny engine, immense model best for?
AI researchers and developers who want to study and improve large MoE models on consumer hardware.
Does colibrì — tiny engine, immense model have an API?
No, colibrì — tiny engine, immense model does not offer a public API.
Can I self-host colibrì — tiny engine, immense model?
Yes, colibrì — tiny engine, immense model can be self-hosted.
Who should use colibrì — tiny engine, immense model?
AI researchers and developers who need to run and study large MoE models without datacenter resources.
Reviews 5.0 ★ · 1 review
colibrì is a game-changer for AI researchers and developers, offering the ability to run and study a 744B MoE model on consumer hardware. Its pure C implementation and zero dependencies make it highly efficient and transparent. The live routing telemetry and per-expert heat tracking provide invaluable insights for optimization. This tool is perfect for those who want to push the boundaries of AI without relying on datacenter resources. Its open-source nature and single C file design make it accessible for modification and improvement by the community.
No reviews yet. Be the first to review colibrì — tiny engine, immense model.