ggml
A C/C++ tensor library for efficient machine learning inference on consumer hardware.
About
ggml.ai is a company founded to support the development of ggml, a tensor library for machine learning that enables large models and high performance on commodity hardware. ggml is used by projects like llama.cpp and whisper.cpp, offering a low-level, cross-platform implementation with integer quantization support and broad hardware compatibility. The company's development is open-source under the MIT license, with a focus on simplicity and encouraging experimentation.
Key facts
At-a-glance profile data
- Website
- ggml.ai
- Founded
- 2021
- Total funding
- -
- Business model
- B2B
- Tags
- Foundation ModelLarge Language ModelAI InfrastructureEdge AIOn-Device AI
Tech stack
Detected on ggml.ai, last scanned 17 Sept 2026. Run your own scan
- CDN
- Fastly
- Email provider
- Google Workspace
“[audio.cpp] 10 hours of audio generated in 3 minutes on RTX 5090 (demo included)! C++/GGML based Supertonic 3, MOSS-TTS, IndexTTS2, and Irodori-TTS released.”
View on Reddit“[audio.cpp] The Sound of GGML — C++/GGML native ACE-Step, Stable Audio, HeartMuLa, RoFormer, HTDemucs released. 10-Minute Music in 60 Seconds!”
View on Reddit“[audio.cpp] VibeVoice 1.5B released — 90-min podcast in 22.95 min, 4.08x real-time, 2.86x faster than Python without quantization. Native C++/ggml.”
View on Reddit“audio.cpp: 12 audio models (Qwen3-TTS, PocketTTS, VeVo2 etc) in 1 C++/ggml runtime — TTS up to 5x faster than Python on CUDA.”
View on Reddit“For llama.cpp/ggml AMD MI50s are now universally faster than NVIDIA P40s. In 2023 I implemented llama.cpp/ggml CUDA support specifically for NVIDIA P40s since they were one of the cheapest options for GPUs with 24 GB VRAM.”
View on RedditSome comments are pulled from public discussions around the web (look for the source icon). Quotes are excerpts; click through to read the full thread.