ggml

ggmlggml
G

ggml

A C/C++ tensor library for efficient machine learning inference on consumer hardware.

DR 492021Active
Rate

About

ggml.ai is a company founded to support the development of ggml, a tensor library for machine learning that enables large models and high performance on commodity hardware. ggml is used by projects like llama.cpp and whisper.cpp, offering a low-level, cross-platform implementation with integer quantization support and broad hardware compatibility. The company's development is open-source under the MIT license, with a focus on simplicity and encouraging experimentation.
Frequently asked

What does ggml do?

ggml.ai is a company founded to support the development of ggml, a tensor library for machine learning that enables large models and high performance on commodity hardware. ggml is used by projects like llama.cpp and whisper.cpp, offering a low-level, cross-platform implementation with integer quantization support and broad hardware compatibility. The company's development is open-source under the MIT license, with a focus on simplicity and encouraging experimentation.

When was ggml founded?

ggml was founded in 2021.

What industry does ggml operate in?

ggml operates in Foundation Model, Large Language Model, AI Infrastructure, Edge AI, On-Device AI, Machine Learning.

Comments
(6)
5 positive1 mixed0 negative
Reddit
r/LocalLLaMAu/Acceptable-Cycle4645Jul 15, 2026Positive

[audio.cpp] 10 hours of audio generated in 3 minutes on RTX 5090 (demo included)! C++/GGML based Supertonic 3, MOSS-TTS, IndexTTS2, and Irodori-TTS released.

View on Reddit
Reddit
r/LocalLLaMAu/Acceptable-Cycle4645Jul 3, 2026Positive

[audio.cpp] The Sound of GGML — C++/GGML native ACE-Step, Stable Audio, HeartMuLa, RoFormer, HTDemucs released. 10-Minute Music in 60 Seconds!

View on Reddit
Reddit
r/LocalLLaMAu/Acceptable-Cycle4645Jul 1, 2026Positive

[audio.cpp] VibeVoice 1.5B released — 90-min podcast in 22.95 min, 4.08x real-time, 2.86x faster than Python without quantization. Native C++/ggml.

View on Reddit
Reddit
r/LocalLLaMAu/Acceptable-Cycle4645Jun 25, 2026Positive

audio.cpp: 12 audio models (Qwen3-TTS, PocketTTS, VeVo2 etc) in 1 C++/ggml runtime — TTS up to 5x faster than Python on CUDA.

View on Reddit
Reddit
r/LocalLLaMAu/mudler_itMay 31, 2026Positive💎

I ported NVIDIA Parakeet (speech-to-text) to ggml: same output as NeMo, faster, GGUF-quantized, no Python.

View on Reddit
Reddit
r/LocalLLaMAu/Remove_AyysSep 27, 2025Mixed

For llama.cpp/ggml AMD MI50s are now universally faster than NVIDIA P40s. In 2023 I implemented llama.cpp/ggml CUDA support specifically for NVIDIA P40s since they were one of the cheapest options for GPUs with 24 GB VRAM.

View on Reddit

Some comments are pulled from public discussions around the web (look for the source icon). Quotes are excerpts; click through to read the full thread.