1 articles with this tag
MicroLLM lab runs 25M, 360M Q4 models in-browser via WebGPU with zero server cost, but its benchmarks test speed, not quality.