Stateofutopia shipped a lab that runs small language models directly in your browser. No server, no account, no API bill.
The pitch is pure edge.
MicroLLM lab loads Q4 quantized models in the 25M to 360M parameter range and executes them with WebGPU, the W3C standard that hits Apple Silicon Metal, DirectX 12 and Vulkan from inside the browser window. Weights are compressed from 16-bit to 4-bit, cutting memory by 75% so a 100M+ model fits in about 50 to 84 MB of browser memory. The site claims sub-10ms time-to-first-token, infinite concurrency on client GPUs, and 100% private prompts that never leave the device.
That last part is the real sell, and the thinnest proof.
The lab gives you three steps: click Load to cache a model to IndexedDB, chat in a panel that shows live tok/s, then run benchmarks and generate a shareable performance certificate with peak and sustained tokens per second. It will run a sustained 256-token decode, show wall time per test, JS heap, GPU buffers, and adapter info, and it falls back to WASM then JS if WebGPU is not available. The downloadable zip is listed at 589 MB, with a current build at 591 MB, and it must be served over HTTP, not file://. For checks it is explicit: the suite uses objective regex and exact-token tests, not writing quality. As the UI notes, a 135M model is allowed to fail, that is the measurement.
