MicroLLM lab - tiny LLMs, Q4, in your browser
MicroLLM lab enables running and benchmarking small language models directly in the browser using WebGPU.
The tool allows developers to execute 25M–360M parameter models locally without server costs or data privacy concerns. By leveraging WebGPU, it provides an efficient edge layer for tasks like intent extraction and query filtering, reducing the need for cloud-based LLM calls.