Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU
FreeToken is a new edge-native serving engine designed to run massive MoE models like GLM-5.2 on single workstation GPUs.
FreeToken addresses the hardware gap for frontier open-weight models by optimizing MoE cache management. It dynamically splits cache misses between PCIe transfers and CPU execution based on real-time bandwidth measurements, allowing developers to run large-scale models locally without requiring datacenter-class GPU clusters.