FreeToken Unlocks Frontier MoE Inference on Consumer Hardware via Dynamic Co-Execution
FreeToken is an open-source inference engine designed to optimize Mixture-of-Experts (MoE) model performance on consumer-grade hardware.
Developed by researchers at UC Berkeley and MIT, FreeToken utilizes dynamic scheduling and weight management to improve decoding speeds. The engine aims to lower the barrier for running complex reasoning systems locally on edge devices.