Hardware Acceleration of Kolmogorov–Arnold Network (KAN) in Large-Scale Systems
Wei-Hsing Huang*, Jianwei Jia*, Yuyao Kong, Faaiq Waqar, Tai-Hao Wen, Meng-Fan Chang, and Shimeng Yu
IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Apr 2026
Recent developments have introduced Kolmogorov– Arnold networks (KANs), an innovative architectural paradigm capable of replicating conventional deep neural network (DNN) capabilities while utilizing significantly reduced parameter counts through the employment of parameterized B-spline functions incorporating trainable coefficients. Nevertheless, the B-spline functional components inherent to KAN architectures introduce distinct hardware acceleration complexities. While B-spline function evaluation can be accomplished through lookup table (LUT) implementations that directly encode functional mappings, thus minimizing computational overhead, such approaches continue to demand considerable circuit infrastructure, including LUTs, multiplexers, decoders, and associated components. This work presents an algorithm-hardware co-design approach for KAN acceleration. At the algorithmic level, techniques include alignment–symmetry and PowerGap KAN hardware-aware quantization, KAN sparsity-aware mapping strategy, and circuit-level techniques include N:1 time modulation dynamic voltage input generator with analog-compute-in-memory (ACIM) circuits. Furthermore, this work conducts comprehensive evaluations on large-scale KAN networks to validate the proposed methodologies. Nonideality factors, including partial sum deviations arising from process variations, have been evaluated with the statistics measured from the TSMC 22-nm RRAM-ACIM prototype chips. Utilizing optimally determined KAN hyperparameters in conjunction with circuit optimizations implemented and evaluated at the 22-nm technology node, despite the model sizes for large-scale tasks in this work increasing by 435 K \times to 756 K \times compared to tiny-scale tasks in previous work, the area overhead increases by only 26 K \times to 40 K \times , with power consumption rising by merely 48\times to 93\times , while accuracy degradation remains minimal at 0.11%–0.22%, thereby demonstrating the scaling potential of our proposed architecture.