MEGA Hub

CAMTA: A Reconfigurable Multi-Region Activation Unit for Nonlinear Function Approximation

Authors

Do you know Carlos Soto-Porras?You can claim authorship or link another user.Do you know Jose Fonseca-Cruz?You can claim authorship or link another user.Do you know Pablo Ramirez-Morera?You can claim authorship or link another user.Do you know Erick Obregon-Fonseca?You can claim authorship or link another user.Do you know Luis G. Leon-Vega?You can claim authorship or link another user.Do you know Jorge Castro-Godinez?You can claim authorship or link another user.

Abstract

Nonlinear activation functions are widely used in machine learning workloads, but their direct hardware implementation is often costly, function-specific, or difficult to reuse across different models. This work introduces CAMTA, a 16-bit reconfigurable multi-region activation unit for nonlinear function approximation in FPGA and ASIC accelerators. CAMTA combines independent region thresholds, per-region polynomial degrees, coefficient sets, and execution modes over a shared Horner-based datapath. Unlike conventional polynomial or piecewise approximation units that mainly reconfigure coefficients or segment selection, CAMTA also reconfigures the computational behavior of each region through HORNER, CONST, ZERO, and IDENTITY modes, enabling the same hardware to support functions with different symmetry and tail behavior without resynthesis. FPGA validation on an AMD Alveo platform shows RMSE as low as \(3.60\times10^{-6}\) for CAMTA-assisted Softmax, outperforming the CORDIC-based Softmax baseline considered in this work by nearly one order of magnitude. FPGA HLS synthesis reports 3 DSPs, 802 FFs, 1756 LUTs, and an 11-cycle datapath latency. ASIC synthesis in TSMC 65~nm at 250~MHz reports \(6632.40~μ\mathrm{m}^2\) total cell area and \(1.3634~\mathrm{mW}\) total power. Compared with a same-node, function-specific PLAC implementation, CAMTA incurs \(2.20\times\) area and \(1.75\times\) power overhead, in exchange for runtime configurability and reuse across multiple nonlinear functions.

Community

00

Publication notes

Author note
Pre-print under review on ICECS 2026