The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference

Authors

DOI:

https://doi.org/10.4108/eetiot.14884

Keywords:

LLM Inference, Dataflow Architecture, Hardware Security, ASIC Design, Processing-in-Memory, Edge AI

Abstract

The deployment of Large Language Models (LLMs) on consumer edge devices is throttled by the “Memory Wall”—the prohibitive bandwidth and energy cost of fetching gigabytes of model weights from DRAM for every token generated. Current architectures (GPUs, NPUs) treat model weights as mutable software data, incurring massive energy penalties to maintain general-purpose programmability. We propose The Immutable Tensor Architecture (ITA), a paradigm shift that treats model weights not as data, but as physical circuit topology. By encoding parameters directly into the metal interconnects and logic of mature- node ASICs (28nm/40nm), ITA eliminates the memory hierarchy entirely. We present a “Split-Brain” system design where a host CPU manages dynamic KV-cache operations while the ITA ASIC acts as a stateless, ROM-embedded dataflow engine. Logic-level simulation indicates a theoretical upper bound of 4.85× reduction in gate count per multiply-accumulate (243 gates vs 1,180 gates), though conservative estimates accounting for routing congestion suggest a 1.62× system-level reduction is more realistic for first-generation implementations. Physical energy modeling projects a 50× improvement in device-level energy efficiency (4.05 pJ/operation vs 201 pJ/operation), while full system power analysis indicates a 10–15× efficiency gain. FPGA prototype validation shows 1.81× LUT reduction for hardwired constant-coefficient multipliers, empirically validating the direction and lower bound of our efficiency claims. For practical deployment, we demonstrate that TinyLlama-1.1B fits on a single 520 mm2 monolithic die at 28nm, while Llama-2-7B requires an 8-chiplet configuration. The architecture is interface-agnostic, supporting deployment via PCIe (M.2 NVMe), Thunderbolt, or USB. Manufacturing cost analysis shows projected unit costs of $52 (1.1B) and $165 (7B) at 100K+ volume. Device power consumption is 1–3 W, with total system power of 7–12 W including host CPU. ITA raises the barrier to model extraction from $2,000 (software dump) to over $50,000 (specialized reverse-engineering equipment), though we acknowledge side-channel vulnerabilities requiring mitigation.

Downloads

Download data is not yet available.

References

[1] Horowitz, M.: Computing’s energy problem (and what we can do about it). In: Proc. IEEE Int. Solid-State Circuits Conf. (ISSCC), pp. 10–14. San Francisco, CA, USA (2014)

[2] JEDEC Solid State Technology Association: JESD209-5: Low power double data rate 5 (LPDDR5) standard. Tech. Rep., Arlington, VA, USA (2019)

[3] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Advances in Neural Information Processing Systems (NeurIPS), pp. 5998–6008. Long Beach, CA, USA (2017)

[4] Shazeer, N.: GLU variants improve transformer. arXiv preprint arXiv:2002.05202 (2020)

[5] Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al.: Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)

[6] Morgan, T.: Accelerating compute by cramming it into DRAM memory. The Next Platform (2022). https://www.upmem.com/nextplatform-com-2019-10-03-accelerating-compute-by-cramming-it-into-dram/

[7] Kwon, Y., et al.: 25.4 A 20nm 6GB function-in-memory DRAM, based on HBM2 with a 1.2TFLOPS 12 The Immutable Tensor Architecture programmable computing unit using bank-level parallelism, for machine learning applications. In: Proc. IEEE Int. Solid-State Circuits Conf. (ISSCC). San Francisco, CA, USA (2021)

[8] Jouppi, N.P., et al.: In-datacenter performance analysis of a tensor processing unit. In: Proc. ACM/IEEE Int. Symp. Comput. Archit. (ISCA), pp. 1–12. Toronto, ON, Canada (2017)

[9] Lam, C.: Hot Chips 34 – Tesla’s Dojo microarchitecture. Chips and Cheese (2022). https://chipsandcheese.com/p/hot-chips-34-teslas-dojo-microarchitecture

[10] Groq, Inc.: The LPU inference engine. White Paper, Mountain View, CA, USA (2023)

[11] Cerebras Systems: Cerebras Systems announces world’s first brain-scale artificial intelligence solution. Press Release (2021). https://www.cerebras.net/news/cerebras-systems-announces-worlds-first-brain-scale-artificial-intelligence-solution/

[12] Graphcore Ltd.: IPU-M2000: Intelligence processing unit. Datasheet, Bristol, UK (2020)

[13] SambaNova Systems Inc.: Reconfigurable dataflow architecture. White Paper, Palo Alto, CA, USA (2021)

[14] Merolla, P.A., et al.: A million spiking-neuron integrated circuit with a scalable communication network and interface. Science 345(6197), 668–673 (2014)

[15] Davies, M., et al.: Loihi: A neuromorphic manycore processor with on-chip learning. IEEE Micro 38(1), 82–99 (2018)

[16] Dettmers, T., Lewis, M., Belkada, Y., Zettlemoyer, L.: LLM.int8(): 8-bit matrix multiplication for transformers at scale. In: Advances in Neural Information Processing Systems (NeurIPS). New Orleans, LA, USA (2022)

[17] Frantar, E., Ashkboos, S., Hoefler, T., Alistarh, D.: GPTQ: Accurate post-training quantization for generative pre-trained transformers. In: Proc. Int. Conf. Learn. Represent. (ICLR). Kigali, Rwanda (2023)

[18] Lin, J., Tang, J., Wang, H., Wang, J., Wang, Y.: AWQ: Activation-aware weight quantization for LLM compression and acceleration. In: Proc. Conf. Mach. Learn. Syst. (MLSys). Santa Clara, CA, USA (2024)

[19] Weste, N.H.E., Harris, D.M.: CMOS VLSI Design: A Circuits and Systems Perspective, 4th edn. Addison-Wesley, Boston, MA, USA (2010)

[20] Reitwiesner, G.W.: Binary arithmetic. Adv. Comput. 1, 231–308 (1960)

[21] Gustafsson, O.: A difference based adder graph heuristic for multiple constant multiplication problems. In: Proc. IEEE Int. Symp. Circuits Syst. (ISCAS). New Orleans, LA, USA (2007)

[22] Synopsys Logic Library IP Team: Harnessing TSMC’s 28HPC+ process with six logic library capabilities. Synopsys IP Tech. Bull. (2015). https://www.synopsys.com/articles/logic-library-capabilities.html

[23] NVIDIA Corporation: A100 tensor core GPU architecture. White Paper, Santa Clara, CA, USA (2020)

[24] Torrance, R., James, D.: The state-of-the-art in semi- conductor reverse engineering. In: Proc. Design Autom. Conf. (DAC), pp. 333–338. San Francisco, CA, USA (2011)

[25] Child, R., Gray, S., Radford, A., Sutskever, I.: Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509 (2019)

[26] Shen, Y., Harris, N.C., Skirlo, S., Prabhu, M., Baehr- Jones, T., Hochberg, M., Sun, X., Zhao, S., Larochelle, H., Englund, D., Soljačić, M.: Deep learning with coherent nanophotonic circuits. Nature Photon. 11, 441–446 (2017)

[27] Ielmini, D., Wong, H.-S.P.: In-memory computing with resistive switching devices. Nature Electron. 1, 333–343 (2018)

[28] Kim, S., Gholami, A., Yao, Z., Mahoney, M.W., Keutzer, K.: I-BERT: Integer-only BERT quantization. In: Proc. Int. Conf. Mach. Learn. (ICML), pp. 5506–5518 (2021)

[29] Yao, Z., Yazdani Aminabadi, R., Zhang, M., Wu, X., Li, C., He, Y.: ZeroQuant: Efficient and affordable post-training quantization for large-scale transformers. In: Advances in Neural Information Processing Systems (NeurIPS), vol. 35, pp. 27168–27183 (2022)

[30] Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., Han, S.: SmoothQuant: Accurate and efficient post-training quantization for large language models. In: Proc. 40th Int. Conf. Mach. Learn. (ICML), pp. 38087–38099. PMLR (2023)

Downloads

Published

21-09-2026

How to Cite

1.
Li F. The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference. EAI Endorsed Trans IoT [Internet]. 2026 Sep. 21 [cited 2026 Sep. 21];11. Available from: https://publications.eai.eu/index.php/IoT/article/view/14884