In just a few years, AI’s explosive growth has made it a mainstay in everyday life, but also sparked major resource consumption concerns. Large language models (LLMs) such as ChatGPT and Claude have significantly larger computational and energy costs versus smaller neural networks: ChatGPT alone drew an estimated 23 gigawatt hours of electricity per month in 2024, exceeding the annual consumption of many countries. Yet, LLMs continue to be widely deployed, making efficiency improvements ever more critical.
Faced with this challenge, Tao Luo, Head of the High Performance Computing Chapter at the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), and Zhehui Wang, A*STAR IAIC Senior Scientist, are adapting novel electronic components known as memristor crossbars—originally trialled to significant success in computer vision models—to LLMs.
“Memristor crossbars are very compact memory-and-computing devices where the memory portions are arranged in crossbar-like structures, allowing other portions to physically perform mathematical operations directly where the data is stored,” said Luo. “The key advantages of these devices are their density and energy efficiency.”
This two-in-one device architecture cuts down a typically energy-hungry portion of computing: data movement. Traditional computers store data in a memory device before moving them to a separate processor device for computation; memristor crossbars remove the need for data to travel, potentially easing a major bottleneck for LLM deployment.
Adapting memristor crossbars for LLMs, however, can be challenging. Besides using volumes of data so large they push the limits of denser on-chip memory technologies, LLMs also contain attention mechanisms and rely on nonlinear operations, which traditional memristor crossbars have not been optimised for.
Together with collaborators from Nanyang Technological University, Singapore; the National University of Singapore; and the Chinese Academy of Sciences, the team developed a memristor crossbar architecture that combines a computation crossbar and a dense crossbar to overcome these hurdles.
“The dense crossbar provides large-capacity storage, while the computation crossbar provides flexible and energy-efficient computing,” Wang said. “By integrating them on the same chip or package, our architecture reduces the need to move data between separate chips or external memory.”
Through simulation experiments with several state-of-the-art LLMs including GPT-3 and LLaMa, the researchers found that their memristor crossbar architecture was as accurate as conventional GPU-based computation, while also being about 68 times more efficient overall in terms of chip space and computation time combined. Their device could also be implemented on chip areas 39 times smaller than traditional memristor crossbars while reducing energy consumption 18-fold.
“We showed a feasible path for deploying LLMs on memristor-based architectures, not just smaller or more regular neural networks,” Wang commented. “This work achieves an architectural synergy between dense storage and flexible computation, providing an important step toward more compact and sustainable AI hardware.”
Moving forward, the team aims to scale their memristor crossbar design to support larger models and move it towards practical deployment.
The A*STAR-affiliated researchers contributing to this research are from the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC).
Facts Only
* Large language models like ChatGPT and Claude have significant computational and energy costs.
* ChatGPT drew an estimated 23 gigawatt hours of electricity per month in 2024.
* Tao Luo and Zhehui Wang are adapting memristor crossbars to LLMs.
* Memristor crossbars are compact memory-and-computing devices with memory arranged in crossbar structures allowing direct mathematical operations where data is stored.
* This architecture reduces energy consumption by cutting down data movement between separate memory and processor devices.
* The team developed a memristor crossbar architecture combining a computation crossbar and a dense crossbar.
* Simulation experiments with GPT-3 and LLaMa showed the architecture was as accurate as conventional GPU computation.
* The architecture was 68 times more efficient overall in terms of chip space and computation time combined compared to conventional methods.
* The device could be implemented on chip areas 39 times smaller than traditional memristor crossbars.
* Energy consumption was reduced by 18-fold with the new architecture.
Executive Summary
Full Take
The core narrative centers on a necessary technological pivot: addressing the unsustainable energy demands of large-scale AI deployment through novel hardware design. The adaptation of memristor technology suggests a systemic tension between the functional requirements of cutting-edge AI (dense storage, complex non-linear operations) and the physical limitations of current computing architectures. The proposed solution—integrating dense storage with flexible computation in a combined crossbar architecture—is not merely an incremental optimization but an architectural synergy that seeks to resolve fundamental bottlenecks.
The pattern observed is the imposition of environmental and resource constraints (energy consumption) as the primary driver for technological innovation, forcing materials science and hardware design to adapt rather than simply scale existing paradigms. The juxtaposition of "denser storage" versus "flexible computation" highlights a recurring theme in technology: optimizing performance often requires merging traditionally separate functional domains. The challenge shifts from simply making AI *bigger* to making it fundamentally *more efficient* within physical constraints.
What assumptions underpin the pursuit of density and efficiency in this context? Are these constraints solely technological, or do they reflect deeper societal expectations regarding resource stewardship? If the path forward is an architectural integration that yields extreme efficiency gains, what new bottlenecks might emerge as models scale further, and how can research proactively anticipate those future limitations rather than merely optimizing current metrics?
