News

NYCU's Professor Tsung-Tai Yeh and Team Sweep ISCA Honors, Breaking AI Power and Memory Bottlenecks

Published on
Author
林珮雯

NYCU's Professor Tsung-Tai Yeh and Team Sweep ISCA Honors, Breaking AI Power and Memory Bottlenecks

In the field of computer science, there is a sanctuary known as the "Olympic Games of Computer Architecture"—ISCA (International Symposium on Computer Architecture). With an extremely high bar for acceptance, this conference represents the most forward-thinking innovations in global hardware design and chip architecture. Over the past decade, Taiwan's presence at ISCA has been scarce, averaging less than one paper per year. However, this academic ceiling has been shattered for two consecutive years by the research team led by Professor Tsung-Tai Yeh from the Department of Computer Science at National Yang Ming Chiao Tung University (NYCU).

Professor Yeh's team has once again advanced to ISCA 2026 with their Omni-LUT

research. This marks the team's second consecutive year planting their flag in this prestigious venue, alongside simultaneous breakthrough results published at ICRA, a top-tier robotics conference. The success stems from two papers that address the tech industry's most pressing headaches: "AI's excessive memory consumption" and "the high power demands of high-definition computing. "

Solving the "Traffic Jam" of Data Transfer

The industry's primary bottleneck is the massive volume of data movement. Professor Yeh explains that whether it is popular Large Language Models (LLMs) or ultra-realistic 3D rendering, the core challenge lies in the frequent transfer of enormous amounts of data between the chip and the memory. Similar to a traffic jam, this process consumes the vast majority of power and time, causing systems to run slowly and overheat. Consequently, moving less data while maintaining precise results has become the most critical technological pivot for the future of the semiconductor and AI industries.

Omni-LUT: Redefining Long-Context AI

In their Omni-LUT research targeting LLMs, the team pushed the limits of long-context processing. While Microsoft Research previously proposed using "Look-up Tables" (LUTs) with ultra-low precision quantization to replace traditional calculations and reduce chip area, that technology struggled with the quantization of the "KV Cache." Since KV Cache is generated dynamically during model execution, it is difficult to pre-process, causing memory demands to skyrocket when handling ultra-long documents. Professor Yeh's team developed a new quantization mechanism—distinct from Google's strategy—that perfectly integrates with LUT accelerators. This breakthrough significantly reduces the memory pressure of the KV Cache, allowing long-context AI operations to run faster and more stably on conventional hardware.

AQB8: Cinematic Graphics with Low Power

In the realm of 3D rendering, Ray Tracing technology produces lifelike images but requires simulating a sea of light rays and frequent data movement. Even the highest-end graphics cards currently on the market struggle with the workload.

Professor Yeh's team developed AQB8 technology, breaking away from traditional high-precision computing frameworks. They discovered a new algorithm capable of compressing data to low precision with virtually no loss in quality. By drastically reducing data movement, this technology enables mobile devices to run movie-quality graphics with low power consumption. Last year, this algorithm garnered significant attention from major global chipmakers, who were impressed by its ability to lower power usage without compromising visual fidelity.

Future Horizons: From Quantum Computing to Cybersecurity

Looking ahead, Professor Yeh is actively expanding these core technologies—"Ray Tracing Acceleration Architecture" and "Hardware-Aware Quantized Compression"—into broader fields. The team is currently adapting ray tracing logic for Quantum Circuit Simulation and applying memory compression to Information Security.

The team's innovative project with Qualcomm was awarded the 2025 Qualcomm Innovation Fellowship in East Asia. The goal is to use quantization compression to optimize the efficiency of "Homomorphic Encryption, " ensuring that AI can process encrypted data at high speeds without compromising privacy.

By securing a place at ISCA for two years running, this NYCU team has proven that Taiwan not only possesses world-class chip manufacturing capabilities but also the expertise in "Hardware-Software Co-Design" to define the future of computing rules across AI, quantum computing, and digital security.