This AI SSD tech makes 8 RTX 5090s perform like 46 GPUs in inference

  • GenStorAIGE AI90 shifts AI memory beyond traditional GPU HBM limitations using SSDs
  • PT200Z SSD supports constant cache updates during demanding inference workloads efficiently
  • Eight RTX 5090 GPUs gain dramatically larger effective inference memory capacity

GenStorAIGE has introduced its AI90 inference acceleration platform at WAIC 2026, taking a storage-centric approach to expanding effective AI memory capacity.

Rather than depending solely on GPU high-bandwidth memory, the platform incorporates PCIe Gen5 solid-state drives directly into the memory hierarchy itself.

This allows portions of the Key-Value Cache used by large language models to sit outside GPU memory entirely.

A three-tier memory architecture built around SSD offloading

AI90 combines HBM, system DRAM, and SSD into a unified three-tier memory structure for handling inference workloads.

By transparently offloading KV Cache data onto SSDs, the platform reduces pressure on GPU memory while supporting significantly larger workloads and longer context windows.

According to GenStorAIGE, this architecture cuts first-token latency from several seconds down to sub-second response times in supported configurations.

That represents up to a 50x improvement, alongside throughput gains reaching 5.1x and a roughly 39% reduction in GPU memory usage.

Combined with intelligent peer-to-peer GPU communication, the company states AI90 can accelerate inference by up to 5.8x on systems running eight Nvidia GeForce RTX 5090 cards.

That multiplier effectively allows an eight-card setup to behave closer to a 46-GPU cluster during sustained inference tasks.

The architecture also supports context windows exceeding 128,000 tokens, enabling far larger document processing and conversation handling without exhausting available memory.

The PT200Z SSD handles the intensive write demands behind the system

To support continuous write workloads generated by constant KV Cache updates, GenStorAIGE paired AI90 with its new PT200Z AI SSD.

Built using pSLC NAND flash and connected through a PCIe Gen5 x4 interface, the drive delivers sequential read speeds reaching 14.8 GB/s.

Random read performance hits approximately 3.1 million IOPS, while read latency sits at just 54 microseconds.

Write latency drops even further to 10 microseconds, supporting the rapid cache updates AI90's architecture depends on constantly.

Endurance ratings reach up to 100 drive writes per day, a figure suited for sustained enterprise AI workloads with constantly shifting cache data.

This design reflects a broader shift across AI infrastructure toward memory tiering, as LLMs increasingly outgrow the practical limits of GPU HBM alone.

Integrating extremely fast SSD storage into inference pipelines offers one method for scaling context length without requiring additional GPUs or larger HBM configurations.

Whether the performance claims hold outside controlled testing conditions remains genuinely unverified at this stage.

As with most vendor announcements, these performance figures come directly from GenStorAIGE and still require independent benchmarking across varied real-world AI workloads.

Via The Guru of 3D

Google logo on a black background next to text reading 'Click to follow TechRadar'



from Latest from TechRadar https://ift.tt/JEwWZ5C