• 2 min read
GenStorAIGE says SSDs make 8 RTX 5090s act like 46
GenStorAIGE’s AI90 uses PCIe Gen5 SSDs to expand inference memory, with claimed gains equivalent to turning eight RTX 5090s into 46 GPUs.

Image: TechRadar
GenStorAIGE says eight Nvidia GeForce RTX 5090 GPUs can perform closer to a 46-GPU cluster during sustained inference with its new AI90 platform, which uses PCIe Gen5 SSDs as part of the memory hierarchy. The company introduced the system at WAIC 2026.
Rather than relying exclusively on GPU high-bandwidth memory (HBM), AI90 moves portions of the large language model Key-Value Cache (KV Cache) onto SSD storage. That gives inference workloads access to a larger effective memory pool without adding GPUs or expanding HBM capacity.
AI90's three-tier memory architecture
AI90 combines three memory tiers:
- GPU HBM
- System DRAM
- PCIe Gen5 SSD storage
The platform transparently offloads KV Cache data to SSDs, reducing GPU memory pressure while supporting longer context windows. GenStorAIGE claims the approach can support context windows exceeding 128,000 tokens.
In supported configurations, the company says AI90 cuts first-token latency from several seconds to under one second—a claimed improvement of up to 50x. It also reports throughput gains of up to 5.1x and approximately 39% lower GPU memory usage.

Recommended reading
A 28.9M-parameter LLM runs on an $8 chip
With intelligent peer-to-peer GPU communication, the system is claimed to deliver up to 5.8x faster inference on machines equipped with eight RTX 5090 cards.
PT200Z SSD targets constant cache writes
GenStorAIGE designed its PT200Z AI SSD to handle the continuous write activity generated by frequent KV Cache updates. The drive uses pSLC NAND flash and a PCIe Gen5 x4 interface, with specifications that include:
- Sequential read speeds of up to 14.8 GB/s
- Random read performance of approximately 3.1 million IOPS
- Read latency of 54 microseconds
- Write latency of 10 microseconds
- Endurance of up to 100 drive writes per day
The approach reflects a broader move toward memory tiering as large language models push beyond the practical limits of GPU HBM. However, all performance figures come from GenStorAIGE and have not yet been independently verified across varied real-world workloads.
Via The Guru of 3D
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via TechRadar


