This AI SSD tech makes 8 RTX 5090s perform like 46 GPUs in inference
Date:
Sat, 25 Jul 2026 18:15:00 +0000
Description:
GenStorAIGE's AI90 combines HBM, DDR, and SSD storage, claiming faster inference while expanding effective GPU memory for large language models.
FULL STORY ======================================================================Copy link Facebook X Whatsapp Reddit Pinterest Flipboard Threads Email Share this article 0 Join the conversation Follow us Add us as a preferred source on Google Newsletter Subscribe to our newsletter GenStorAIGE AI90 shifts AI memory beyond traditional GPU HBM limitations using SSDs PT200Z SSD supports constant cache updates during demanding inference workloads efficiently Eight RTX 5090 GPUs gain dramatically larger effective inference memory capacity GenStorAIGE has introduced its AI90 inference acceleration platform at WAIC 2026, taking a storage-centric approach to expanding effective AI memory capacity.
Rather than depending solely on GPU high-bandwidth memory, the platform incorporates PCIe Gen5 solid-state drives directly into the memory hierarchy itself. This allows portions of the Key-Value Cache used by large language models to sit outside GPU memory entirely. Latest Videos From TechRadar Watch full video here: A three-tier memory architecture built around SSD offloading AI90 combines HBM, system DRAM, and SSD into a unified three-tier memory structure for handling inference workloads.
By transparently offloading KV Cache data onto SSDs, the platform reduces pressure on GPU memory while supporting significantly larger workloads and longer context windows. You may like Nvidia RTX 5090 GPUs are so expensive that Intel's Arc Pro B70 is now a genuine bargain for AI This tiny AMD PC
just ran a massive 397B AI Model that required a server room full of GPUs a year ago This SIM-card-sized 8TB PCIe 5.0 SSD hits 11GB/s, but AI firms will likely hoard them all
According to GenStorAIGE, this architecture cuts first-token latency from several seconds down to sub-second response times in supported
configurations.
That represents up to a 50x improvement, alongside throughput gains reaching 5.1x and a roughly 39% reduction in GPU memory usage. Are you a pro?
Subscribe to our newsletter Sign up to the TechRadar Pro newsletter to get
all the top news, opinion, features and guidance your business needs to succeed! Contact me with news and offers from other Future brands Receive email from us on behalf of our trusted partners or sponsors By submitting
your information you agree to the Terms & Conditions and Privacy Policy and are aged 16 or over.
Combined with intelligent peer-to-peer GPU communication, the company states AI90 can accelerate inference by up to 5.8x on systems running eight Nvidia GeForce RTX 5090 cards.
That multiplier effectively allows an eight-card setup to behave closer to a 46-GPU cluster during sustained inference tasks.
The architecture also supports context windows exceeding 128,000 tokens, enabling far larger document processing and conversation handling without exhausting available memory. What to read next Tiny company steals AMD's thunder and challenges Nvidia with old-tech PCIe AI accelerator Inference needs memory: how context is becoming AI infrastructure Anthropic's Claude wants to help Micron design better HBM, DRAM, and SSD for AI The PT200Z SSD handles the intensive write demands behind the system To support continuous write workloads generated by constant KV Cache updates, GenStorAIGE paired AI90 with its new PT200Z AI SSD.
Built using pSLC NAND flash and connected through a PCIe Gen5 x4 interface, the drive delivers sequential read speeds reaching 14.8 GB/s.
Random read performance hits approximately 3.1 million IOPS, while read latency sits at just 54 microseconds.
Write latency drops even further to 10 microseconds, supporting the rapid cache updates AI90's architecture depends on constantly.
Endurance ratings reach up to 100 drive writes per day, a figure suited for sustained enterprise AI workloads with constantly shifting cache data.
This design reflects a broader shift across AI infrastructure toward memory tiering, as LLMs increasingly outgrow the practical limits of GPU HBM alone.
Integrating extremely fast SSD storage into inference pipelines offers one method for scaling context length without requiring additional GPUs or larger HBM configurations.
Whether the performance claims hold outside controlled testing conditions remains genuinely unverified at this stage.
As with most vendor announcements, these performance figures come directly from GenStorAIGE and still require independent benchmarking across varied real-world AI workloads.
Via The Guru of 3D Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
======================================================================
Link to news story:
https://www.techradar.com/pro/this-ai-ssd-tech-makes-8-rtx-5090s-perform-like- 46-gpus-in-inference
--- Mystic BBS v1.12 A49 (Linux/64)
* Origin: tqwNet Technology News (1337:1/100)