End IO Bottlenecks for Agentic LLM Inference

By running directly on all your GPU nodes, Quobyte forms a massive, distributed storage system that can aggregate every available network and NIC to deliver unprecedented throughput.

This eliminates all storage bottlenecks, making it the perfect fit for multi-turn agentic LLM inference where fast context loading is now the dominant performance requirement.

Quobyte GPU Converged is the high-performance foundation for AI training, fine-tuning, and large-scale inference workloads.

Turn Surplus GPU Resources into Reliable Storage

Turn idle CPU and local NVMe inside your GPU nodes into storage by running Quobyte directly on the fleet. Quobyte runs on any x86 or ARM CPU, including NVIDIA Grace, and supports Hopper and Blackwell based systems, right alongside your existing processes.

The result is low-latency access that speeds up training and inference. And with bulletproof resiliency, data stays available and accessible even as GPU nodes restart, fail, or are added to or removed from the cluster.

Lower power consumption and cost

Quobyte delivered the best performance in MLPerf Storage (unverified) with less resources and power than the competition. By leveraging the underutilized resources on your existing GPU nodes, you can dramatically reduce your storage’s power consumption and cost.

GPU Converged Storage builds on that efficiency by using surplus CPU and local NVMe in your GPU nodes to reduce separate storage hardware, lowering cost and power while keeping latency low.

Storage That Scales with Your GPU Fleet

With Quobyte GPU Converged Storage, every GPU node you add contributes storage capacity and throughput automatically. There is no separate storage tier to size, deploy, or scale independently.

The Best of Both: Combine GPU nodes and dedicated storage servers in one cluster

With Quobyte’s policy-based data management, you can optimize data placement and move data transparently between these different deployment models. This powerful flexibility lets you leverage the high-performance of your GPU nodes for active workloads while using dedicated storage servers for less-frequent data access, ensuring you have the right balance of speed and cost for every task.

See Your Savings Instantly

Enter your cluster details to instantly see how much AI infrastructure cost you can reduce by using GPU nodes as storage. Compare 3-year savings across hardware, power, rack space, and networking.

Calculate Savings

Resources

Resource Type
Quobyte Delivers on the Promise of GPU Converged

GPU converged storage that stays reliable as AI infrastructure scales

Read More

Resource Type
Zoox Scales AI Storage to 30PB+

Discover how Zoox reached 30+ PB with Quobyte hybrid-tier AI storage.

Read More

Resource Type
Powering Siemens Healthineers' AI Factory

Siemens Healthineers relies on Quobyte to run 1,600+ AI experiments daily.

Read More

Resource Type
Quobyte's Bulletproof Resiliency

Built so anything can fail, Quobyte delivers resiliency at data center scale.

Read More

Resource Type
Native Multi-Tenancy

Quobyte transforms shared infrastructure into a secure SaaS platform

Read More

GPU Converged Storage

Massive scalable storage for AI that uses the hardware you already have. Turn underutilized CPU and NVMe resources into a high-performance storage tier, slashing infrastructure complexity and costs by eliminating the need for separate storage hardware.

Quobyte GPU Converged Storage runs directly on your GPU nodes, harnessing the existing CPU and NVMe resources inside the node to create a high-performance, ultra-reliable storage layer. By eliminating the need for external storage you remove the massive complexity and cost associated with managing a storage hardware and networking.

This streamlined architecture delivers up to 78% lower TCO and 87% lower power consumption while keeping your infrastructure lean and your GPUs fully utilized at any scale.

Frequently Asked Questions

What is Quobyte GPU Converged storage?

Quobyte GPU Converged is a software-defined parallel file system that runs natively on the CPU of GPU nodes (x86 or ARM). By clustering local NVMe drives and utilizing surplus CPU cycles and RAM—resources that typically sit idle during training or inference—it creates a high-performance, shared-nothing storage tier. This eliminates the need for expensive external storage appliances, turning a GPU fleet into its own massively scalable and resilient storage system.

How are my workloads isolated from Quobyte on the node?

Quobyte runs entirely in user space. The Quobyte services run in cgroups or containers, making it easy to achieve performance, resource, and secure isolation from any jobs running on the GPU nodes. Quobyte services can be assigned to specific CPU.

On the networking side, Quobyte can communicate over Ethernet or Infiniband, and can easily be limited to specific networks – like the ethernet frontend network.

How do you ensure security and data protection in shared GPU nodes?

Your data is end-to-end encrypted. This means that your data leaves the machine where your job is writing it only in encrypted form and is never decrypted until it is read by an application. You don’t have to trust the network, admins, or anyone on the GPU nodes. This also includes at rest encryption.

Quobyte also provides secure multi-tenancy isolating multiple users or groups from each other on the same Quobyte GPU Converged cluster.

Does Quobyte provide secure multi-tenancy to share the nodes among multiple customers?

Yes, Quobyte provides highly secure multi-tenancy built into every aspect of the storage system. Tenants are fully isolated from each other and can even be isolated on a disk or node basis!

The Quobyte webconsole, API and command line tools offer full self-service for tenant admins and users.

Quobyte clusters support dual-mode tenants, regular and high-security tenants, on the same cluster. High security tenants provide the encryption, certificate authentication and TLS connections necessary for use-cases like handling PII, human genome data, or medical records.

How does the Quobyte CSI driver map Kubernetes namespaces to secure storage tenants for multi-provider GPU clouds?

The Quobyte CSI driver can automatically map kubernetes namespaces onto Quobyte tenants. For enhanced end-to-end security for persistent volumes, Quobyte provides mechanisms for users and service accounts to authenticate against the storage system.

Does this work with Nvidia DGX B100 or DGX H200?

Yes, the DGX B100 and H100 have very powerful x86 CPUs with significantly more cores than most AI workloads require, plus plenty of NVMe drives. These systems are perfectly suited to run Quobyte GPU converged storage and turn those unused resources into one of the world’s most reliable and powerful storage systems for AI.

Does this work with Nvidia Grace Hopper and Grace Blackwell?

Yes, these systems have a NVIDIA Grace ARM CPU with a high core count. Quobyte supports ARM and has been supporting ARM since 2018. Quobyte GPU converged benefits from the massive number of cores on these CPUs.

Does Quobyte GPU Converged work with Infiniband or Ethernet?

Quobyte runs over any IP fabric and can take advantage of RDMA capabilities over Ethernet (ROCEv2), Infiniband, and OmniPath. Quobyte even runs in mixed mode where servers and clusters can have multiple interfaces with diverging technologies.

Does Quobyte GPU Converged work with Kubernetes?

Yes, Quobyte runs natively in containers and is tightly integrated with Kubernetes. We were among the first storage vendors to offer a Kubernetes driver in 2016!

Read more about our extensive capabilities for kubernetes clusters.

How does Quobyte GPU Converged storage speed up checkpointing and training?

Quobyte runs on every GPU node, meaning that you have a massive scale-out storage system with the throughput and low latency to speed up checkpointing and training. And with every GPU node you add, your storage system’s performance grows in lockstep so that you never outgrow your storage.

How does Quobyte GPU Converged reduce AI infrastructure spend and cost?

Quobyte GPU converged saves you money in a number of ways:

  1. No extra hardware, use what you already paid for – the unused CPU and NVMes in the GPU nodes
  2. Lower power consumption and less cooling. The extra power consumption by running Quobyte on the nodes is negligible, especially compared to external storage appliances
  3. Many high-speed networking ports saved. Your GPU nodes are already connected. External storage requires a large number of extra network ports that cost you lots of money and you need bigger switches. In addition, many dated appliances require and extra storage network on top of this.
  4. Less space used in the data center
  5. Easier operations for your team. No need to learn new technology and hardware. Quobyte runs on regular Linux and standard hardware, everything your team already know.

Can Quobyte maintain linear performance scaling when expanding a Blackwell B200 cluster to 1,000+ nodes?

Yes, the performance of a Quobyte cluster grows with every node added. We designed the system to be highly scalable so you don’t have diminishing returns and can scale clusters to 10,000s of nodes with linear performance scaling. Best of all, you can add as many nodes as you like at the same time, 1 or 1000s, and adding and removing nodes is fully non-disruptive.

Can I perform non-disruptive software updates on a Quobyte GPU Converged cluster while a multi-week training job is running?

Yes, everything in Quobyte is non-disruptive so your GPUs stay busy at all times. From updates to kernel patches, random reboots, adding or removing nodes. The cluster never goes down thanks to Quobyte’s unique and patented bulletproof resiliency. And this is not a sham like other storage systems where you “quickly swap a container” for selling you a quick interruption as “non-disruptive” updates. Quobyte is truly non-disruptive and the most robust storage system thanks to its hyperscaler design.

What happens to data availability if 20% of the nodes in a converged storage cluster go offline simultaneously due to a rack-level power failure?

Quobyte supports failure domains natively and the loss of a full failure domain, up to cluster and data center level, is not a problem at all! Thanks to Quobyte’s bulletproof resiliency you can operate like a hyperscaler, even when someone trips over a cable.

Can Quobyte’s Global Namespace be used to transparently move AI training datasets between on-prem Blackwell clusters and public cloud instances?

Yes, Quobyte can be deployed anywhere: on prem, on your own servers in a colo, on the edge, and in the public clouds. The Quobyte data mover helps you and your users to quickly move data between Quobyte servers, as well as to and from cloud storage.

In addition, you can mix architectures like x86 and ARM even inside the same Quobyte cluster!

What are the advantages of using a software-defined converged storage model over dedicated AI storage appliances for generative AI startups?

If saving lots of money isn’t enough of an argument, the quick and automatic scaling should convince you. With Quobyte GPU converged, your storage performance automatically scales with every node you add. Unlike the storage appliance that you can quickly outgrow, Quobyte grows with your needs instantly!

And unlike a storage appliance, you can also scale down a Quobyte cluster when you need less GPU nodes. Quobyte GPU Converged gives you the cost advantage and true flexibility that AI startups need. You can start at 4 GPU nodes and scale to 10,000s of thousands, without having to plan ahead.

Does GPU Converged work on Neoclouds/GPU clouds?

Yes, when you get the whole GPU node you have access to the entire CPU and NVMe and can take advantage of a Quobyte GPU Converged cluster. Why waste money on the slow and expensive external NFS storage the cloud providers offer?

Can Quobyte GPU Converged achieve the 1 GB/s per GPU storage performance requirement for NVIDIA DGX SuperPODs

Yes, Quobyte can easily saturate several 200Gbit ethernet links on DGX machines, providing more than 20GB/s to the machine. We have demonstrated with the MLPerf storage benchmark (unverified) that Quobyte can saturate more GPUs over a 200GBit link than the competitors.

What is the TCO impact of replacing an external AI storage appliance with Quobyte for a 1,000-node DGX H200 cluster?

A 1000 node DGX H200 cluster with Quobyte can save you up to $22M in TCO versus an external storage appliance, free up 800 high speed network ports, and reduce your power consumption by up to 3.4 MWh over three years. In addition to these savings, it also outperforms the external storage by providing 80TB/s bandwidth from the NVMes.