
Bin Fan
Founder and VP, Technology
Alluxio, Inc.
Enterprises investing in artificial intelligence and machine learning infrastructure often seek to improve price performance by increasing GPU utilization while managing infrastructure and data-access costs.
Oracle Cloud Infrastructure (OCI) provides high-performance GPU compute for demanding AI workloads. For I/O-intensive workloads, however, latency and bandwidth between compute and remote object storage can affect overall accelerator utilization. When GPUs wait for data, organizations may pay for accelerator capacity that is not being fully used.
This reference architecture shows how deploying Alluxio as a compute-side distributed caching layer on OCI can help:
- Improve GPU utilization
- Reduce effective infrastructure cost for certain training or inference workloads
- Reduce data-access latency for frequently accessed datasets
- Work with Oracle Cloud Infrastructure Object Storage and other supported object-storage backends
By helping reduce storage-related bottlenecks, this architecture can help improve effective GPU price performance for supported workloads.
Business objective: Improve AI infrastructure price performance
GPU clusters can represent a significant component of AI infrastructure cost. Price performance can improve when:
- GPUs maintain higher utilization
- I/O bottlenecks are reduced
- Unnecessary remote data movement is reduced
- Persistent storage remains cost-efficient
Adding Alluxio on the compute side can help keep GPUs supplied with frequently accessed data more consistently, supporting improved utilization and infrastructure efficiency.
Benchmark scope and methodology: H100 accelerator utilization was evaluated using MLPerf Storage v2.0’s standard accelerator-emulation methodology; no physical H100 GPUs were used. Storage I/O, network traffic, and caching were exercised in the physical test environment described below, where latency, IOPS, and throughput were measured.
Architecture overview
1. Compute-side acceleration
Modern GPU instances can include high-performance local NVMe storage. Alluxio is a software-based distributed caching layer that can run alongside OCI GPU compute and use local NVMe capacity across participating nodes. For repeatedly accessed data, this architecture can reduce dependence on remote-storage latency and available network bandwidth between compute and persistent storage.
2. Storage-agnostic backend
Persistent data can remain in supported object-storage backends, including:
- Oracle Cloud Infrastructure Object Storage
- Amazon S3
- Google Cloud Storage
- Microsoft Azure Blob Storage
- Supported on-premises object storage
This approach can support architectures in which compute and persistent storage are managed independently.
3. Throughput scaling
AI workloads can be I/O intensive. Alluxio uses a distributed architecture in which additional worker nodes can contribute cache capacity and bandwidth. In the tested configuration described later in this article, throughput increased as nodes were added.
4. No bulk dataset migration or proprietary storage format required
At petabyte scale, bulk dataset migration can add cost and operational complexity. With this architecture, object storage can remain the system of record while Alluxio caches frequently accessed data closer to compute. This approach does not require customers to bulk-migrate or reformat the source dataset solely to use the caching layer.
Data flow
- AI/ML workloads run on OCI GPU instances.
- Workloads access data through Alluxio.
- Alluxio retrieves requested data from object storage when it is not already cached.
- Retrieved data is cached on NVMe storage distributed across participating workers.
- Subsequent reads of cached data can be served from the compute-side cache with lower latency than remote reads in the tested configuration.
Illustrative reference architecture with Alluxio deployed near OCI GPU compute

Note: This diagram is illustrative and does not depict the physical WARP test platform. Alluxio uses the front-end TCP/IP network. In the tested configuration, traffic used 100 GbE; RDMA was not used.
Architecture components
1. OCI region
Compute resources can be deployed within a single OCI region to take advantage of regional networking.
2. GPU compute layer
OCI GPU shapes, including supported H100-based configurations, can host:
- Training workloads
- Fine-tuning workloads
- Inference workloads
- Kubernetes-based or direct bare-metal deployment
3. Alluxio cluster
An Alluxio cluster consists of Alluxio worker processes that contribute capacity and bandwidth to a distributed cache. Workers can run on:
- Separate OCI Dense I/O nodes provisioned for caching, in a dedicated configuration
- OCI GPU nodes that also host the compute workload, in a co-located configuration
Depending on configuration, workers can use:
- Local NVMe
- DRAM
- Network bandwidth
- CPU resources
Workloads can access data through supported Alluxio POSIX or S3-compatible interfaces. Caching frequently accessed data closer to compute can help reduce remote-read latency. Alluxio Enterprise 3.6 uses a decentralized architecture in which data and metadata are distributed across workers. Node-failure behavior was not tested as part of this benchmark.
4. Persistent object-storage layer
Persistent data can remain in OCI Object Storage or another supported object-storage backend, which continues to serve as the system of record. Once requested data is cached, repeated reads can rely less on the location and latency of the backend storage.
Choosing an Alluxio deployment model on OCI
Co-located mode: Cost-efficient configuration
Alluxio runs on the provisioned GPU instances.
Potential fit:
- Smaller clusters
- Rapid deployment
- Cost-sensitive environments
Potential benefits:
- Does not require separate caching nodes
- Can use existing local NVMe capacity
- Uses configurable CPU, memory, and NVMe resources alongside the GPU workload
- Can help improve price performance in appropriate compact environments
In co-located mode, Alluxio workers run on the same OCI GPU instances as the AI workload.

Dedicated mode: Performance-oriented configuration
Alluxio runs on separate Dense I/O nodes.
Potential fit:
- Large multi-GPU clusters
- Shared infrastructure
- High-throughput environments
Potential benefits:
- Additional cache and throughput capacity
- Independent scaling of the caching tier and GPU cluster
- The potential for higher sustained accelerator utilization in validated configurations

In dedicated mode, Alluxio workers run on separate OCI Dense I/O instances while GPU workloads run on OCI GPU instances.
How to deploy Alluxio on OCI
1. Provision GPU compute on OCI
- Deploy supported GPU shapes in the target region.
- Configure networking appropriate for the required bandwidth.
2. Identify persistent object storage
- Use Oracle Cloud Infrastructure Object Storage or another supported backend.
3. Deploy Alluxio on the compute side
- Install Alluxio on GPU nodes for a co-located configuration or on separate Dense I/O nodes for a dedicated configuration.
- Configure the supported object-storage backend for Alluxio access.
4. Point AI workloads to Alluxio
- Use a supported S3-compatible endpoint or POSIX interface.
- This approach can help reduce the need for bulk dataset migration and application changes, depending on the workload and existing access method.
5. Warm the cache, if appropriate
- Preload frequently accessed datasets using supported Alluxio functionality.
- Cache warming can help increase the amount of requested data available locally when a workload begins.
Benchmark test configuration
The results and examples in this post are based on Alluxio Enterprise 3.6. The physical WARP test environment used OCI BM.DenseIO.E5.128 instances, with WARP client workloads running as Kubernetes jobs connected to Alluxio cluster workers and OCI Object Storage as the persistent storage backend.

Figure: WARP performance-testing environment used for the benchmark results discussed in this post
Performance characteristics
Benchmark results discussed in this post include:
- Approximately 90–98% accelerator utilization reported by MLPerf Storage v2.0 using emulated H100 accelerators.[1]
- Approximately 0.3–0.6 ms average warm-cache data-access latency measured with WARP
- Up to approximately 57 GB/s aggregate I/O throughput measured in the six-node MLPerf Storage configuration using 100 Gbps networking
The key takeaway is that reducing time spent waiting for frequently accessed data can help increase sustained GPU utilization and improve effective GPU price performance for I/O-constrained workloads. See the related OCI + Alluxio benchmark article.
Price-performance impact
Without a compute-side cache, I/O-constrained workloads may experience:
- More GPU time waiting for remote reads
- Higher effective infrastructure cost for a given amount of useful accelerator work
With an appropriately configured Alluxio cache:
- GPUs can spend less time waiting for repeatedly accessed data
- Repeated remote-storage access, and associated access costs where applicable, may be reduced
- Higher infrastructure utilization can help improve effective price performance
Performance comparison
| Metric | Baseline | Alluxio | Improvement |
| Illustrative GPU utilization* | 60% | 95% | 58% higher |
| WARP benchmark — 10 KiB objects, high-frequency small-object access | |||
| Average latency (16 concurrent connections) | 24.1 ms | 0.45 ms | ~53X lower |
| IOPS (48 concurrent connections, 6-node cluster) | 11,800 | 240,800 | ~20X |
| WARP benchmark — 1 GiB objects, ~500 GB dataset | |||
| Throughput | 6.5 GiB/s | 33.3 GiB/s average (40.9 GiB/s peak) | ~5X |
* GPU-utilization values are illustrative model inputs. Storage-performance results are based on the WARP test conditions shown in the table.
What to consider before deployment
Cache sizing
Size the Alluxio cluster based on the workload and target data rates.
Dataset working set: Performance can improve when frequently requested data fits within the distributed NVMe cache. Reads served from cache can avoid remote object-storage access and its associated latency.
Network bandwidth: For throughput-sensitive workloads, deploy sufficient worker capacity and network bandwidth to support target data rates.
Concurrency: Highly concurrent workloads may require additional worker capacity and I/O parallelism to reduce cache contention and network bottlenecks.
Multicloud strategy
Because Alluxio supports multiple storage backends, organizations can, subject to supported configurations:
- Keep persistent data in OCI Object Storage, supported on-premises object storage, or another supported object-storage backend
- Run GPU compute on OCI
- Help reduce the need for large-scale migration of persistent datasets solely to place data near compute
Once frequently accessed data is cached, Alluxio can reduce the degree to which repeated-read performance depends on the location of persistent storage.
Operational monitoring
Monitor metrics such as:
- GPU utilization
- Cache-hit rate
- Network saturation
- Storage-backend latency
For workloads that benefit from caching, increases in cache-hit rate may correlate with improved accelerator utilization and price performance.
Conclusion
Improving AI/ML infrastructure price performance often requires keeping GPUs supplied with data while reducing I/O bottlenecks.
By deploying Alluxio as a compute-side caching layer with OCI GPU instances, organizations can help reduce data-access latency, improve accelerator utilization, and improve effective price performance for appropriately configured I/O-constrained workloads.
[1] MLPerf® v2.0 Training ResNet-v2.0. Result not verified by MLCommons Association. Unverified results have not been through an MLPerf review and may use measurement methodologies and/or workload implementations that are inconsistent with the MLPerf specification for verified results. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See www.mlcommons.org for more information.

Bin Fan
Founder and VP, Technology
Alluxio, Inc.
