CXL has become one of the most important technologies shaping AI infrastructure. As hyperscalers race to deploy larger AI models, longer context windows and increasingly memory-intensive inference workloads, memory capacity and bandwidth have emerged as critical constraints on performance, efficiency and scaling. At the same time, CXL adoption is reaching an inflection point, moving from evaluation into real-world deployment across hyperscale environments.
Marvell is leading this transition with Structera™ X memory expansion solutions developed alongside the world’s leading hyperscalers. The story begins with the shipping of Structera X 2404 and 2504 platforms, which have enabled hyperscalers to expand memory resources more efficiently, including extending the useful life of existing DDR4 investments while powering demanding AI workloads.
Structera X is not a series of disconnected product eras—it is a single, continuous architectural evolution. Today’s generation is already delivering real hyperscaler deployments, ecosystem maturity and a compelling TCO advantage. From that foundation, Marvell is extending the architecture toward the next phase of AI infrastructure innovation, adding capabilities enabled by the evolving CXL 3.2 ecosystem, PCIe Gen 6 connectivity and more advanced multi-host memory sharing architectures. These advancements will create larger, more flexible memory pools, enabling more efficient sharing of resources across servers and improving infrastructure utilization at hyperscale. As the architecture advances, Marvell is driving it toward higher bandwidth, deeper data optimization and increasingly disaggregated memory environments built to meet the growing demands of AI workloads.
Why does this matter for hyperscalers?
AI infrastructure economics are increasingly determined by how efficiently memory can be deployed, shared and utilized. Memory stranded behind individual servers reduces overall infrastructure efficiency, while shortages of next-generation memory can slow expansion plans. CXL provides a path to disaggregate and optimize memory resources at scale, improving utilization and maximizing the value of existing investments.
Modern AI inference workloads, from LLM serving to recommendation systems, remain fundamentally memory bound. KV-cache capacity directly influences throughput, latency and accelerator utilization. As model sizes increase, the ability to intelligently expand and share memory resources becomes a strategic infrastructure advantage.
Structera X has been optimized specifically for hyperscale deployments through deep customer collaboration, workload-aware architecture decisions and broad ecosystem integration. The platform is designed to address hyperscaler priorities around scale, utilization, operational efficiency and total cost of ownership. Strong customer engagement, expanding deployment activity and growing design-win momentum underscore the increasing role of CXL-based memory architectures in next-generation AI infrastructure.
A key element of this evolution is the ultra-high-speed interface connecting compute and memory resources. As AI clusters continue to grow, faster interconnect technologies will ensure memory expansion solutions keep pace with increasingly demanding training and inference workloads, enabling memory resources to operate more like a flexible, scalable pool than isolated server-attached assets.
Marvell continues to offer the industry’s most comprehensive CXL portfolio spanning memory expansion, near-memory acceleration and memory pooling technologies. This breadth is a decisive advantage: rather than stitching together point products, customers gain a single, cohesive architecture with Structera X, Structera A and Structera S at its core—that scales and evolves alongside the most demanding hyperscale AI infrastructure requirements. It is this end-to-end portfolio, backed by deep hyperscaler collaboration, that positions Marvell as a driving force in the industry’s shift toward disaggregated, composable memory.
Complementing Structera X memory expansion, the Marvell Structera A near-memory accelerator brings compute directly to the data. The Structera A 2504 integrates 16 Arm® Neoverse® V2 cores running at 3.2 GHz and acts as a “server within a server,” offloading time-sensitive, memory-bound workloads such as AI inference, deep-learning recommendation models (DLRM), vector search and in-memory databases from the host CPU. With support for up to 4TB of capacity and 200 GB/s of memory bandwidth, plus inline LZ4 compression and XTS-AES 256-bit encryption, Structera A processes data where it resides to reduce data movement, boost throughput and free host cores for other tasks. In Marvell demonstrations, Structera A accelerators have cut vector-search times by more than 5x, while a trio of devices processed more queries per second than a leading server CPU at lower latency. Together, Structera X memory expansion and Structera A near-memory acceleration give hyperscalers a cohesive toolkit for scaling and optimizing memory across increasingly disaggregated AI infrastructure.
Completing the portfolio, the Marvell Structera S CXL switch makes rack-scale memory pooling a reality. The Structera S 30260 is a 260-lane, PCIe Gen 6 / CXL 3.x switch that connects 16 or 32 CPUs or GPUs to as much as 48TB of shared memory, delivering up to 4TB/s of aggregate bandwidth at under 460ns round-trip latency, with support for both DDR5 and DDR4. By enabling memory to be dynamically pooled, shared and disaggregated across the rack, Structera S drives dramatically higher utilization—Marvell benchmarks show up to 16x better scalability than local memory, 72% lower latency than RDMA-based pooling and, in GPU configurations, up to 4.8x higher inference throughput and an 82.7% reduction in time-to-first-token. Working in concert with Structera X memory expanders and Structera A near-memory accelerators, Structera S turns memory into a truly composable, rack-level resource, the foundation of the disaggregated AI infrastructure this architecture is built to enable.
The Marvell CXL ecosystem continues to expand across the full data center stack:
# # #
This blog contains forward-looking statements within the meaning of the federal securities laws that involve risks and uncertainties. Forward-looking statements include, without limitation, any statement that may predict, forecast, indicate or imply future events or achievements. Actual events or results may differ materially from those contemplated in this blog. Forward-looking statements are only predictions and are subject to risks, uncertainties and assumptions that are difficult to predict, including those described in the “Risk Factors” section of our Annual Reports on Form 10-K, Quarterly Reports on Form 10-Q and other documents filed by us from time to time with the SEC. Forward-looking statements speak only as of the date they are made. Readers are cautioned not to put undue reliance on forward-looking statements, and no person assumes any obligation to update or revise any such forward-looking statements, whether as a result of new information, future events or otherwise.
Tags: AI infrastructure, Data Center, Data Processing, server connectivity, Cloud