news

WEKA and Oracle Cloud Infrastructure Validate 10x Throughput Gains for Long-Context AI Inference

Jun 9, 2026 · Source: cision
WEKA and Oracle Cloud Infrastructure Validate 10x Throughput Gains for Long-Context AI Inference

Joint benchmarks on OCI H100 infrastructure showed 10x more concurrent users, 10x higher token throughput, and 7x more tokens served without adding GPUs

CAMPBELL, Calif., June 9, 2026 -- WEKA, the AI data and memory infrastructure company, today announced production-scale benchmarks that show how organizations can improve the economics of long-context AI inference by serving more users and tokens on the same GPU footprint. The benchmarks show that WEKA's NeuralMesh™ platform with Augmented Memory Grid™ on Oracle Cloud Infrastructure (OCI) serves 10x more concurrent users, delivers 10x higher token throughput, and produces 7x more tokens per GPU than DRAM-only configurations without adding infrastructure. The results were validated on a nine-node OCI bare-metal H100 cluster with 100,000-token context windows.

(PRNewsfoto/WekaIO)

"Enterprise AI workloads are pushing context windows and GPU utilization to new limits," said Pablo Selem, senior director, software development, Oracle Cloud Infrastructure. "These benchmarks show how WEKA's NeuralMesh platform with Augmented Memory Grid on OCI helps remove memory bottlenecks so customers can support larger, more demanding inference workloads without simply adding more GPUs."

Three Outcomes That Change the Math on Inference

Validated at production scale on a bare-metal H100 cluster (nine nodes, 72 GPUs, 100,000-token context windows, thousands of concurrent users), NeuralMesh with Augmented Memory Grid on OCI delivered:

  • 10x more concurrent users served, without adding infrastructure. NeuralMesh with Augmented Memory Grid scaled past 5,000 concurrent users vs. about 600 for DRAM-only configurations. This eliminates the failure cliff that hits when cache saturates by expanding the active cache working set from 8.64 TiB of DRAM to 287 TiB of usable NVMe. In addition, more users per GPU means the same investment stretches further.
  • 10x higher token throughput. More output from every GPU in the cluster. On OCI, NeuralMesh with Augmented Memory Grid reached approx. two million tokens per second, compared to under 200,000 for the DRAM-only baseline. For product teams running real-time AI features, including search, summarization, code assist, and multi-turn agents, the throughput determines the ceiling for how many users can be served, how fast features respond, and how much revenue the infrastructure can support.
  • 7x more tokens served. Lower cost per token at scale. NeuralMesh with Augmented Memory Grid served five billion tokens, compared to 700 million for the DRAM-only baseline, in a single one-hour, 2,400-user test. For organizations running agentic workflows, DRAM saturation quietly drains GPU capacity through constant recomputation, creating a direct hit on cost per token and ROI.

"Inference is bottlenecked by how much effective memory is available to GPUs," said Liran Zvibel, CEO of WEKA. "These results prove that AI token economics aren't solved by hardware alone; they're solved by eliminating the memory wall that has been the real ceiling on what existing hardware can do. NeuralMesh with Augmented Memory Grid running on OCI brings orders of magnitude more tokens to customers in an extremely cost-efficient way."

Transforming AI Economics with Context Memory Infrastructure

As inference demand grows, AI infrastructure inefficiencies compound. Every key-value (KV) cache eviction is a tax: on GPU cycles, latency, user experience, and the cost of every token served. For long-context and agentic workloads, where inputs routinely run to 100,000 tokens or more, that tax is not a rounding error. It is a direct hit on the unit economics of every organization running production AI.

Augmented Memory Grid, a capability of NeuralMesh, solves the problem at the architectural level by decoupling KV cache from local GPU memory and storing it in a high-performance token warehouse accessible across the cluster. Any host can serve any session with cache hits intact, eliminating rigid session stickiness while delivering superior performance to DRAM, improving load balancing, and enabling clean horizontal scaling as concurrency grows. The result is persistent context memory for AI agents and the cost lever that makes long-context inference economical to run at scale.

Production-Grade Proof

OCI published the full benchmark methodology, system configuration, and results on its AI & Data Science blog on May 13, 2026. The benchmarks, executed on a nine-node OCI bare-metal H100 cluster, move beyond the prior phase of validation, which demonstrated 1000x more KV cache capacity and up to 20x faster time to first token at 128,000 tokens. This latest phase tests the full economics of inference in production: concurrency density, sustained throughput, cache persistence, and service level objective (SLO) stability when demand spikes under high load.

Available on Oracle Marketplace

NeuralMesh with Augmented Memory Grid is generally available to WEKA customers and on the Oracle Marketplace, with OCI as WEKA's exclusive cloud launch partner. Organizations running long-context inference on OCI can deploy a validated, production-ready architecture today. For more on the OCI and WEKA Augmented Memory Grid benchmark, read the OCI blog: https://blogs.oracle.com/ai-and-datascience/scaling-long-context-inference-on-oci-with-wekas-augmented-memory-grid.

About WEKA

WEKA is the AI data and memory infrastructure company transforming the economics of agentic AI. Its NeuralMesh™ platform unifies high-performance data storage with extended GPU memory, giving enterprises, AI cloud providers, and AI builders a single foundation for training, inference, and agentic workloads. With Augmented Memory Grid, NeuralMesh extends GPU memory capacity by 1000x, accelerates time to first token by up to 20x, and delivers 10x more concurrent users from the same GPU footprint, proven in production benchmarks. Trusted by 30% of the Fortune 50, WEKA enables organizations to scale AI faster, optimize GPU utilization, and reduce the cost of every token served. Learn more at www.weka.io or connect with us on LinkedIn and X.

WEKA and the W logo are registered trademarks of WekaIO, Inc. Other trade names herein may be trademarks of their respective owners.

Cision View original content to download multimedia:https://www.prnewswire.co.uk/news-releases/weka-and-oracle-cloud-infrastructure-validate-10x-throughput-gains-for-long-context-ai-inference-302793740.html

Comments (0)

Login to join the conversation

Login / Register

No comments yet. Be the first to comment!

More to Read

Tredence Launches Domain Native Forward Deployed Engineering to Close the Last Mile of Enterprise AI
news
Jul 27, 2026 1 min

Tredence Launches Domain Native Forward Deployed Engineering to Close the Last Mile of Enterprise AI

The FDE practice builds an elite class of engineers at the intersection of domain expertise and data & AI, solving enterprises hardest business problems. BENGALURU, India and SAN JOSE, Calif. , July 27, 2026 -- Tredence, the world s leading data & AI services company, today announced the launch of its Forward Deployed Engineering (FDE) practice, committing to build a dedicated pool of 200 FDEs over the next 12-18 months. Through this practice, the company intends to build the most domain-native engineering capability in the market, helping clients move faster from problem to impact.

Mapex AI Accelerates Global Growth in Geospatial Intelligence through Strategic Government Projects and International Partnerships
news
Jul 27, 2026 1 min

Mapex AI Accelerates Global Growth in Geospatial Intelligence through Strategic Government Projects and International Partnerships

NOIDA, India , July 27, 2026 -- Mapex AI Private Limited, a leading geospatial technology company, continues to strengthen its global presence by delivering innovative, enterprise-grade geospatial solutions that transform complex spatial data into actionable intelligence. The Company has developed deep expertise in delivering GIS-based Master Planning, Land Information Systems, Property Tax Solutions, Utility Mapping, Digital Twins, Navigation, High-Precision Surveys, Enterprise GIS, and Geospatial Data Infrastructure to support India s flagship digital governance and infrastructure programmes.

Highstar Launches Full-Chain Battery Cell Portfolio for AI Data Centers
news
Jul 27, 2026 1 min

Highstar Launches Full-Chain Battery Cell Portfolio for AI Data Centers

NANTONG, China , July 27, 2026 -- Highstar unveiled its full-chain battery cell solution for artificial intelligence data centers (AIDCs) at the 2026 GGII Energy Storage Industry Summit. The portfolio spans three critical power layers: grey-space UPS and high-voltage direct current (HVDC) systems, white-space battery backup units (BBUs), and grid-side energy storage. Targeted cells for grey-space and white-space challenges As AI workloads increase rack density and load volatility, data centers require faster response, dependable backup, controlled temperature rise and stronger lifecycle performance. Highstar s solution centers on two products for the layers closest to computing loads.

Ookla Study in Manila: Carrier VoLTE Networks Prove to Outperform OTT Apps in Voice Quality and Reliability
news
Jul 27, 2026 1 min

Ookla Study in Manila: Carrier VoLTE Networks Prove to Outperform OTT Apps in Voice Quality and Reliability

MANILA, Philippines , July 27, 2026 -- Ookla s comprehensive controlled network test reveals that traditional mobile operator networks deliver a measurably superior voice experience compared to Over-the-Top (OTT) applications like WhatsApp, particularly in critical areas such as audio fidelity, weak signal resilience, and call reliability. The study, which evaluated the networks of the Philippines three major mobile operators Smart Communications, Globe Telecom, and DITO Telecommunity , demonstrates that carrier-managed VoLTE infrastructure remains the gold standard for consistent, high-quality communication, outperforming OTT voice services.

Aolani and Rafay Collaborate on One of the Industry's First NVIDIA DSX OS Deployments on NVIDIA GB200 NVL72 Infrastructure
news
Jul 27, 2026 1 min

Aolani and Rafay Collaborate on One of the Industry's First NVIDIA DSX OS Deployments on NVIDIA GB200 NVL72 Infrastructure

The collaboration demonstrates how next-generation AI infrastructure can be transformed into production-ready AI platforms for enterprise and cloud providers. SUNNYVALE, Calif. and SINGAPORE , July 26, 2026 -- Rafay Systems , a leading platform provider for modern infrastructure and AI workloads and a member of NVIDIA Inception , today announced a strategic collaboration with Aolani to deliver one of the industry s first deployments of NVIDIA DSX OS running on NVIDIA GB200 NVL72 infrastructure, helping demonstrate how next-generation AI infrastructure can move from installation to production-ready AI services.

TECNO Unveiled as Title Sponsor of The SAFF Championship Bangladesh 2026, Bringing AI Innovation to South Asian Football
news
Jul 25, 2026 1 min

TECNO Unveiled as Title Sponsor of The SAFF Championship Bangladesh 2026, Bringing AI Innovation to South Asian Football

As Official Title Sponsor, TECNO joins hands with SAFF to inspire the next generation through football, innovation, and meaningful fan experiences. DHAKA, Bangladesh , July 25, 2026 -- TECNO, an AI-driven innovative technology brand, officially announced its title sponsorship of the SAFF Championship Bangladesh 2026, South Asia s premier international football tournament, during the tournament s official launch ceremony in Dhaka. Scheduled to take place from 4 17 November 2026, the championship will bring together South Asia s leading national teams, celebrating the region s passion for football while strengthening friendship, sporting excellence, and regional unity.

Tony Jaa Becomes GAC's 30-Millionth Customer - GAC Wins Global Trust with "True Craftsmanship"
news
Jul 25, 2026 1 min

Tony Jaa Becomes GAC's 30-Millionth Customer - GAC Wins Global Trust with "True Craftsmanship"

GUANGZHOU, China , July 25, 2026 -- On July 16, at the roll-off ceremony for GAC s 30-millionth vehicle, Feng Xingya, Chairman of GAC Group, handed over the key to the right-hand-drive GAC M8 PHEV (named GN8 overseas) to Tony Jaa. The milestone vehicle is headed straight for overseas markets. Thai action superstar Tony Jaa s choice reflects the trust of 30 million customers worldwide. That trust is built not on showmanship, but on GAC s solid manufacturing true craftsmanship.

From China Mobile's Call Upgrade to the Commercial Launch of "Calling + AI" by Leading Operators: AI Is Reshaping the Value of Native Calling
news
Jul 25, 2026 1 min

From China Mobile's Call Upgrade to the Commercial Launch of "Calling + AI" by Leading Operators: AI Is Reshaping the Value of Native Calling

BEIJING , July 25, 2026 -- On June 15, 2026, China Mobile announced a comprehensive upgrade to its traditional calling services, ushering in a next-generation calling experience defined by HD, intelligence, and security. This milestone not only marks a major leap in telecommunication innovation but also reflects a global, inevitable shift: the transformation of basic communication into intelligent, inclusive services. Breaking Experience Barriers and Redefining the Paradigm of Basic Calling