Akamai Bets on Distributed Cloud to Handle Asia-Pacific AI Inference Surge
Asia-Pacific enterprises are shifting artificial intelligence from pilot projects into production, and the transition is rewiring cloud infrastructure across the region. Akamai says the change demands a distributed computing model that brings inference closer to users and data, rather than routing everything through centralized hyperscale facilities.
Sean Li, managing director for Asia Pacific and senior vice president of sales at Akamai, said companies across the region are moving AI budgets from experimental work into live deployments. The shift creates infrastructure demands that traditional cloud architectures were not built to handle.
"Cloud infrastructure has generally been designed for websites, applications and centralized computing," said Jay Jenkins, chief technology officer for cloud computing at Akamai. "Production AI introduces different requirements, including low-latency inference and GPU-accelerated processing."

From Centralized Training to Distributed Inference
The distinction matters. Training large models still benefits from massive GPU clusters in centralized data centers. But inference -- the phase where models actually serve users -- generates different traffic patterns. Instead of large bulk transfers between GPU clusters, Akamai is seeing smaller, continuous requests distributed across many locations. These workloads are sensitive to latency and need processing near the data and users they serve.
Jenkins said production inference is creating a new kind of network demand. The company plans to deploy thousands of NVIDIA Blackwell GPUs across its distributed infrastructure to meet it. Akamai has also built what it calls Inference Cloud, described as the first global implementation of NVIDIA AI Grid. The platform routes inference workloads across Akamai's edge network.
"This allows us to bring compute closer to where the data and users are located, helping customers improve latency while also making AI workloads more cost efficient," Jenkins said.
The partnership with NVIDIA covers Blackwell GPUs on Akamai Cloud and NVIDIA AI Grid intelligent orchestration. Both companies aim to place and manage AI workloads across distributed GPU infrastructure as demand shifts.
"Inference has become the most compute-intensive phase of AI -- demanding real-time reasoning at planetary scale," said Jensen Huang, chief executive officer of NVIDIA. "Together, NVIDIA and Akamai are moving inference closer to users everywhere, delivering faster, more scalable generative AI and unlocking the next generation of intelligent applications."
Asia-Pacific's Complex Infrastructure Picture
The region presents a fragmented picture. Markets vary widely in cloud maturity, connectivity and data center capacity. AI workloads are placing additional pressure on computing and energy resources. Data sovereignty adds another layer. Companies must balance application performance with local requirements for data residency, security and compliance.
Processing AI workloads closer to users could improve fraud detection, personalized banking, commerce recommendations and media processing. Keeping inference near the underlying data can also reduce the movement of sensitive or regulated information.
Akamai is positioning its distributed cloud network as an infrastructure layer spanning three tiers. Centralized facilities handle model training and large-scale experimentation. Regional cloud resources address market-level performance and data residency needs. Edge infrastructure handles real-time inference closer to users, devices and applications.
The largest changes are emerging among digital-native businesses, particularly in e-commerce. Jenkins expects multi-agent systems to increase demand for different types of computing capacity across multiple locations.

Security Expands With Autonomous Agents
AI adoption expands the enterprise attack surface across applications, APIs, data pipelines, infrastructure and autonomous agents. Reuben Koh, security technology and strategy director for Asia-Pacific and Japan at Akamai, identified identity and access management as a key issue as AI systems become more autonomous.
Sensitive information can enter AI services through prompts, processing pipelines and third-party tools. AI applications also rely heavily on APIs connecting models with business data and operational systems. Other risks include prompt injection, data poisoning and model misuse. Access controls become more consequential when agents inherit user privileges or gain permission to act across business systems.
Autonomous agents can generate continuous interactions across software services, mobile applications and connected devices. Machine-to-machine exchanges may also occur without direct human involvement. Security systems therefore need to attribute actions to individual users or AI agents. Organizations also require visibility into prompts, responses and data transfers as they occur.
Akamai Workforce Protector, formerly LayerX, is designed to govern interactions across AI, software-as-a-service, web and private applications. Akamai acquired LayerX and integrated its technology into the security portfolio. The platform uses a browser extension for Chrome, Edge, Firefox and Safari, allowing companies to retain their existing browsers and network architecture.
It captures the context surrounding prompts, text entries, clipboard activity, file transfers and browser plug-ins. Policies can allow, warn, redact or block an interaction based on the user, application, data type and assessed risk. A management console centralizes policy creation, investigations and reporting. The platform's cloud intelligence component analyses applications, identities, extensions and AI activity.
Workforce Protector can discover AI applications, desktop tools and agents used across an organization. It can also attribute activity to human and AI identities, creating records for investigations and compliance reporting. Companies can enforce read-only sessions and watermarking on unmanaged devices used by employees or contractors. The system also assesses browser extensions based on their permissions, behavior and risk scores.
These interaction-level controls complement network and endpoint security. Secure access service edge and security service edge platforms focus on network traffic and connectivity. Endpoint tools cover files and device processes. Browser-level monitoring adds context for prompts, clipboard activity and other fileless interactions.
Real-World Deployment Shows Early Gains
GoVeda, a patent search provider, has migrated AI workloads to Akamai Cloud to support searches across more than 220 million patent publications worldwide. The company uses G8 CPU compute for inference and plans to add NVIDIA RTX 6000 Pro GPUs. The deployment also uses managed containers, object storage, block storage and managed databases.
GoVeda reported a 30 percent performance improvement and a 20 percent reduction in infrastructure costs after the migration.
"Akamai's infrastructure just works," said Cheng Tai, chief executive officer and co-founder of GoVeda. "It's stable, reliable, and allows us to scale as our needs evolve."
The case illustrates the practical impact of the distributed model. Rather than routing patent-search queries to a distant centralized cluster, GoVeda's workloads run closer to its users and data. The result is lower latency and lower cost.
The Edge Becomes the Default
The shift Akamai describes aligns with broader industry movements. Cloudflare, Fastly and other edge-platform providers have also been building out GPU-equipped edge locations for inference. The pattern suggests that for production AI, the edge is no longer an add-on -- it is becoming the primary compute tier for inference workloads.
For enterprises in Asia-Pacific, the implications are direct. Data sovereignty laws in countries including Indonesia, Australia and India require certain data to stay within national borders. A distributed cloud that can place inference in specific jurisdictions while maintaining a unified control plane offers a compliance path that pure hyperscale clouds struggle to provide.
Energy constraints add urgency. Data center capacity in major Asia-Pacific hubs -- Singapore, Tokyo, Sydney -- is tight. Distributing inference to edge sites in secondary markets can relieve pressure on core facilities while improving response times for local users.
Akamai's bet is that this model will generalize beyond Asia-Pacific. The same forces -- latency sensitivity, data sovereignty, energy constraints, agent-driven traffic -- are appearing in Europe and the Americas. If the pattern holds, the next wave of cloud infrastructure investment will flow to the edge, not just the core.
For more on how cloud infrastructure is adapting to AI workloads, see our Cloud & Edge Computing coverage. The original Akamai briefing is available at SecurityBrief Asia.