×
Community Blog Alibaba Cloud ESA at Apsara 2026: The Full Edge-for-AI Lineup - Seven Shipped Capabilities and Four Stage Directions, Under Three Pillars

Alibaba Cloud ESA at Apsara 2026: The Full Edge-for-AI Lineup - Seven Shipped Capabilities and Four Stage Directions, Under Three Pillars

ESA at Apsara 2026: eleven capabilities - seven shipped, four from the conference - grouped under AI acceleration, AI security, and agent runtime.

Introduction

The edge was built around a simple assumption: content is mostly static, requests are short, sessions don't hold state, and a hit ratio is the number that matters. Everything in the classic CDN - caching, TTLs, purge rules, prefetch, origin shield - was engineered against that shape of traffic.

AI traffic breaks every one of those assumptions. An inference session streams for tens of seconds. An agent loop fires dozens of small calls per turn, each metered in tokens. Requests carry prompts and completions that should not land verbatim in an access log. The unit of cost is tokens, not bytes, so "cache hit" stops being the right word. A new class of visitor shows up who is neither a buyer nor a hacker - a crawler assembling a training corpus. And workloads that once quietly tolerated handing a private key to the edge now have to fend off store-now-decrypt-later quantum attacks and stricter key-custody rules at the same time.

The answer ESA gave at Apsara 2026 (September 22, Hangzhou) was not incremental. The keynote, "ESA's full evolution: a global runtime foundation for AI Agents," framed the shift as moving from the CDN you know toward a foundation that runs AI applications at the edge across four dimensions - acceleration, security, compute, and network. The 13:50 new-product session then cut the shipping work into three pillars: AI acceleration, AI security, and the Agent runtime. That three-pillar split is the backbone of this article.

The lineup at a glance

Xnip2026_10_06_17_52_39

Pillar Capability Status (2026) One-line role
AI acceleration POST-request caching Product updates - March Cache reusable AI POST responses (embeddings, idempotent inference) at the edge
AI acceleration Global network · last mile Stage direction Extend delivery into cross-border, emerging, and constrained-network regions
AI security AI Gateway (transport & perf slice) Product updates - April Edge-proximate entry, connection reuse, streaming-friendly handling
AI security AI Crawler Management Product updates - April Identify and govern AI crawlers as a distinct traffic class
AI security Post-Quantum Encryption (PQC) Product updates - April Hybrid X25519MLKEM768 key exchange, on by default, site-wide
AI security Keyless Certificates Product updates - April HTTPS acceleration without uploading your private key
AI security Account Takeover (ATO) protection Product updates - March AI/ML detection of credential abuse and account takeover
AI security SNI allowlist Product updates - July Validate SNI against the Host header at TLS handshake
Agent runtime AI Gateway (control-plane slice) Product updates - April Unified proxy, observability, guardrails, failover at the edge
Agent runtime Edge Containers Stage direction Run containerized agent tools next to the POP that ends the user session
Agent runtime Domain as the entry Stage direction DNS → edge routing → delivery engineered as one path, one policy
Agent runtime ESA Logs + AI Agents Published capability (blog) Natural-language ops over logs, config, and incidents

Seven of these are already "supported" entries in ESA's product updates. The last mile, Edge Containers, and Domain-as-the-entry were shown on the Apsara stage and are not yet release-notes entries - treat them as directions, not shipping features. ESA Logs + AI Agents sits between: described as a published capability in an international ESA blog, but not among the seven 2026 release-notes items above.


Pillar 1 - AI Acceleration

1.1 POST-request caching

Solution

For years, CDN caching effectively meant GET. POST was treated as non-cacheable by default, which was correct for form submits but wrong for a growing share of AI traffic: embedding calls, semantic-search hits, idempotent inference, model-capability probes, and repeated tool calls inside an agent loop all arrive as POST yet return responses that are cacheable in practice. Applications re-paid full round-trip and provider cost for the same result every time.

What's New

Listed as supported in ESA's product updates for March 2026. The release-notes entry reads: "Once POST caching is enabled, nodes can cache POST response bodies and serve them to subsequent identical requests, lowering origin load and improving response speed."

Key Capabilities

  • Edge nodes cache POST response bodies and replay them to matching requests
  • Cuts repeated origin/provider round-trips for idempotent AI calls
  • Composable with existing cache rules keyed on header, path, and body
  • Works for the request/response shape real AI APIs use, not just static assets

Benefit for Your Architecture

Embedding and similarity-search endpoints get the single biggest win - the same query repeated across users and agents stops re-crossing the ocean on every hit. Agent loops that fire the same tool or classifier call many times per turn collapse most of those calls into a cache replay. And because it's an edge rule, it applies uniformly to every AI endpoint you front, not one hand-patched service.

Design Pattern

Cache POST only where responses are genuinely idempotent for a given request body, and key the cache on the fields that actually change the answer (model name, prompt/version, retrieval index). Set short TTLs and keep a purge path for model or index updates. Treat anything carrying per-user or personalized output as out of scope, and make sure cached bodies never leak identifiers. This rule pairs directly with the AI Gateway below: route through the gateway, let it decide cacheable vs. pass-through.


Pillar 2 - AI Security

2.1 AI Gateway (the transport and performance slice)

The AI Gateway is a single release-notes item that spans pillars; its edge-proximity and streaming handling are an acceleration story, and its control plane is an Agent-runtime story (2.5). Both are covered so nothing is counted twice - here is the fast-path half.

Solution

Today, an AI application that talks to more than one model provider usually runs its own gateway - a slim pass-through service (essentially a small proxy) sitting in front of OpenAI, Anthropic, Google, an open-weights backend, and a couple of embedding or safety models, doing little more than rotating API keys, throttling, retrying, and logging. That gateway lives in one region and reinvents the rate limiting, retry semantics, and observability the CDN already had years ago. The result is a second, less-matured traffic plane running parallel to the site.

Xnip2026_10_06_17_56_13

What's New

Listed as supported in ESA's product updates for April 2026, described in the release-notes entry as: "The performance half of that sentence is the acceleration story: AI requests now enter at the edge node nearest the user, and the transport is engineered for streaming."

Key Capabilities

  • One edge entry point for every provider an application calls
  • Connection reuse and keepalive for many small, repeated calls
  • Streaming-friendly handling - SSE, chunked transfer, HTTP/2, visible first byte
  • Failover happens at the edge rather than in a client SDK retry block

Benefit for Your Architecture

The agent loop stops paying a cross-region tax on every hop. A chat turn calls one model; an agent turn may call retrieval, tools, classifiers, and a synthesizer - a dozen small calls. A centralized gateway routes every one back to a single region; putting the entry point at the user's edge node removes that multiplier.

Security and observability are not built twice. Model APIs come with their own auth and limits, but teams cheaply skip a layer at the app tier assuming "the gateway handles it." When the gateway is the edge, and the edge already runs WAF and DDoS, the defense exists in one place.

Design Pattern

Point clients and agent orchestrators at the gateway endpoint only; keep provider secrets in edge configuration, never on the client. Enable request/response logging on the AI path with retention tied to data sensitivity - prompts and completions often carry personal or proprietary content, so scope log access apart from normal site logs. Restrict upstream providers or self-hosted inference endpoints so only ESA's converged back-to-origin addresses can reach them (see Origin Protection in this series, Part 4).

2.2 AI Crawler Management

Solution

Between "legitimate SEO crawling" and "malicious scraping" sits a new class of visitor - the AI crawler. GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBbot and their peers each have a user-agent, a politeness policy, and an intent, usually to assemble or refresh a training corpus. A content site now has to answer each family separately: is this visitor bringing me customers, or just taking my content? Should the referral path an AI answer engine sends me be treated differently?

Classic bot management had two buckets - allow known bots, challenge or block unknown ones - and that classification is no longer enough.

What's New

Listed as supported in ESA's product updates for April 2026. The entry reads: "A dedicated detection engine plus flexible access-control policies precisely identify mainstream AI crawlers, manage their permissions differentially, and analyze access data for IP protection and resource optimization."

Key Capabilities

  • A dedicated AI-crawler detection engine, not a bare user-agent string match
  • Per-category, per-host, per-path policies that compose into allow / throttle / challenge / block
  • Access analytics suited to licensing, compliance, and legal conversations
  • Coexists with existing bot management, WAF, and rate limits on the same edge

Benefit for Your Architecture

"Should I care about AI traffic?" becomes a policy decision instead of a vibe. Every site can now state a position - fully open, conditional (throttle plus a licensing threshold for heavy users), default-deny with named exceptions, or metered (allow first, license from the access data later). The engine and the policies are the primitives that make those positions enforceable.

Reachability and IP protection stop being a false trade-off. Blocking all crawlers costs you visibility in AI answer engines. Allowing everything hands your corpus to training with no licensing trail. You can split by category - welcome the ones that return customers, govern the ones that only take.

Bulk scraping stops hitting your origin. A bot 404ing hundreds of paths a second isn't "a little noisy," it's what takes down autoscaling at 3 a.m. Decide at the edge and the cost never reaches the origin.

Design Pattern

Content sites decide the stance first, then encode the policy; SaaS and API-first products should treat crawler management as an endpoint automation-access policy, not just a page-level one. Combine with SNI allowlist (2.5, July 2026) so crawlers can't dodge classification through a TLS-layer mismatch. Keep robots.txt consistent with the enforced policy - the detection engine is more authoritative, but a semantic conflict weakens the defensibility of any good-faith claim.

2.3 Post-Quantum Encryption

Solution

The quantum threat to today's TLS traffic is not "broken suddenly tomorrow." It is record-now, decrypt-later: capture ciphertext today, store it, and unravel it when a strong enough quantum machine exists. Anything with a confidentiality horizon beyond roughly a decade is already in range - medical records, financial instructions, government filings, legal discovery, RAG source corpora, model weights, and long-lived sessions.

The defense is swapping the TLS handshake to a quantum-resistant key exchange. What historically stopped adoption was never the algorithm - it was the cost: upgrade server libraries, gain client support, re-sign certificates, roll out gradually. ESA pushes that cost near zero because the handshake happens at the edge, and changing the edge changes every site at once.

What's New

Listed as supported in ESA's product updates for April 2026, and on by default. The entry reads: "Native post-quantum cryptography (PQC): the X25519MLKEM768 hybrid key exchange protects client-to-edge traffic against quantum attacks. Enabled by default site-wide, zero configuration. "

Xnip2026_10_06_17_57_37

Key Capabilities

  • X25519MLKEM768 hybrid key exchange (classic X25519 + NIST-standardized ML-KEM)
  • Site-level default on - no action required
  • Transparent negotiation for new browsers, safe fallback to classic suites for old clients
  • Covers the client ↔ edge segment - the "first mile," and the segment most exposed to passive capture

Benefit for Your Architecture

Closes the record-now window on the segment most likely to be recorded. The user-to-edge leg crosses the public internet, hostile networks, and physical links - the natural surface for passive capture. Making the first mile quantum-resistant is where PQC has its highest density of benefit.

"PQC support" on procurement questionnaires is a yes in advance. More enterprise and regulatory questionnaires ask for this line. On ESA it's simply on; no scheduling, no rollout campaign.

No three-way change to client, certificate, and origin code. On the edge-terminated segment the painful parts of a PQC migration dissolve: no re-issuing certificates, no forced client upgrades, no origin library swap. Old clients still complete the handshake, just on a classic suite.

Design Pattern

For any site with a long confidentiality horizon - finance, health, public sector, model weights, RAG corpora - keep the negotiated cipher name (X25519MLKEM768) from the access log as the audit artifact that "we already run PQC." For true end-to-end coverage, the edge ↔ origin leg is a separate transport that needs origin-side PQC to extend - the default site-wide setting covers the first mile only. Deep-dive and enablement steps: "What is Post-Quantum Encryption - and How to Enable It on ESA".

2.4 Keyless Certificates

Solution

A whole class of business was never standing at the CDN door - not because the edge wasn't fast, but because a compliance rule said "the private key does not leave this boundary." Finance, public sector, healthcare, defense, and the strictest tier of enterprise procurement all require the TLS private key to stay inside a self-managed HSM/KMS. The traditional onboarding flow asks you to upload both certificate and key to the platform; that single step disqualifies the whole service.

Keyless separates "the certificate is at the edge" from "the private key is at the edge." ESA uses your certificate for HTTPS acceleration at the edge; whenever a handshake needs a signature, it calls a KeyServer you deploy yourself. The private key never leaves your trust boundary - it does not enter Alibaba Cloud's platform either.

What's New

Listed as supported in ESA's product updates for April 2026. The entry reads: "With a self-hosted KeyServer and keyless certificate configured in the console; HTTPS acceleration works private key to a third-party platform — minimizing key leakage and compliance risks"

Key Capabilities

  • You run the KeyServer; the edge calls it whenever it needs a signature
  • The private key stays in your HSM/KMS/datacenter - your choice
  • Full edge HTTPS acceleration is preserved - signing is small and fast, not a traffic path
  • Materially lowers the compliance risk of a key leaking from a third-party platform
  • Configured in the ESA console; the KeyServer address is a value you provide, not one we issue

Benefit for Your Architecture

Unlocks the tier of workloads that previously could not go to the edge at all. That's the direct commercial effect: finance/public-sector/healthcare customers once excluded by key-custody rules can now take edge acceleration, WAF, DDoS, and PQC at the same time.

A compromised platform never yields your key. The key is not on a third party, so a third-party breach is not your incident. The audit story is simpler and the incident playbook is shorter.

No rip-and-replace of your CA/PKI lifecycle. Heavily regulated enterprises already have designated CAs and long-horizon certificate governance. Keyless uses the certificate you have today; you don't stand up a shadow PKI on a cloud platform.

Design Pattern

A hardened host on a low-exposure subnet inside your network is enough for the KeyServer - it answers signature challenges, it doesn't carry traffic, so its throughput and network surface are small. Register it as a Keyless configuration in the console, then treat it like an origin: allow only ESA's converged back-to-origin address ranges to reach the KeyServer (whitelist procedure in Part 4, Origin Protection). Before rollout, take an end-to-end handshake-latency baseline - the signing hop adds one internal round-trip to TLS setup. Note Keyless and PQC are orthogonal and stack cleanly: PQC guards the transport, Keyless guards key custody.

2.5 Account Takeover (ATO) protection

Solution

AI applications multiply the attack surface on accounts. Agent sessions and API keys are long-lived, programmatic, and easy to reuse; credential stuffing, brute force, phishing, and infostealer malware all target them. A single leaked key can quietly drive a large token bill or impersonate a user through an agent loop. Traditional WAF rules see valid requests and let them through; the abuse is behavioral, across sessions and devices, not in any one request.

What's New

Listed as supported in ESA's product updates for March 2026. The entry reads: "Account Security uses AI and machine-learning to detect account-takeover attacks, defending against credential stuffing, brute force, phishing, and infostealer compromise."

Key Capabilities

  • AI/ML risk scoring that reasons across login and access patterns rather than single-request signatures
  • Detection of credential stuffing, brute-force, phishing-derived, and malware-stealer abuse
  • Runs on the same edge as WAF, bot management, and rate limits
  • Applies to the account, checkout, and API-auth entry points your AI app exposes

Benefit for Your Architecture

Abuse detection lands where the traffic already terminates. Because ATO scoring runs at the edge, you don't have to ship login telemetry to a separate security vendor to get a risk verdict. AI apps with user accounts and API keys get a defense built for exactly their failure mode - a stolen key replayed from a new geography is an ATO signal, not a WAF signature. And the scoring is per-request but stateful, so it complements (does not replace) your rate limits.

Design Pattern

Turn ATO protection on for authentication, password reset, token-exchange, and high-value API endpoints. Feed it real client IP (behind the edge, preserve it properly) so geographic and device signals are accurate. Pair with rate limits and device/bot signals already available at the edge; treat ATO verdicts as one input into step-up auth or key revocation. For agent platforms, scope per-key spend alarms so a hijacked key is caught by cost anomaly as well as behavioral score.

2.6 SNI allowlist

Solution

Automated clients - crawlers included - increasingly exploit mismatches at the TLS layer: a handshake announces one SNI but the HTTP request carries a different Host header, or the client simply probes with an SNI that doesn't belong to your site. That mismatch lets abusive traffic slip past host-based routing and classification.

What's New

Listed as supported in ESA's product updates for July 2026. The entry reads: "A new SNI allowlist validates the SNI used in the SSL/TLS handshake against the Host header in the HTTP request, giving sites a wider range of protection policies."

Key Capabilities

  • Enforces SNI ↔ Host consistency at handshake time
  • Rejects or flags mismatched probes before they consume origin resources
  • A primitive you can layer under bot, crawler, and access policies
  • Same edge, same policy plane as WAF and crawler management

Benefit for Your Architecture

SNI validation closes the TLS-layer evasion route that AI-crawler and bot policies depend on staying honest - an abusive agent that mislabels its SNI can no longer present itself as traffic for a permitted hostname. It's a small, cheap primitive, but it makes the enforcement of everything above it (crawler policy, bot rules, routing) materially harder to bypass.

Design Pattern

Turn on SNI-to-Host consistency for domains that should only ever be reached with a matching SNI; use it under AI Crawler Management (2.2) so crawler classification can't be dodged at the handshake. Because it's an allowlist, enumerate the legitimate SNI/Host pairs for multi-tenant or apex/CNAME setups before enforcing, to avoid false rejects on correctly-configured clients.


Pillar 3 - Agent Runtime

3.1 AI Gateway (the control-plane slice)

Same April 2026 release-notes item as 2.1; this is the "unified API proxy + observability + guardrails" half.

Solution

Once the AI path is terminated at the edge, the natural next step is to run the model control plane there too - a single place that fronts every provider, sees every call, and applies policy consistently. That removes the parallel application-side gateway and the drift between "what the site enforces" and "what the AI path enforces."

What's New

The April 2026 AI Gateway entry is the same text as above - "one edge proxy for AI calls, with request/response observability, inherited edge security, and provider-agnostic policy."

Key Capabilities

  • Auth, quota, and routing policy identical across providers - swap vendors without re-coding
  • Failover decided at the edge, not in client SDK retry logic
  • Per-request token and cost signal captured close to the user
  • Shares the site's single log, dashboard, and alert pipeline

Benefit for Your Architecture

Switching models is a config change, not a code migration. Moving from vendor A to B for price, quality, or compliance no longer touches the SDK or rewrites retry logic - you change one edge policy. Token-level economics get collected nearest the user. Edge instrumentation lets product and finance see cost by region and cohort for the first time, not by datacenter.

Design Pattern

Covered with 2.1: single gateway endpoint, provider keys at the edge, AI-path logging scoped by sensitivity, upstream locked to ESA's converged origin ranges.

3.2 – 3.4 Stage directions (no release-notes entry yet)

The product session and the customer talks highlighted three runtime directions that are on the roadmap and demoed on stage(but not yet officially supported).

Edge Containers. Shown by Shanghai Qianyao Information Technology founder & CTO Zhenyi,Yang, in "Building global applications on ESA AI Gateway + Edge Containers." The idea is containerized workloads - agent tools, retrieval functions, small-model inference, glue - scheduled onto the same edge substrate that already carries delivery, security, and the AI Gateway. Its value is completing what the gateway started: if the agent's tools and retrieval still live in one region, the "per-hop cross-region tax" is only partly removed; pulling compute next to the session that ends the user removes the multiplier fully.

Global network · last mile. Product Director, Jimmy Wang's session "ESA global network: extending cloud and AI applications to the last mile" pushed the transport layer into places the public internet handles poorly - cross-border, emerging markets, constrained networks. This is the same engineering posture as Part 5 of this series (mainland-China access optimization), applied to other geographies. The direction is to turn "globally reachable" from an adjective into something you can put behind an SLA.

Domain as the entry. Xiaojie, Sun's session "Domain as the entry: an end-to-end path for ESA application delivery" argued that DNS, edge routing, and application delivery should stop being four separate products stitched together and become one engineered path with one policy surface.

ESA Logs + AI Agents. Adjacent to the above and already described as published in an international ESA blog (though not one of the seven release-notes items here): a natural-language operations layer for asking about logs, configuration drift, and incident summaries - so "which sites saw rising error rates in the last hour" becomes a question, not a SQL editor task, and headcount doesn't scale linearly with site count.


What the lineup adds up to

Read across the three pillars, the direction is consistent: AI Gateway pulls the model control plane from the application region back to the edge; POST caching makes repeated AI calls cheap where they're idempotent; AI Crawler Management gives content sites an enforceable stance toward a new class of automated visitor; PQC is on by default and closes the record-now window on the first mile; ATO and SNI allowlist harden the account and TLS layers that AI abuse targets; Keyless unlocks the workloads key-custody rules once barred from the edge; and the runtime directions (Edge Containers, last mile, domain-as-entry, NL ops) keep pulling compute, network, DNS, and operations onto the same edge substrate.

Two non-product signals on the same stage make this read as an industry move rather than a feature drop:

A joint standard, not a vendor claim. ESA and CAICT (China Academy of Information and Communications Technology) co-published the first edge-cloud AI-native capability industry standard (session by Enran,Dong, Deputy Director, Gov & Enterprise Digital Intelligence Dept.). For buyers, that converts "AI-ready edge" from a slide adjective into a checklist you can ask any vendor to be evaluated against.

Real production customers on the same stage. Alibaba's own QwenWork (Chief Architect, Zhenyu,Pan), Adidas (Senior Staff Engineer, Xuemei,Zhao, on one-stop edge acceleration and security), and ecosystem partner Shanghai Qianyao Information Technology (CTO Zhengyi,Yang) each presented. A national-scale AI productivity app, a global consumer brand, and a partner - running combinations of the capabilities above - is the strongest available evidence that this portfolio is in production, not on a roadmap.

Getting Started

0 0 0
Share on

Bryan, Zhang

8 posts | 1 followers

You may also like

Comments

Bryan, Zhang

8 posts | 1 followers

Related Products

  • Edge Security Acceleration (Original DCDN)

    Edge Security Acceleration (ESA) provides capabilities for edge acceleration, edge security, and edge computing. ESA adopts an easy-to-use interactive design and accelerates and protects websites, applications, and APIs to improve the performance and experience of access to web applications.

    Learn More
  • Edge Node Service

    An all-in-one service that provides elastic, stable, and widely distributed computing, network, and storage resources to help you deploy businesses on the edge nodes of Internet Service Providers (ISPs).

    Learn More
  • Edge Network Acceleration

    Establish high-speed dedicated networks for enterprises quickly

    Learn More
  • Secure Access Service Edge

    An office security management platform that integrates zero trust network access, office data protection, and terminal management.

    Learn More