The edge was built around a simple assumption: content is mostly static, requests are short, sessions don't hold state, and a hit ratio is the number that matters. Everything in the classic CDN - caching, TTLs, purge rules, prefetch, origin shield - was engineered against that shape of traffic.
AI traffic breaks every one of those assumptions. An inference session streams for tens of seconds. An agent loop fires dozens of small calls per turn, each metered in tokens. Requests carry prompts and completions that should not land verbatim in an access log. The unit of cost is tokens, not bytes, so "cache hit" stops being the right word. A new class of visitor shows up who is neither a buyer nor a hacker - a crawler assembling a training corpus. And workloads that once quietly tolerated handing a private key to the edge now have to fend off store-now-decrypt-later quantum attacks and stricter key-custody rules at the same time.
The answer ESA gave at Apsara 2026 (September 22, Hangzhou) was not incremental. The keynote, "ESA's full evolution: a global runtime foundation for AI Agents," framed the shift as moving from the CDN you know toward a foundation that runs AI applications at the edge across four dimensions - acceleration, security, compute, and network. The 13:50 new-product session then cut the shipping work into three pillars: AI acceleration, AI security, and the Agent runtime. That three-pillar split is the backbone of this article.

| Pillar | Capability | Status (2026) | One-line role |
|---|---|---|---|
| AI acceleration | POST-request caching | Product updates - March | Cache reusable AI POST responses (embeddings, idempotent inference) at the edge |
| AI acceleration | Global network · last mile | Stage direction | Extend delivery into cross-border, emerging, and constrained-network regions |
| AI security | AI Gateway (transport & perf slice) | Product updates - April | Edge-proximate entry, connection reuse, streaming-friendly handling |
| AI security | AI Crawler Management | Product updates - April | Identify and govern AI crawlers as a distinct traffic class |
| AI security | Post-Quantum Encryption (PQC) | Product updates - April | Hybrid X25519MLKEM768 key exchange, on by default, site-wide |
| AI security | Keyless Certificates | Product updates - April | HTTPS acceleration without uploading your private key |
| AI security | Account Takeover (ATO) protection | Product updates - March | AI/ML detection of credential abuse and account takeover |
| AI security | SNI allowlist | Product updates - July | Validate SNI against the Host header at TLS handshake |
| Agent runtime | AI Gateway (control-plane slice) | Product updates - April | Unified proxy, observability, guardrails, failover at the edge |
| Agent runtime | Edge Containers | Stage direction | Run containerized agent tools next to the POP that ends the user session |
| Agent runtime | Domain as the entry | Stage direction | DNS → edge routing → delivery engineered as one path, one policy |
| Agent runtime | ESA Logs + AI Agents | Published capability (blog) | Natural-language ops over logs, config, and incidents |
Seven of these are already "supported" entries in ESA's product updates. The last mile, Edge Containers, and Domain-as-the-entry were shown on the Apsara stage and are not yet release-notes entries - treat them as directions, not shipping features. ESA Logs + AI Agents sits between: described as a published capability in an international ESA blog, but not among the seven 2026 release-notes items above.
For years, CDN caching effectively meant GET. POST was treated as non-cacheable by default, which was correct for form submits but wrong for a growing share of AI traffic: embedding calls, semantic-search hits, idempotent inference, model-capability probes, and repeated tool calls inside an agent loop all arrive as POST yet return responses that are cacheable in practice. Applications re-paid full round-trip and provider cost for the same result every time.
Listed as supported in ESA's product updates for March 2026. The release-notes entry reads: "Once POST caching is enabled, nodes can cache POST response bodies and serve them to subsequent identical requests, lowering origin load and improving response speed."
Embedding and similarity-search endpoints get the single biggest win - the same query repeated across users and agents stops re-crossing the ocean on every hit. Agent loops that fire the same tool or classifier call many times per turn collapse most of those calls into a cache replay. And because it's an edge rule, it applies uniformly to every AI endpoint you front, not one hand-patched service.
Cache POST only where responses are genuinely idempotent for a given request body, and key the cache on the fields that actually change the answer (model name, prompt/version, retrieval index). Set short TTLs and keep a purge path for model or index updates. Treat anything carrying per-user or personalized output as out of scope, and make sure cached bodies never leak identifiers. This rule pairs directly with the AI Gateway below: route through the gateway, let it decide cacheable vs. pass-through.
The AI Gateway is a single release-notes item that spans pillars; its edge-proximity and streaming handling are an acceleration story, and its control plane is an Agent-runtime story (2.5). Both are covered so nothing is counted twice - here is the fast-path half.
Today, an AI application that talks to more than one model provider usually runs its own gateway - a slim pass-through service (essentially a small proxy) sitting in front of OpenAI, Anthropic, Google, an open-weights backend, and a couple of embedding or safety models, doing little more than rotating API keys, throttling, retrying, and logging. That gateway lives in one region and reinvents the rate limiting, retry semantics, and observability the CDN already had years ago. The result is a second, less-matured traffic plane running parallel to the site.

Listed as supported in ESA's product updates for April 2026, described in the release-notes entry as: "The performance half of that sentence is the acceleration story: AI requests now enter at the edge node nearest the user, and the transport is engineered for streaming."
The agent loop stops paying a cross-region tax on every hop. A chat turn calls one model; an agent turn may call retrieval, tools, classifiers, and a synthesizer - a dozen small calls. A centralized gateway routes every one back to a single region; putting the entry point at the user's edge node removes that multiplier.
Security and observability are not built twice. Model APIs come with their own auth and limits, but teams cheaply skip a layer at the app tier assuming "the gateway handles it." When the gateway is the edge, and the edge already runs WAF and DDoS, the defense exists in one place.
Point clients and agent orchestrators at the gateway endpoint only; keep provider secrets in edge configuration, never on the client. Enable request/response logging on the AI path with retention tied to data sensitivity - prompts and completions often carry personal or proprietary content, so scope log access apart from normal site logs. Restrict upstream providers or self-hosted inference endpoints so only ESA's converged back-to-origin addresses can reach them (see Origin Protection in this series, Part 4).
Between "legitimate SEO crawling" and "malicious scraping" sits a new class of visitor - the AI crawler. GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBbot and their peers each have a user-agent, a politeness policy, and an intent, usually to assemble or refresh a training corpus. A content site now has to answer each family separately: is this visitor bringing me customers, or just taking my content? Should the referral path an AI answer engine sends me be treated differently?
Classic bot management had two buckets - allow known bots, challenge or block unknown ones - and that classification is no longer enough.
Listed as supported in ESA's product updates for April 2026. The entry reads: "A dedicated detection engine plus flexible access-control policies precisely identify mainstream AI crawlers, manage their permissions differentially, and analyze access data for IP protection and resource optimization."
"Should I care about AI traffic?" becomes a policy decision instead of a vibe. Every site can now state a position - fully open, conditional (throttle plus a licensing threshold for heavy users), default-deny with named exceptions, or metered (allow first, license from the access data later). The engine and the policies are the primitives that make those positions enforceable.
Reachability and IP protection stop being a false trade-off. Blocking all crawlers costs you visibility in AI answer engines. Allowing everything hands your corpus to training with no licensing trail. You can split by category - welcome the ones that return customers, govern the ones that only take.
Bulk scraping stops hitting your origin. A bot 404ing hundreds of paths a second isn't "a little noisy," it's what takes down autoscaling at 3 a.m. Decide at the edge and the cost never reaches the origin.
Content sites decide the stance first, then encode the policy; SaaS and API-first products should treat crawler management as an endpoint automation-access policy, not just a page-level one. Combine with SNI allowlist (2.5, July 2026) so crawlers can't dodge classification through a TLS-layer mismatch. Keep robots.txt consistent with the enforced policy - the detection engine is more authoritative, but a semantic conflict weakens the defensibility of any good-faith claim.
The quantum threat to today's TLS traffic is not "broken suddenly tomorrow." It is record-now, decrypt-later: capture ciphertext today, store it, and unravel it when a strong enough quantum machine exists. Anything with a confidentiality horizon beyond roughly a decade is already in range - medical records, financial instructions, government filings, legal discovery, RAG source corpora, model weights, and long-lived sessions.
The defense is swapping the TLS handshake to a quantum-resistant key exchange. What historically stopped adoption was never the algorithm - it was the cost: upgrade server libraries, gain client support, re-sign certificates, roll out gradually. ESA pushes that cost near zero because the handshake happens at the edge, and changing the edge changes every site at once.
Listed as supported in ESA's product updates for April 2026, and on by default. The entry reads: "Native post-quantum cryptography (PQC): the X25519MLKEM768 hybrid key exchange protects client-to-edge traffic against quantum attacks. Enabled by default site-wide, zero configuration. "

Closes the record-now window on the segment most likely to be recorded. The user-to-edge leg crosses the public internet, hostile networks, and physical links - the natural surface for passive capture. Making the first mile quantum-resistant is where PQC has its highest density of benefit.
"PQC support" on procurement questionnaires is a yes in advance. More enterprise and regulatory questionnaires ask for this line. On ESA it's simply on; no scheduling, no rollout campaign.
No three-way change to client, certificate, and origin code. On the edge-terminated segment the painful parts of a PQC migration dissolve: no re-issuing certificates, no forced client upgrades, no origin library swap. Old clients still complete the handshake, just on a classic suite.
For any site with a long confidentiality horizon - finance, health, public sector, model weights, RAG corpora - keep the negotiated cipher name (X25519MLKEM768) from the access log as the audit artifact that "we already run PQC." For true end-to-end coverage, the edge ↔ origin leg is a separate transport that needs origin-side PQC to extend - the default site-wide setting covers the first mile only. Deep-dive and enablement steps: "What is Post-Quantum Encryption - and How to Enable It on ESA".
A whole class of business was never standing at the CDN door - not because the edge wasn't fast, but because a compliance rule said "the private key does not leave this boundary." Finance, public sector, healthcare, defense, and the strictest tier of enterprise procurement all require the TLS private key to stay inside a self-managed HSM/KMS. The traditional onboarding flow asks you to upload both certificate and key to the platform; that single step disqualifies the whole service.
Keyless separates "the certificate is at the edge" from "the private key is at the edge." ESA uses your certificate for HTTPS acceleration at the edge; whenever a handshake needs a signature, it calls a KeyServer you deploy yourself. The private key never leaves your trust boundary - it does not enter Alibaba Cloud's platform either.
Listed as supported in ESA's product updates for April 2026. The entry reads: "With a self-hosted KeyServer and keyless certificate configured in the console; HTTPS acceleration works private key to a third-party platform — minimizing key leakage and compliance risks"
Unlocks the tier of workloads that previously could not go to the edge at all. That's the direct commercial effect: finance/public-sector/healthcare customers once excluded by key-custody rules can now take edge acceleration, WAF, DDoS, and PQC at the same time.
A compromised platform never yields your key. The key is not on a third party, so a third-party breach is not your incident. The audit story is simpler and the incident playbook is shorter.
No rip-and-replace of your CA/PKI lifecycle. Heavily regulated enterprises already have designated CAs and long-horizon certificate governance. Keyless uses the certificate you have today; you don't stand up a shadow PKI on a cloud platform.
A hardened host on a low-exposure subnet inside your network is enough for the KeyServer - it answers signature challenges, it doesn't carry traffic, so its throughput and network surface are small. Register it as a Keyless configuration in the console, then treat it like an origin: allow only ESA's converged back-to-origin address ranges to reach the KeyServer (whitelist procedure in Part 4, Origin Protection). Before rollout, take an end-to-end handshake-latency baseline - the signing hop adds one internal round-trip to TLS setup. Note Keyless and PQC are orthogonal and stack cleanly: PQC guards the transport, Keyless guards key custody.
AI applications multiply the attack surface on accounts. Agent sessions and API keys are long-lived, programmatic, and easy to reuse; credential stuffing, brute force, phishing, and infostealer malware all target them. A single leaked key can quietly drive a large token bill or impersonate a user through an agent loop. Traditional WAF rules see valid requests and let them through; the abuse is behavioral, across sessions and devices, not in any one request.
Listed as supported in ESA's product updates for March 2026. The entry reads: "Account Security uses AI and machine-learning to detect account-takeover attacks, defending against credential stuffing, brute force, phishing, and infostealer compromise."
Abuse detection lands where the traffic already terminates. Because ATO scoring runs at the edge, you don't have to ship login telemetry to a separate security vendor to get a risk verdict. AI apps with user accounts and API keys get a defense built for exactly their failure mode - a stolen key replayed from a new geography is an ATO signal, not a WAF signature. And the scoring is per-request but stateful, so it complements (does not replace) your rate limits.
Turn ATO protection on for authentication, password reset, token-exchange, and high-value API endpoints. Feed it real client IP (behind the edge, preserve it properly) so geographic and device signals are accurate. Pair with rate limits and device/bot signals already available at the edge; treat ATO verdicts as one input into step-up auth or key revocation. For agent platforms, scope per-key spend alarms so a hijacked key is caught by cost anomaly as well as behavioral score.
Automated clients - crawlers included - increasingly exploit mismatches at the TLS layer: a handshake announces one SNI but the HTTP request carries a different Host header, or the client simply probes with an SNI that doesn't belong to your site. That mismatch lets abusive traffic slip past host-based routing and classification.
Listed as supported in ESA's product updates for July 2026. The entry reads: "A new SNI allowlist validates the SNI used in the SSL/TLS handshake against the Host header in the HTTP request, giving sites a wider range of protection policies."
SNI validation closes the TLS-layer evasion route that AI-crawler and bot policies depend on staying honest - an abusive agent that mislabels its SNI can no longer present itself as traffic for a permitted hostname. It's a small, cheap primitive, but it makes the enforcement of everything above it (crawler policy, bot rules, routing) materially harder to bypass.
Turn on SNI-to-Host consistency for domains that should only ever be reached with a matching SNI; use it under AI Crawler Management (2.2) so crawler classification can't be dodged at the handshake. Because it's an allowlist, enumerate the legitimate SNI/Host pairs for multi-tenant or apex/CNAME setups before enforcing, to avoid false rejects on correctly-configured clients.
Same April 2026 release-notes item as 2.1; this is the "unified API proxy + observability + guardrails" half.
Once the AI path is terminated at the edge, the natural next step is to run the model control plane there too - a single place that fronts every provider, sees every call, and applies policy consistently. That removes the parallel application-side gateway and the drift between "what the site enforces" and "what the AI path enforces."
The April 2026 AI Gateway entry is the same text as above - "one edge proxy for AI calls, with request/response observability, inherited edge security, and provider-agnostic policy."
Switching models is a config change, not a code migration. Moving from vendor A to B for price, quality, or compliance no longer touches the SDK or rewrites retry logic - you change one edge policy. Token-level economics get collected nearest the user. Edge instrumentation lets product and finance see cost by region and cohort for the first time, not by datacenter.
Covered with 2.1: single gateway endpoint, provider keys at the edge, AI-path logging scoped by sensitivity, upstream locked to ESA's converged origin ranges.
The product session and the customer talks highlighted three runtime directions that are on the roadmap and demoed on stage(but not yet officially supported).
Edge Containers. Shown by Shanghai Qianyao Information Technology founder & CTO Zhenyi,Yang, in "Building global applications on ESA AI Gateway + Edge Containers." The idea is containerized workloads - agent tools, retrieval functions, small-model inference, glue - scheduled onto the same edge substrate that already carries delivery, security, and the AI Gateway. Its value is completing what the gateway started: if the agent's tools and retrieval still live in one region, the "per-hop cross-region tax" is only partly removed; pulling compute next to the session that ends the user removes the multiplier fully.
Global network · last mile. Product Director, Jimmy Wang's session "ESA global network: extending cloud and AI applications to the last mile" pushed the transport layer into places the public internet handles poorly - cross-border, emerging markets, constrained networks. This is the same engineering posture as Part 5 of this series (mainland-China access optimization), applied to other geographies. The direction is to turn "globally reachable" from an adjective into something you can put behind an SLA.
Domain as the entry. Xiaojie, Sun's session "Domain as the entry: an end-to-end path for ESA application delivery" argued that DNS, edge routing, and application delivery should stop being four separate products stitched together and become one engineered path with one policy surface.
ESA Logs + AI Agents. Adjacent to the above and already described as published in an international ESA blog (though not one of the seven release-notes items here): a natural-language operations layer for asking about logs, configuration drift, and incident summaries - so "which sites saw rising error rates in the last hour" becomes a question, not a SQL editor task, and headcount doesn't scale linearly with site count.
Read across the three pillars, the direction is consistent: AI Gateway pulls the model control plane from the application region back to the edge; POST caching makes repeated AI calls cheap where they're idempotent; AI Crawler Management gives content sites an enforceable stance toward a new class of automated visitor; PQC is on by default and closes the record-now window on the first mile; ATO and SNI allowlist harden the account and TLS layers that AI abuse targets; Keyless unlocks the workloads key-custody rules once barred from the edge; and the runtime directions (Edge Containers, last mile, domain-as-entry, NL ops) keep pulling compute, network, DNS, and operations onto the same edge substrate.
Two non-product signals on the same stage make this read as an industry move rather than a feature drop:
A joint standard, not a vendor claim. ESA and CAICT (China Academy of Information and Communications Technology) co-published the first edge-cloud AI-native capability industry standard (session by Enran,Dong, Deputy Director, Gov & Enterprise Digital Intelligence Dept.). For buyers, that converts "AI-ready edge" from a slide adjective into a checklist you can ask any vendor to be evaluated against.
Real production customers on the same stage. Alibaba's own QwenWork (Chief Architect, Zhenyu,Pan), Adidas (Senior Staff Engineer, Xuemei,Zhao, on one-stop edge acceleration and security), and ecosystem partner Shanghai Qianyao Information Technology (CTO Zhengyi,Yang) each presented. A national-scale AI productivity app, a global consumer brand, and a partner - running combinations of the capabilities above - is the strongest available evidence that this portfolio is in production, not on a roadmap.
8 posts | 1 followers
FollowJohnny Mai - August 27, 2026
Alibaba Cloud Community - September 7, 2026
Alibaba Cloud Community - July 17, 2026
Alibaba Cloud Community - September 25, 2026
Bryan, Zhang - October 5, 2026
Alibaba Cloud Native Community - April 15, 2026
8 posts | 1 followers
Follow
Edge Security Acceleration (Original DCDN)
Edge Security Acceleration (ESA) provides capabilities for edge acceleration, edge security, and edge computing. ESA adopts an easy-to-use interactive design and accelerates and protects websites, applications, and APIs to improve the performance and experience of access to web applications.
Learn More
Edge Node Service
An all-in-one service that provides elastic, stable, and widely distributed computing, network, and storage resources to help you deploy businesses on the edge nodes of Internet Service Providers (ISPs).
Learn More
Edge Network Acceleration
Establish high-speed dedicated networks for enterprises quickly
Learn More
Secure Access Service Edge
An office security management platform that integrates zero trust network access, office data protection, and terminal management.
Learn MoreMore Posts by Bryan, Zhang