All Products
Search
Document Center

Platform For AI:Release notes

Last Updated:Jul 06, 2026

2026

June

Release date

Feature

Description

Related documentation

June 30, 2026

DSW AgentBox released with built-in mainstream AI Coding Agents, ready to use out of the box

Use AI Agent to assist with code writing and debugging, rapid prototyping during training task intervals, and AI-assisted full-stack development on GPU instances. For team managers, unified PAI Token configuration eliminates the need for individual API Key applications, lowering the barrier for teams to adopt AI programming tools.

June 30, 2026

PAI Agentic Readiness V1.0: Core CLI and Core Skills

Users can operate PAI in their preferred Agent without learning or using the PAI console.

May

Release date

Feature

Description

Related documentation

May 30, 2026

Deep integration of commercial large models from Model Studio

PAI integrates commercial large models from Model Studio, expanding the selection of ready-to-use models.

April

Release date

Feature

Description

Related documentation

April 29, 2026

Intelligent Assistant (Xiao PAI) MCP service for DLC diagnostics

Intelligent Assistant (Xiao PAI) now exposes its diagnostic and Q&A capabilities as an MCP service for agent ecosystem integration. Query DLC task restart counts, analyze restart reasons, and check blocklisted node status to identify root causes of training failures.

Intelligent Assistant (Xiao PAI)

April 23, 2026

DSW DockerBoard for sub-container management on shared machines

DSW DockerBoard is now GA on Lingjun Cluster. Create and manage multiple sub-containers within a single DSW instance, enabling multiple developers to share a machine with better resource utilization.

Manage Sub-containers with DockerBoard

April 17, 2026

Dynamic parameters for EAS services

EAS services now support dynamic parameters that you can hot-update in real time and perform real-time create, read, update, and delete (CRUD) operations on without restarting the service or its instances.

April 15, 2026

DLC task templates

DLC now supports task templates for different experimental tracks. Pre-fill configurations, lock critical settings, and define standard description formats to capture submission best practices, archive tasks within the same experiment, and improve iteration efficiency.

Task Templates

April 13, 2026

DLC flexible storage mounting for Lingjun Cluster v1.0

DLC on Lingjun Cluster now supports flexible storage mounting v1.0. It auto-detects network characteristics and switches between VPC and VCS mounts, enabling cross-HPN-Zone access when computing power and CPFS storage are in different HPNs.

April 13, 2026

GPU resource utilization metrics for DLC tasks

DLC now reports GPU utilization metrics based on busy time—the percentage of time a GPU actively computes—helping you manage computing resources more precisely.

April 9, 2026

Linux graphical desktop for DSW

DSW instances now include a Linux graphical desktop (TurboVNC + noVNC) accessible through your browser without extra configuration. Ideal for scenarios requiring a GUI, such as autonomous driving simulations and embodied AI debugging.

March

Release date

Feature

Description

Related documentation

March 20, 2026

Model warm-up cache for PAI-EAS

The model warm-up cache preloads model caches and provides a high-speed source for cache-accelerated inference services. Suited for LLMs and AI image/video generation where large models are mounted from OSS or NAS.

Deploy a model warm-up cache service

March 19, 2026

Global search

Search across PAI resources—DSW instances, DLC tasks, EAS services, and models—by keyword to quickly locate assets.

February

Release date

Feature

Description

Related documentation

February 27, 2026

Intelligent configuration for PD decoupling in PAI-EAS

When deploying distributed inference with PD decoupling, Intelligent Assistant (Xiao PAI) recommends optimal PD ratios and parallelism strategies based on model characteristics, instance configurations, and I/O lengths to improve resource utilization and end-to-end performance.

February 26, 2026

DSW integration with OpenClaw

PAI-DSW automates OpenClaw installation with system control, persistent memory, and scheduled notifications. Interact through a web UI or DingTalk while the AI agent runs directly on its compute resources.

January

Release date

Feature

Description

Related documentation

January 30, 2026

AutoML available in Singapore

The PAI-AutoML service is now generally available in the Singapore region.

January 23, 2026

Resource diagnostics for inference services

EAS now provides resource quota and priority queuing diagnostics for both distributed service instances that use multiple machines and GPUs, and regular service instances, helping you identify and resolve resource issues.

January 8, 2026

PAI Intelligent Assistant diagnostics (Phase 2): Support for DLC tasks and EAS services

Phase 2 adds real-time diagnosis, root cause analysis, and repair suggestions for DLC training failures and EAS performance bottlenecks. Supports interactive, natural-language troubleshooting.

2025

December

Release date

Feature

Description

Related documentation

2025-12-22

Model evaluation supports dual-model offline competition

PAI ModelEval now supports dual-model offline competitions. Compare two models on the same Q&A dataset to select the better performer.

2025-12-19

KV Store global context cache released

EAS LLM deployments now support a global context cache for the KV Store to improve inference throughput.

2025-12-05

PAI-DLC releases Ray 2.0: RayQuota

PAI-DLC now supports Ray's native scheduling and runtime on PAI resource quotas, delivering stable and efficient performance for AI data processing and reinforcement learning in LLM and autonomous driving workloads.

2025-12-01

DLC displays all IP addresses of running pods

DLC training jobs now expose pod IP addresses across computing resource types (Lingjun and general-purpose), including head/tail node IPs, network IPs, and RDMA IPs.

View training details

November

Release date

Feature

Description

Related documentation

2025-11-28

Ray on DLC supports dynamic scaling

Ray on DLC now supports dynamic scaling. You can configure minimum and maximum instance counts and use the quota preemption mechanism for task scaling. This helps balance global tasks and resource usage.

Quickly deploy a WebUI service

2025-11-28

PAI-EAS releases scenario-based deployment for CosyVoice

PAI-EAS now supports the CosyVoice high-fidelity speech synthesis model, which is suitable for various scenarios, including customer service conversations, audiobook narration, and short video dubbing.

2025-11-21

DSW OpenAPI supports calls over PrivateLink

You can now call the DSW OpenAPI over PrivateLink, allowing you to manage instances and integrate features from your office network or on-premises data center. A connection to an Alibaba Cloud VPC must be established beforehand.

2025-11-07

PAI-EAS supports instance-level real-time logs

PAI-EAS now supports viewing instance-level real-time logs with streaming output. This feature is ideal for interactive debugging and other scenarios that require low-latency visibility.

October

Release date

Feature

Description

Related documentation

2025-10-17

Lingjun intelligent computing GU7 series supports the 570 driver

The Lingjun intelligent computing GU7 series now supports the 570 driver. For the GU7 series, the PAI training service supports multiple drivers, including 530, 550, and 570. You can select a driver when submitting a task without reinstallation, leveraging the flexibility of the Serverless platform.

2025-10-17

Scenario-based deployment of Dify on PAI-EAS is released

Dify is an open-source development platform for LLM applications that helps developers, enterprises, and non-technical personnel quickly build, deploy, and manage generative AI-based applications. You can deploy the open-source version of the Dify platform on EAS with a single click. This supports WebUI usage and API calls, allowing you to quickly build, deploy, and manage your applications.

Deploy Dify, an LLM application platform

2025-10-17

DSW instance lifecycle events support message notifications

1. Workspace administrators can configure message notification rules in Workspace Settings > Event Notification. You can select instance status and image save status events and specify notification targets, such as DingTalk, SMS, phone calls, or WeCom.

2. When a status event in a notification rule occurs, you receive a real-time message notification, allowing you to take prompt action.

2025-10-17

Model distillation feature is released

Based on PAI's proprietary EasyDistll algorithm library, this feature provides productized model distillation capabilities. It transfers the capabilities of a large teacher model to a smaller student model through distillation. The feature also supports online synthesis of distillation data, including instruction data for non-inference models and chain-of-thought data for inference models. This enables the student model to approach or match the teacher model's performance on specific tasks, helping you improve model performance and reduce deployment costs.

September

Release date

Feature

Description

Related documentation

2025-09-19

ArtLab releases the design agent

The Design Agent allows users to generate high-quality images, produce videos, and perform fine-grained image edits using natural language instructions. This feature streamlines the AIGC design workflow and makes advanced creative tasks more accessible.

Unleash the creativity of natural language and redefine the AIGC design pipeline.

2025-09-19

EAS releases compute detection and fault tolerance

The compute detection and fault tolerance feature for EAS comprehensively inspects all resources involved in inference. It automatically isolates faulty nodes and triggers background automated O&M processes, which reduces the likelihood of issues during initial inference and improves deployment success rates.

Compute detection and fault tolerance

2025-09-17

Data development supports direct execution of PAI Flow nodes

PAI Flow provides end-to-end machine learning workflow development capabilities, offering the same workflow features as PAI's visual modeling Designer and enabling periodic workflow scheduling.

PAI Flow node

August

Release date

Feature

Description

Related documentation

2025-08-25

EAS supports expert parallelism deployment

EAS now supports expert parallelism (EP) for deploying Mixture-of-Experts (MoE) models, such as DeepSeek-R1. This feature works with inference engines like vLLM and SGLang to overcome hardware limitations, improve resource utilization, and increase system throughput.

Deploy MoE models by using expert parallelism and PD separation

2025-08-15

Model evaluation center v1.0 released

This out-of-the-box feature lets you run a complete model evaluation pipeline without writing code, helping you quickly determine if a model meets your business needs.

2025-08-13

AI resource groups (Lingjun intelligent computing) support pay-as-you-go and savings plans

AI resource groups (Lingjun intelligent computing) now support pay-as-you-go purchases. When combined with savings plans, the service automatically applies discounts based on the subscription duration. The longer the subscription (1, 3, or 5 years), the greater the discount. This provides a more flexible and cost-effective way to use resources.

Create a resource group and purchase Lingjun intelligent computing resources

2025-08-13

DataJuicer on DLC is generally available

DLC now supports submitting DataJuicer framework tasks. By using numerous operators (over 100), multi-scale processing (single-node and multi-node), and high availability (self-healing), DataJuicer efficiently performs large-scale data cleaning, filtering, transformation, and enhancement for LLM text and multimodal data processing.

Quickly submit a DataJuicer job

2025-08-11

DLC's proprietary custom task framework (v1.0) is generally available

DLC's proprietary distributed framework, Custom, supports PAI scheduling policies and self-healing capabilities. It also provides advanced features such as custom roles, success policies, and extended ports. This meets the computing demands of various business scenarios, including post-training for large models and autonomous driving.

2025-08-08

Model weight service is released

The model weight service significantly reduces cold start and scale-out times. It addresses the common issue of slow model loading times, overcoming a key performance bottleneck in ultra-large-scale LLM deployment.

Model weight service

2025-08-07

EAS releases Prefill-Decode (PD) separation

EAS has released the Prefill-Decode (PD) separation feature. It includes multiple deployment modes, such as static and dynamic PD separation, and is compatible with various inference engines like vLLM, SGLang, and BladeLLM to help you reduce inference latency.

July

Release date

Feature

Description

Related documentation

2025-07-10

DSW supports distributed development and debugging environments

This feature helps you debug and verify distributed tasks, creating a more efficient development and training workflow.

  1. Built-in environment variables, such as communication library configurations, adapted for different resources and network architectures.

  2. A DNS-based method for instances to access each other by using their instance IDs.

  3. Support for RDMA and eRDMA.

Use interconnected instances for distributed training

June

Release date

Feature

Description

Related documentation

2025-06-10

ArtLab supports building and sharing AIGC applications based on ComfyUI

PAI-ArtLab enhances its enterprise-grade AIGC application capabilities.

1. Build and publish custom AIGC applications based on ComfyUI workflows.

2. Share AIGC applications from the ArtLab platform as out-of-the-box AIGC applications for PCs and H5 mobile clients.

Model Gallery

2025-06-05

Data development supports PAI Flow

This feature unifies the entry points for big data development and AI products, and enhances the deep integration between PAI Flow and big data engines to achieve integrated big data and AI development.

  1. Data Studio PAI Flow supports more operator components.

  2. Supports end-to-end operations for PAI Flow, such as running, publishing, and O&M.

  3. The user experience of PAI Flow is optimized, such as by enhancing the configuration linkage of required properties for PAI Flow operator components.

  4. You can use PAI Flow as a node within a larger workflow and run the entire workflow.

May

Release date

Feature

Description

Related documentation

2025-05-30

Quick start > Model Gallery is available in the Malaysia (Kuala Lumpur) region

Model Gallery is now available in the Malaysia (Kuala Lumpur) region. Model Gallery integrates high-quality pretrained models from various open-source AI communities to help you quickly get started with model training and deployment on PAI.

Model Gallery

April

Release date

Feature

Description

Related documentation

2025-04-04

Image management supports building custom images

PAI now offers a custom image building feature. You can flexibly install dependencies on an existing image or build a custom image by using a Dockerfile. The image is automatically pushed to ACR and registered on the PAI platform, which meets custom requirements and eliminates the tedious process of local building and uploading.

2025-04-01

DSW is generally available in US (Virginia) and US (Silicon Valley)

DSW is now generally available in the US (Virginia) and US (Silicon Valley) regions.

March

Release date

Feature

Description

Related documentation

2025-03-28

LangStudio 1.0 is released

This version adds the following features based on LangStudio 0.1:

1. Knowledge base creation and management: You can create and synchronize knowledge bases on the console and use them when building application flows.

2. Use PAI-DSW as an application development environment: Provides DSW-based Notebook and WebIDE environments for application flow development.

3. Application flow performance evaluation: Prebuilt evaluation templates support offline performance evaluation and online service performance evaluation for applications.

4. Conversation history for deployed applications: You can use a cloud database or local storage for conversation history, and manage and export the history.

2025-03-28

DSW supports dynamic mounting of NAS datasets and storage paths

DSW now supports dynamic storage mounting. You can mount or unmount NAS datasets without restarting the instance.

2025-03-28

DLC supports mounting OSS data sources by using ossfs

DLC now supports mounting OSS data sources by using ossfs. This provides good OSS read and write performance for compute-intensive tasks such as autonomous driving, which typically involve sequential and random reads, and sequential append writes.

Use cloud storage

2025-03-27

AI compute node statuses upgraded

The compute node statuses are optimized. A new status code that indicates scheduling is prohibited is added to improve your user experience.

Nodes

2025-03-19

DLC supports custom roles for Ray tasks

When you submit Ray framework tasks in DLC, you can customize the Worker role to enable hybrid execution on heterogeneous resources.

Create a training job

2025-03-19

Resource quotas support scaling of specified nodes

Resource quota scaling now supports node-level operations. This provides greater flexibility for managing, reallocating, and transferring computing power between quotas.

Manage resource quotas

2025-03-07

PAI training service is generally available in US (Silicon Valley)

DLC and resource quotas are now available in the US (Silicon Valley) region. You can use resource quotas and public resources (pay-as-you-go) to submit training jobs.

Regions and zones

February

Release date

Feature

Description

Related documentation

2025-02-28

DLC supports read/write permissions when mounting storage instances

When you mount Alibaba Cloud storage instances (such as OSS, NAS, and CPFS) in DLC, you can configure read/write permissions. This supports fine-grained permission management for your storage instances.

Use cloud storage

2025-02-21

AI scheduling engine v2.0 implements multi-level task preemption

Based on resource quotas, the PAI AI scheduling engine uses task classification and a dynamic priority algorithm to trigger preemption, ensuring high-priority jobs can execute quickly. Combined with AIMaster's preemptive rollback technology, interrupted jobs automatically save their intermediate state and enter a queue. They are then prioritized for resumption after resources are released. This achieves efficient scheduling in resource-constrained scenarios.

Preemption strategy

2025-02-10

PAI-DLC supports nconnect when mounting Alibaba Cloud file storage

When mounting Alibaba Cloud file storage (such as NAS and CPFS) for PAI-DLC jobs, you can now configure multiple connections (nconnect). This provides fine-grained control over the number of mount connections, which optimizes concurrent access performance from multiple nodes and ensures the stability of large-scale training jobs.

Use cloud storage

2025-02-07

EAS supports multi-node distributed inference

With the advent of ultra-large-scale Mixture-of-Experts (MoE) models like Qwen-Max and DeepSeek, a single device can no longer handle their vast number of parameters. To address this, EAS introduces a multi-node distributed inference solution that overcomes hardware limitations and efficiently supports the deployment and operation of ultra-large-scale models. EAS distributed inference supports various parallelism methods, including pipeline parallelism, tensor parallelism, and data parallelism, and is compatible with high-performance inference engine frameworks such as BladeLLM, vLLM, and SGLang.

Multi-node distributed inference

January

Release date

Feature

Description

Related documentation

2025-01-21

DLC supports training timeout alerts

DLC now supports the configuration of training timeout alerts. You can customize timeout alert rules for training jobs during the environment preparation, queuing, and running stages. When a rule is triggered, an alert notification is sent, making it easier for you to monitor abnormal training progress.

Message notification

2025-01-21

DLC supports training status notifications

DLC now supports subscriptions to training status notifications. New status events such as queuing, preempting, preparing, and running are added. This makes it easier for you to track training progress and enhances the message notification capabilities of the training service.

Message notification

2025-01-20

DLC supports direct mounting of storage services during job submission

When you submit a training job by using DLC, you can directly select different storage instances in the submission form. A variety of Alibaba Cloud storage instances are supported, including OSS, General-Purpose NAS, Extreme NAS, General-Purpose CPFS, and AI-Computing CPFS. This lowers the entry barrier and simplifies usage.

Create a training job

2025-01-20

Ray on DLC supports idle resources

DLC now supports submitting Ray jobs by using idle resources. This allows you to run multiple types of jobs with a single set of resources, enabling resource sharing between jobs and improving resource utilization.

Use idle resources

2025-01-20

ArtLab launches industry tools

ArtLab has launched industry tools. The first phase includes applications such as realistic e-commerce product rendering (for home appliances and furniture), corporate-style poster generation, and creative footwear design. More prebuilt applications will be continuously added.

2025-01-20

ArtLab launches the AIGC Application Zone

PAI-ArtLab has launched the AIGC Application Zone module. This allows you to use online applications that encapsulate ComfyUI workflows for operations like text-to-image and image-to-image generation. It lowers the barrier to using AIGC production tools and reduces costs through a Serverless service model.

1. Out-of-the-box: Start applications with one click, with no environment configuration required.

2. Platform-integrated enterprise-grade AIGC applications, such as generating corporate-style posters and creating profile pictures for corporate events.

3. Serverless application mode: Billing occurs only during GPU inference, significantly reducing user costs.

2025-01-20

Model Gallery supports model inference acceleration

Pretrained models in PAI-Model Gallery can be matched with supported inference acceleration capabilities (such as vLLM and BladeLLM) based on the selected machine type.

2025-01-16

EAS upgrades BladeLLM high-performance deployment service

PAI-EAS now supports scenario-based deployment of BladeLLM, achieving faster response times and higher throughput for LLM inference.

BladeLLM is a PAI-proprietary inference engine that provides an efficient runtime, high-performance operator implementation, and hybrid quantization. PAI-EAS fully integrates with BladeLLM to launch a high-performance LLM inference service. It supports the deployment of prebuilt and custom models and lets you enable advanced options like model parallelism and speculative sampling with a single click, providing an efficient LLM deployment solution.

Get started with BladeLLM

2025-01-02

Model Gallery is generally available in more regions

PAI-Model Gallery integrates pretrained models from fields such as LLM, computer vision (CV), natural language processing (NLP), and speech, providing one-stop, no-code features for model training, compression, evaluation, and deployment. The service is now available in the following regions: China (Hong Kong), Japan (Tokyo), Indonesia (Jakarta), Germany (Frankfurt), and US (Virginia).

2024

December

Release date

Feature

Description

References

2024-12-23

Billing for DLC pay-as-you-go jobs now supports filtering by job type

DLC training jobs now support a system tag (key:acs:pai:payType) to distinguish between pay-as-you-go jobs and preemptible jobs. This tag lets you quickly identify and filter job types in your billing system, providing a clearer view of your consumption and discounts.

View billing details

2024-12-16

Machine Learning Designer now supports aggregated execution for LLM data preprocessing pipelines

Machine Learning Designer now supports merging and executing multiple serial large model data preprocessing (on DLC) nodes. This avoids time-consuming disk writes and the overhead from repeatedly starting and stopping distributed tasks, which improves execution efficiency. The feature also supports automatic aggregation.

Group LLM data processing components

2024-12-16

DLC now supports preemptible jobs on general-purpose computing resources

DLC now supports running preemptible jobs on general-purpose computing resources. This offers a more cost-effective option for AI computing power.

2024-12-10

PAI training service is now available in the Germany (Frankfurt) region

DLC and resource quotas are now available in the Germany (Frankfurt) region. You can use resource quotas to submit training jobs in this region.

2024-12-09

DLC job status lifecycle updated to v2.0

To simplify the job lifecycle for jobs that use a resource quota, the "Queuing" and "PreAllocation" statuses are now consolidated into a single "Queuing" status. This makes job statuses clearer and easier to track.

2024-12-06

DLC SanityCheck now supports custom check items

The DLC SanityCheck feature now supports over 15 check items, including compute performance checks, node communication checks, cross-checks for compute and communication, and simulation verification. This improves troubleshooting and fault diagnosis for issues related to computing power and networks. You can also select specific checks to run, giving you greater control.

SanityCheck: Health check

November

Release date

Feature

Description

References

November 20, 2024

DSW supports dynamic mounting of OSS datasets

1. Dynamically mount and unmount OSS datasets without restarting your instance for rapid data access.

2. A user-friendly SDK lets you mount or unmount a dataset with simple configurations or a single line of code.

3. Mount AI asset datasets, such as PAI public or custom datasets, or directly mount an OSS storage path.

Mounting datasets, OSS, NAS, or CPFS

November 20, 2024

DSW instances support custom service access configuration

To support the growing use of WebUI and application frameworks in AIGC development, PAI-DSW offers a custom service access configuration. Use this feature to securely share services with collaborators for testing and verification at any point during development.

Access instance services over the internet

October

Release date

Feature

Description

Related documentation

2024-10-17

AI general computing resource groups in international regions support L20 instances

PAI AI general computing resource groups in international regions support L20 (gn8is series) instances.

2024-10-12

DLC job statuses upgraded to v1.0

DLC job statuses have been updated for improved clarity and simplicity. This update applies to both jobs and instances across various resource types (such as quota and bidding resource) and billing models (like subscription and pay-as-you-go). A new EnvPreparing status is added, a Bidding status is introduced for bidding jobs, and the Created status is simplified. The Queuing and PreAllocation statuses for pay-as-you-go jobs are also simplified.

2024-10-11

ArtLab adds serverless ComfyUI tool

The PAI-ArtLab toolbox includes a serverless ComfyUI tool, allowing you to perform operations such as text-to-image and image-to-image. The serverless model reduces costs by billing you only for model inference.

AI Design (ArtLab)

2024-10-10

QuickStart supports DPO and CPT for LLM training

PAI QuickStart's Model Gallery offers enhanced LLM training capabilities. In addition to the existing Supervised Fine-Tuning (SFT), this update introduces Direct Preference Optimization (DPO) and Continued Pre-training (CPT) methods.

QuickStart

September

Release date

Feature

Description

Related documentation

September 29, 2024

DSW now integrates Tongyi Lingma

DSW now integrates the AI coding assistant Tongyi Lingma (Personal Edition), providing features such as real-time, line-level and function-level code completion, natural language to code, unit test generation, code optimization, comment generation, code explanation, developer Q&A, and error troubleshooting. You can use these features immediately—no installation or login required—to code more efficiently.

Code Intelligently with Tongyi Lingma

October 8, 2024

PAI Training Service is now available in China (Hong Kong) and Indonesia (Jakarta)

DLC and AI resource quota are now available in the China (Hong Kong) and Indonesia (Jakarta) regions. You can submit training jobs in these regions using resource quotas or pay-as-you-go public resources.

August

Date

Feature

Description

References

November 11, 2024

The judge model service is now generally available (GA)

The PAI judge model service uses an LLM fine-tuned on Qwen2 as a judge to score the outputs of evaluated models. It is designed for open-ended and complex Q&A scenarios. Its key benefits are as follows:

1. Accuracy: The judge model excels at evaluating subjective questions. It intelligently classifies questions into various scenarios, such as open-ended chats, consultations, recommendations, creative writing, code generation, and role-playing. It then applies scenario-specific evaluation criteria, significantly improving evaluation accuracy.

2. Efficiency: The judge model eliminates the need for manual data labeling. Simply provide the prompts and model responses, and the service will automatically analyze and evaluate the LLM, significantly increasing evaluation efficiency.

3. Ease of use: You can create evaluation tasks through multiple methods, including the console, API calls, and SDK calls. This supports both quick-start experiences and flexible integration for developers.

4. Cost-effectiveness: At a lower cost, the service provides evaluation performance comparable to that of ChatGPT-4 in Chinese-language evaluation scenarios.

Judge model overview

September 3, 2024

DSW launches its lightweight edition (Notebook Lab)

1. Lightweight notebook editing: You can develop directly in your browser without pre-provisioning resources.

2. Notebooks as assets: Notebooks are decoupled from instance resources, making it easier to archive and share them as technical documents or code.

Notebook Lab

August 26, 2024

EAS launches LLM Intelligent Router to improve LLM inference efficiency

When you deploy an LLM-based EAS service, you can associate it with an LLM Intelligent Router. The router intelligently distributes requests to ensure that the compute and GPU memory usage across backend instances remain balanced, improving overall cluster resource utilization.

Deploy an LLM Intelligent Router

August 26, 2024

DLC training jobs on general-purpose compute resources now support CPU affinity

For workloads whose performance is significantly affected by CPU cache affinity and scheduling, PAI DLC general-purpose compute resource groups now support CPU core binding. This feature improves job performance.

August 15, 2024

EAS launches the dedicated gateway feature

The EAS dedicated gateway is designed to meet inference requirements for security isolation and access control and to reduce network risk in high-concurrency and high-throughput scenarios. You can configure access allowlists for both internet and intranet traffic, enabling fine-grained management. Dedicated gateway resources ensure stable service connections. You can use PrivateLink to connect the intranet to your enterprise VPC. You can also enable or disable internet access at any time, giving you full control over your network.

Use a dedicated gateway

August 15, 2024

Workspaces now support custom user roles

A workspace is the top-level concept in PAI. It provides enterprises and teams with unified management of computing resources and user permissions, along with end-to-end development tools that support team collaboration and AI asset management. As usage scenarios become more advanced, built-in roles may not meet the specific management needs of customers. For example, you might want a role that can use DSW but not DLC. To address such needs, PAI workspaces now support custom roles, allowing you to configure permissions with greater flexibility.

Manage members of a workspace

August 5, 2024

Discontinuation of early-version PAI-PyTorch algorithm components (PyTorch100 and PyTorch131)

Dear PAI users,

Due to a system-wide upgrade, Platform for AI (PAI) will officially discontinue early-version PyTorch algorithm components in all clusters on August 30, 2024. If you submit PyTorch jobs to MaxCompute using the PAI command pai -name pytorch100/pytorch131, please migrate them promptly. We recommend that you use PAI-DLC to submit PyTorch jobs. Create a training job. Starting August 31, 2024, existing jobs that use these early-version PyTorch algorithm components will no longer be covered by the Service Level Agreement (SLA).

If you have any questions or need assistance, please contact us through your dedicated DingTalk group or by submitting a ticket.

Thank you for your cooperation.

July

Release date

Feature

Description

References

July 3, 2024

EAS now supports GPU sharing

When you deploy a model in EAS, you can partition GPU resources based on compute share and memory size. This feature reduces resource costs and improves utilization. On the deployment page, you can schedule instances by either memory or compute power, enabling multiple instances to run on a single GPU.

EAS Overview

July 3, 2024

Health checks now available for EAS service instances

The EAS health check feature helps maintain high availability by providing rapid fault detection and automatic recovery, enabling enterprise-grade inference service deployments. By using the native health check mechanism of Kubernetes, EAS automatically detects and recovers failed containers. This ensures that traffic is routed only to healthy instances and prevents resources from being allocated to unhealthy ones.

EAS Overview

June

Release date

Feature

Description

Related documentation

2024-07-01

QuickStart now supports model evaluation for LLMs

PAI-QuickStart offers a model evaluation feature for LLMs. You can assess models by using authoritative public datasets (such as CMMLU, C-Eval, and MMLU) or your own custom datasets. This helps you determine a model's suitability for your business scenarios and compare the performance of different models.

Model evaluation

2024-06-19

PAI general computing resources support CPFS for Lingjun (invitational preview)

The PAI training service running on general computing resources (ECI) can now mount CPFS for Lingjun, providing a cost-effective solution for data storage and computing in large-scale model scenarios.

2024-06-12

Machine Learning Designer is now available in the China (Ulanqab) region

You can now use Machine Learning Designer in the China (Ulanqab) region from the PAI console.

2024-06-11

Machine Learning Designer adds the Notebook component

Machine Learning Designer now includes the Notebook component, which connects to DSW instances. This allows you to write, debug, and run code directly within a pipeline, preserving its context and state.

Notebook

May

Release date

Feature

Description

Related documentation

2024-07-01

QuickStart now supports QLoRA, LoRA, and full-parameter fine-tuning for LLMs

For LLMs in PAI-QuickStart, you can now choose from full-parameter fine-tuning or the more cost-effective LoRA and QLoRA methods. This allows you to select the best approach for your needs and reduce training costs.

QuickStart: Fine-tune, evaluate, and deploy Qwen2.5 series models

2024-06-07

DSW now supports configuring a RAM role for an instance

After you grant the default PAI RAM role to an instance, you can develop and debug in DSW without configuring an AccessKey pair (AK). Authentication is handled securely through the credential chain in the following scenarios:

  • Submit a training job to the current workspace using the PAI SDK.

  • Submit a training job to the current workspace using the DLC SDK.

  • Submit a job to a MaxCompute project where the instance owner has execution permissions, using the ODPS SDK.

  • Access data in the bucket configured as the default storage path for the current workspace using the OSS SDK.

  • Use the Tongyi Lingma service in the WebIDE.

If a custom role is injected, the DSW instance uses the role's temporary access credentials to access specified Alibaba Cloud services, such as OSS and RDS. This ensures secure communication between the DSW instance and other Alibaba Cloud services.

Configure a RAM role for a DSW instance

April

Release date

Feature

Description

References

April 29, 2024

EAS introduces serverless model deployment for AI image generation

EAS offers a serverless model deployment option ideal for model service workloads with intermittent or unpredictable traffic. With this serverless option, services incur no costs while idle. You are billed only for the GPU computation time used during image generation.

Quickly deploy Stable Diffusion for text-to-image generation in EAS

March

Release date

Feature

Description

References

March 25, 2024

DSW supports integrated AI and big data development

DSW allows you to submit data analysis and preprocessing jobs to MaxCompute or E-MapReduce using Python code. After processing, you can use the data for model training on the DSW instance's GPU or in DLC.

Connect to E-MapReduce to process big data

March 25, 2024

File transfer station is available in DSW

DSW provides a file transfer station to accelerate uploading large files, such as large models, from your local computer to a DSW instance. This allows you to upload a large file once and use it across multiple DSW instances within the same RAM account.

File transfer station

March 15, 2024

PAI-Lingjun Intelligent Computing Service is available in the Singapore region

PAI-Lingjun Intelligent Computing Service, a next-generation intelligent computing product from Alibaba Cloud, provides deeply optimized heterogeneous computing cluster instances. Proven through extensive use in AI applications, it delivers high performance, efficiency, and resource utilization. The service meets the demands of various industries, including autonomous driving, scientific research, drug discovery, finance, and the metaverse, by providing accessible intelligent computing power to accelerate technological innovation. PAI-Lingjun Intelligent Computing Service is now available in the Singapore region, and you can activate it on demand in the console.

June 6, 2024

Discontinuation of GPU servers and related algorithm components in Machine Learning Designer

Due to the expiration of the service warranty for the V100 and P100 server clusters, PAI discontinued the TensorFlow(GPU), MXNet, and PyTorch algorithm components in Machine Learning Designer on March 1, 2024. You can continue to use the cloud-native versions of these components by submitting training jobs to PAI-DLC. We recommend using the Python component in Machine Learning Designer to submit DLC jobs, which provides full feature parity and supports more framework versions. As of June 1, 2024, jobs that use these components are no longer covered by the SLA. All related clusters will be fully decommissioned by June 30, 2024.

Python script

February

Release date

Feature

Description

References

2024-02-28

Machine Learning Designer supports data preprocessing operators and common templates for LLMs

High-quality data preprocessing is crucial for successfully applying LLMs. Machine Learning Designer provides common, high-performance data preprocessing operators, including deduplication, standardization, and sensitive information masking. It leverages the large-scale distributed computing power of MaxCompute to significantly improve data preprocessing efficiency in LLM scenarios and enhance model reliability and performance.

Component reference: Foundation model data processing

2024-02-04

EAS serverless deployment is in invitational preview

EAS offers a serverless deployment option, which is ideal for model services with intermittent or unpredictable traffic. With this option, services have no startup cost, and you are billed only when GPU computing occurs. For example, with an AI painting model, you are charged only for the duration of image generation.

EAS overview

January

Release date

Feature

Description

References

2024-02-04

QuickStart is now available on the International Site

QuickStart is now available in the Singapore region.

2024-02-04

EAS now supports simple deployment

EAS provides a simplified deployment method for common use cases, including models from ModelScope and Hugging Face, as well as Triton, TFServing, LLM, and SD web application deployments. For these deployments, you only need to provide the model's storage directory to launch the service or application with a single click.

2024-02-01

One-click deployment of AI video generation applications in EAS

EAS now lets you deploy AI video generation web applications based on ComfyUI and Stable Video Diffusion models with a single click. This feature provides a quick solution for text-to-video and image-to-video generation, and helps customers in industries such as short video and live streaming platforms, gaming and entertainment, and animation production rapidly adopt AIGC.

Deploy an AI video generation application in EAS in 5 minutes

2023

December

Release date

Feature

Description

Related documentation

2023-12-13

Machine Learning Designer is now available in the Indonesia (Jakarta) region

You can now use Machine Learning Designer on demand in the Indonesia (Jakarta) region through the PAI console.

2023-12-06

DSW instances now support direct SSH access

You can now access your DSW instances from machines within your VPC or from your local development environment for easier development and training.

Remote connection: Direct SSH

November

Release date

Feature

Description

Related documentation

2023-11-20

PAI launches the AutoML platform

PAI provides AutoML, a service that enhances the machine learning capabilities of the PAI platform. AutoML integrates a variety of algorithms and distributed computing resources and supports multiple access methods. For hyperparameter tuning, it automatically finds optimal hyperparameter values, significantly improving model tuning efficiency.

How AutoML works

October

Release date

Feature

Description

Related documentation

2023-10-27

Subscription-based AI training is now available in five international regions

PAI's subscription-based AI training is now available in the China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), and Singapore regions.

September

Release date

Feature

Description

Related documentation

2023-09-28

One-click deployment of Tongyi Qianwen models on EAS

You can use PAI-EAS to deploy a web application based on the open-source Tongyi Qianwen model with a single click and perform model inference using the WebUI and API. Qwen-7B is a 7-billion-parameter model in the Tongyi Qianwen large language model series developed by Alibaba Cloud. Qwen-7B is a Transformer-based model trained on a large-scale and diverse pre-training dataset that includes web text, professional books, and code. Based on Qwen-7B, an AI assistant named Qwen-7B-Chat was also developed by using alignment mechanisms.

Use EAS to deploy LLM applications in 5 minutes

2023-09-18

DLC now supports subscriptions and alerts for monitoring metrics

PAI-DLC provides comprehensive monitoring metrics to help you understand resource load. You can use these metrics to view the resource status of your jobs and configure alert rules for them.

Training monitoring and alerts

2023-09-18

EasyCKPT, a high-performance checkpointing framework, is now available

PAI-EasyCKPT is a high-performance checkpoint framework for large model training in PyTorch. It uses strategies such as asynchronous hierarchical saving, overlapping model copying and computation, and network-aware asynchronous storage to achieve a nearly zero-overhead checkpointing mechanism with lossless model save and restore capabilities throughout the training process. It supports mainstream large model training frameworks such as Megatron and DeepSpeed and can be integrated with minimal code changes.

EasyCkpt: High-performance state saving and recovery for large AI models

August

Release date

Feature

Description

Related documentation

2023-09-04

Support for fine-tuning and deploying Stable Diffusion models

  • One-click deployment of Stable Diffusion models.

  • Solutions for fine-tuning Stable Diffusion models.

  • Quick start and deployment of Stable Diffusion WebUI.

  • Application for fine-tuning Kohya Stable Diffusion models.

AIGC

July

June

May

Release date

Feature

Description

Related documentation

2023-05-20

PAI launches the PAI SDK for Python

The PAI SDK for Python provides high-level APIs to help machine learning engineers easily train models, deploy them on PAI, and connect their end-to-end machine learning workflow.

PAI SDK for Python

April

Release date

Feature

Description

Related documentation

2023-04-19

EAS now supports an elastic resource pool

EAS supports an elastic resource pool: for a service deployed in a dedicated resource group, if resources are insufficient during a scale-out event, EAS automatically launches new instances on pay-as-you-go public resources, which are billed as a public resource group. During a scale-in event, service instances in the public resource group are scaled in first.

Elastic resource pool

2023-04-04

The EAS console launches a new quick service deployment wizard

EAS service deployment supports three methods: deploying a service from an image, deploying an AI web application from an image, and deploying a service from a model processor. You can use one-click deployment to quickly deploy AI services or applications to EAS, which simplifies the deployment process.

EAS overview

March

Release date

Feature

Description

Related documentation

2023-03-23

Model management feature upgrade

Machine Learning Designer can quickly register trained models with the model management service. PAI model management adds a model version admission mechanism. Changes to a model's admission status can trigger downstream events, such as sending automatic messages to DingTalk group chatbots or calling a specified HTTP or HTTPS service.

Model version admission and event triggers

February

Release date

Feature

Description

Related documentation

2023-02-13

EAS now supports preemptible instances

When you use a public resource group to deploy an EAS service, you can use preemptible instances to reduce running costs.

EAS spot instances

2023-02-06

EAS now supports specifying multiple instance specifications

During EAS deployment, you can specify multiple instance specifications. The system then provisions resources by iterating through the list of specifications, which significantly reduces the deployment risk caused by insufficient inventory of a single specification.

Specify multiple instance specifications

January

Release date

Feature

Description

Related documentation

2023-01-13

Machine Learning Designer now supports one-click deployment of end-to-end pipelines as online services

Machine Learning Designer can package an offline data processing pipeline that includes data preprocessing, feature engineering, and model prediction into a pipeline model, then deploy it as an EAS online service with one click.

Deploy a pipeline as an online service

2022

December

Feature

Description

Release date

Region

Related documentation

EAS adds support for Yitian 710 series computing resources

EAS now supports computing resources from the Yitian 710 series. These resources provide superior cost-effectiveness, reducing your model deployment and inference costs while improving efficiency.

2022-12-08

All regions

EAS overview

Designer adds new algorithm components

Designer now provides several new algorithm components, including the Prophet time-series algorithm, MTable Expander, MTable Assembler, and Time Window SQL. These components are available in the component tree on the left side of the Designer UI.

2022-12-05

All regions

Designer component overview

Designer adds support for custom templates

Designer now supports creating custom templates from successfully run pipelines. Use these templates to build similar pipelines faster and more efficiently.

2022-12-01

All regions

Create a pipeline: Custom template

November

Feature

Description

Release date

Region

Related documentation

EAS adds self-service O&M for machine nodes

EAS now provides self-service operations and maintenance (O&M) for machine instances through resource groups. Supported operations include viewing basic machine information, stopping and restarting node scheduling, and clearing service instances from a node.

2022-11-30

All regions

EAS overview

DSW instance updates

DSW now displays the instance lifecycle, allowing you to track status changes.

You can also view instance details and modify configurations.

2022-11-18

All regions

Create a DSW instance

September

Feature

Description

Release date

Region

Related documentation

EAS adds support for service groups and asynchronous inference

When creating an EAS service, you can assign it to a service group. A service group uses a unified ingress to distribute traffic to its services based on a defined allocation policy. You can specify the traffic ratio for each service to optimize resource utilization.

PAI provides a queue service and an asynchronous inference feature. You can perform inference by distributing requests, subscribing to pushed results, or periodically querying results.

2022-09-30

All regions

August

Feature

Description

Release date

Region

Related documentation

Designer adds new algorithm components

Designer has added several new training and prediction components, including XGBoost, DBSCAN, Gaussian Mixture Model, Ridge Regression, and Lasso Regression. These components are available in the component tree on the left side of the Designer UI.

2022-08-02

All regions

July

Feature

Description

Release date

Region

Related documentation

Designer adds custom Python Script component

Designer now includes a custom Python Script component. Use this component to develop custom algorithms and combine them with PAI's built-in algorithms for greater flexibility.

2022-07-15

Python Script

Designer is now available in US (Virginia)

Designer is now available in the US (Virginia) region. Select this region in the PAI console and create a workspace to begin using Designer features.

2022-07-05

US (Virginia)

None

EAS-benchmark for automatic service stress testing is now available

EAS now provides EAS-benchmark, a distributed, general-purpose stress testing tool. This tool enables you to create stress testing tasks and perform one-click stress tests on prediction services deployed in EAS.

2022-07-04

None

June

Feature

Description

Release date

Region

Related documentation

New visualization and analysis capabilities

Designer's visual deep learning components now integrate with TensorBoard for visualization and analysis. A dedicated dashboard also displays feature importance evaluations, correlation analyses, and scatter plots.

2022-06-22

Use TensorBoard to view analysis reports

Designer is now available in China (Hong Kong)

Designer is now available in the China (Hong Kong) region, providing hundreds of PAI's proprietary machine learning algorithms and dozens of industry templates. You can access these resources on demand in the PAI console.

2022-06-20

China (Hong Kong)

None

May

Feature

Description

Release date

Region

Related documentation

Designer is now available in Singapore and US (Silicon Valley)

Designer is now available in the Singapore and US (Silicon Valley) regions, providing hundreds of PAI's proprietary machine learning algorithms and dozens of industry templates. You can access these resources on demand in the PAI console.

2022-05-10

  • Singapore

  • US (Silicon Valley)

None

April

Feature

Description

Release date

Region

Related documentation

Fully managed Flink resources are now available

You can now purchase fully managed Flink resources and associate them with a workspace. Once associated, you can build AI pipelines for large-scale distributed model training in Designer by using a drag-and-drop interface or a PyAlink Script.

2022-04-30

Germany (Frankfurt)

Manage fully managed Flink resources

Designer adds components for anomaly detection, recommendation, data sources, and custom algorithms

Designer has introduced several new components: PyAlink Script, Read CSV File, IForest Anomaly Detection, Local Outlier Factor Anomaly Detection, One-Class SVM Anomaly Detection, and Swing Recommendation. Using the PyAlink Script component, you can call hundreds of algorithms from the Alink framework.

2022-04-16

Germany (Frankfurt)

March

Feature

Description

Release date

Region

Related documentation

Designer is now available in Germany (Frankfurt)

Designer is now available in the Germany (Frankfurt) region, providing hundreds of PAI's proprietary machine learning algorithms and dozens of industry templates. You can access these resources on demand in the PAI console.

2022-03-30

Germany (Frankfurt)

None

PAI-Blade adds support for TensorFlow 2.7

PAI-Blade now supports TensorFlow 2.7.

2022-03-27

All regions

None

DSW is now available in four new regions

You can now create DSW instances and use DSW features to build and train models in these regions.

2022-03-21

  • Singapore

  • Malaysia (Kuala Lumpur)

  • Indonesia (Jakarta)

  • Germany (Frankfurt)

None

EAS adds scheduled auto scaling and support for deploying images for gRPC or WebSocket services.

EAS now supports scheduled auto scaling, which allows you to scale instances for deployed services on a schedule. EAS also supports deploying services by using open-source TFServing or Triton.

2022-03-21

Scheduled auto scaling