2026
June
|
Release date |
Feature |
Description |
Related documentation |
|
June 30, 2026 |
DSW AgentBox released with built-in mainstream AI Coding Agents, ready to use out of the box |
Use AI Agent to assist with code writing and debugging, rapid prototyping during training task intervals, and AI-assisted full-stack development on GPU instances. For team managers, unified PAI Token configuration eliminates the need for individual API Key applications, lowering the barrier for teams to adopt AI programming tools. |
|
|
June 30, 2026 |
PAI Agentic Readiness V1.0: Core CLI and Core Skills |
Users can operate PAI in their preferred Agent without learning or using the PAI console. |
May
|
Release date |
Feature |
Description |
Related documentation |
|
May 30, 2026 |
Deep integration of commercial large models from Model Studio |
PAI integrates commercial large models from Model Studio, expanding the selection of ready-to-use models. |
April
|
Release date |
Feature |
Description |
Related documentation |
|
April 29, 2026 |
Intelligent Assistant (Xiao PAI) MCP service for DLC diagnostics |
Intelligent Assistant (Xiao PAI) now exposes its diagnostic and Q&A capabilities as an MCP service for agent ecosystem integration. Query DLC task restart counts, analyze restart reasons, and check blocklisted node status to identify root causes of training failures. |
|
|
April 23, 2026 |
DSW DockerBoard for sub-container management on shared machines |
DSW DockerBoard is now GA on Lingjun Cluster. Create and manage multiple sub-containers within a single DSW instance, enabling multiple developers to share a machine with better resource utilization. |
|
|
April 17, 2026 |
Dynamic parameters for EAS services |
EAS services now support dynamic parameters that you can hot-update in real time and perform real-time create, read, update, and delete (CRUD) operations on without restarting the service or its instances. |
|
|
April 15, 2026 |
DLC task templates |
DLC now supports task templates for different experimental tracks. Pre-fill configurations, lock critical settings, and define standard description formats to capture submission best practices, archive tasks within the same experiment, and improve iteration efficiency. |
|
|
April 13, 2026 |
DLC flexible storage mounting for Lingjun Cluster v1.0 |
DLC on Lingjun Cluster now supports flexible storage mounting v1.0. It auto-detects network characteristics and switches between VPC and VCS mounts, enabling cross-HPN-Zone access when computing power and CPFS storage are in different HPNs. |
|
|
April 13, 2026 |
GPU resource utilization metrics for DLC tasks |
DLC now reports GPU utilization metrics based on busy time—the percentage of time a GPU actively computes—helping you manage computing resources more precisely. |
|
|
April 9, 2026 |
Linux graphical desktop for DSW |
DSW instances now include a Linux graphical desktop (TurboVNC + noVNC) accessible through your browser without extra configuration. Ideal for scenarios requiring a GUI, such as autonomous driving simulations and embodied AI debugging. |
March
|
Release date |
Feature |
Description |
Related documentation |
|
March 20, 2026 |
Model warm-up cache for PAI-EAS |
The model warm-up cache preloads model caches and provides a high-speed source for cache-accelerated inference services. Suited for LLMs and AI image/video generation where large models are mounted from OSS or NAS. |
|
|
March 19, 2026 |
Global search |
Search across PAI resources—DSW instances, DLC tasks, EAS services, and models—by keyword to quickly locate assets. |
February
|
Release date |
Feature |
Description |
Related documentation |
|
February 27, 2026 |
Intelligent configuration for PD decoupling in PAI-EAS |
When deploying distributed inference with PD decoupling, Intelligent Assistant (Xiao PAI) recommends optimal PD ratios and parallelism strategies based on model characteristics, instance configurations, and I/O lengths to improve resource utilization and end-to-end performance. |
|
|
February 26, 2026 |
DSW integration with OpenClaw |
PAI-DSW automates OpenClaw installation with system control, persistent memory, and scheduled notifications. Interact through a web UI or DingTalk while the AI agent runs directly on its compute resources. |
January
|
Release date |
Feature |
Description |
Related documentation |
|
January 30, 2026 |
AutoML available in Singapore |
The PAI-AutoML service is now generally available in the Singapore region. |
|
|
January 23, 2026 |
Resource diagnostics for inference services |
EAS now provides resource quota and priority queuing diagnostics for both distributed service instances that use multiple machines and GPUs, and regular service instances, helping you identify and resolve resource issues. |
|
|
January 8, 2026 |
PAI Intelligent Assistant diagnostics (Phase 2): Support for DLC tasks and EAS services |
Phase 2 adds real-time diagnosis, root cause analysis, and repair suggestions for DLC training failures and EAS performance bottlenecks. Supports interactive, natural-language troubleshooting. |
2025
December
|
Release date |
Feature |
Description |
Related documentation |
|
2025-12-22 |
Model evaluation supports dual-model offline competition |
PAI ModelEval now supports dual-model offline competitions. Compare two models on the same Q&A dataset to select the better performer. |
|
|
2025-12-19 |
KV Store global context cache released |
EAS LLM deployments now support a global context cache for the KV Store to improve inference throughput. |
|
|
2025-12-05 |
PAI-DLC releases Ray 2.0: RayQuota |
PAI-DLC now supports Ray's native scheduling and runtime on PAI resource quotas, delivering stable and efficient performance for AI data processing and reinforcement learning in LLM and autonomous driving workloads. |
|
|
2025-12-01 |
DLC displays all IP addresses of running pods |
DLC training jobs now expose pod IP addresses across computing resource types (Lingjun and general-purpose), including head/tail node IPs, network IPs, and RDMA IPs. |
November
|
Release date |
Feature |
Description |
Related documentation |
|
2025-11-28 |
Ray on DLC supports dynamic scaling |
Ray on DLC now supports dynamic scaling. You can configure minimum and maximum instance counts and use the quota preemption mechanism for task scaling. This helps balance global tasks and resource usage. |
|
|
2025-11-28 |
PAI-EAS releases scenario-based deployment for CosyVoice |
PAI-EAS now supports the CosyVoice high-fidelity speech synthesis model, which is suitable for various scenarios, including customer service conversations, audiobook narration, and short video dubbing. |
|
|
2025-11-21 |
DSW OpenAPI supports calls over PrivateLink |
You can now call the DSW OpenAPI over PrivateLink, allowing you to manage instances and integrate features from your office network or on-premises data center. A connection to an Alibaba Cloud VPC must be established beforehand. |
|
|
2025-11-07 |
PAI-EAS supports instance-level real-time logs |
PAI-EAS now supports viewing instance-level real-time logs with streaming output. This feature is ideal for interactive debugging and other scenarios that require low-latency visibility. |
October
|
Release date |
Feature |
Description |
Related documentation |
|
2025-10-17 |
Lingjun intelligent computing GU7 series supports the 570 driver |
The Lingjun intelligent computing GU7 series now supports the 570 driver. For the GU7 series, the PAI training service supports multiple drivers, including 530, 550, and 570. You can select a driver when submitting a task without reinstallation, leveraging the flexibility of the Serverless platform. |
|
|
2025-10-17 |
Scenario-based deployment of Dify on PAI-EAS is released |
Dify is an open-source development platform for LLM applications that helps developers, enterprises, and non-technical personnel quickly build, deploy, and manage generative AI-based applications. You can deploy the open-source version of the Dify platform on EAS with a single click. This supports WebUI usage and API calls, allowing you to quickly build, deploy, and manage your applications. |
|
|
2025-10-17 |
DSW instance lifecycle events support message notifications |
1. Workspace administrators can configure message notification rules in Workspace Settings > Event Notification. You can select instance status and image save status events and specify notification targets, such as DingTalk, SMS, phone calls, or WeCom. 2. When a status event in a notification rule occurs, you receive a real-time message notification, allowing you to take prompt action. |
|
|
2025-10-17 |
Model distillation feature is released |
Based on PAI's proprietary EasyDistll algorithm library, this feature provides productized model distillation capabilities. It transfers the capabilities of a large teacher model to a smaller student model through distillation. The feature also supports online synthesis of distillation data, including instruction data for non-inference models and chain-of-thought data for inference models. This enables the student model to approach or match the teacher model's performance on specific tasks, helping you improve model performance and reduce deployment costs. |
September
|
Release date |
Feature |
Description |
Related documentation |
|
2025-09-19 |
ArtLab releases the design agent |
The Design Agent allows users to generate high-quality images, produce videos, and perform fine-grained image edits using natural language instructions. This feature streamlines the AIGC design workflow and makes advanced creative tasks more accessible. Unleash the creativity of natural language and redefine the AIGC design pipeline. |
|
|
2025-09-19 |
EAS releases compute detection and fault tolerance |
The compute detection and fault tolerance feature for EAS comprehensively inspects all resources involved in inference. It automatically isolates faulty nodes and triggers background automated O&M processes, which reduces the likelihood of issues during initial inference and improves deployment success rates. |
|
|
2025-09-17 |
Data development supports direct execution of PAI Flow nodes |
PAI Flow provides end-to-end machine learning workflow development capabilities, offering the same workflow features as PAI's visual modeling Designer and enabling periodic workflow scheduling. |
August
|
Release date |
Feature |
Description |
Related documentation |
|
2025-08-25 |
EAS supports expert parallelism deployment |
EAS now supports expert parallelism (EP) for deploying Mixture-of-Experts (MoE) models, such as DeepSeek-R1. This feature works with inference engines like vLLM and SGLang to overcome hardware limitations, improve resource utilization, and increase system throughput. |
Deploy MoE models by using expert parallelism and PD separation |
|
2025-08-15 |
Model evaluation center v1.0 released |
This out-of-the-box feature lets you run a complete model evaluation pipeline without writing code, helping you quickly determine if a model meets your business needs. |
|
|
2025-08-13 |
AI resource groups (Lingjun intelligent computing) support pay-as-you-go and savings plans |
AI resource groups (Lingjun intelligent computing) now support pay-as-you-go purchases. When combined with savings plans, the service automatically applies discounts based on the subscription duration. The longer the subscription (1, 3, or 5 years), the greater the discount. This provides a more flexible and cost-effective way to use resources. |
Create a resource group and purchase Lingjun intelligent computing resources |
|
2025-08-13 |
DataJuicer on DLC is generally available |
DLC now supports submitting DataJuicer framework tasks. By using numerous operators (over 100), multi-scale processing (single-node and multi-node), and high availability (self-healing), DataJuicer efficiently performs large-scale data cleaning, filtering, transformation, and enhancement for LLM text and multimodal data processing. |
|
|
2025-08-11 |
DLC's proprietary custom task framework (v1.0) is generally available |
DLC's proprietary distributed framework, Custom, supports PAI scheduling policies and self-healing capabilities. It also provides advanced features such as custom roles, success policies, and extended ports. This meets the computing demands of various business scenarios, including post-training for large models and autonomous driving. |
|
|
2025-08-08 |
Model weight service is released |
The model weight service significantly reduces cold start and scale-out times. It addresses the common issue of slow model loading times, overcoming a key performance bottleneck in ultra-large-scale LLM deployment. |
|
|
2025-08-07 |
EAS releases Prefill-Decode (PD) separation |
EAS has released the Prefill-Decode (PD) separation feature. It includes multiple deployment modes, such as static and dynamic PD separation, and is compatible with various inference engines like vLLM, SGLang, and BladeLLM to help you reduce inference latency. |
July
|
Release date |
Feature |
Description |
Related documentation |
|
2025-07-10 |
DSW supports distributed development and debugging environments |
This feature helps you debug and verify distributed tasks, creating a more efficient development and training workflow.
|
June
|
Release date |
Feature |
Description |
Related documentation |
|
2025-06-10 |
ArtLab supports building and sharing AIGC applications based on ComfyUI |
PAI-ArtLab enhances its enterprise-grade AIGC application capabilities. 1. Build and publish custom AIGC applications based on ComfyUI workflows. 2. Share AIGC applications from the ArtLab platform as out-of-the-box AIGC applications for PCs and H5 mobile clients. |
|
|
2025-06-05 |
Data development supports PAI Flow |
This feature unifies the entry points for big data development and AI products, and enhances the deep integration between PAI Flow and big data engines to achieve integrated big data and AI development.
|
May
|
Release date |
Feature |
Description |
Related documentation |
|
2025-05-30 |
Quick start > Model Gallery is available in the Malaysia (Kuala Lumpur) region |
Model Gallery is now available in the Malaysia (Kuala Lumpur) region. Model Gallery integrates high-quality pretrained models from various open-source AI communities to help you quickly get started with model training and deployment on PAI. |
April
|
Release date |
Feature |
Description |
Related documentation |
|
2025-04-04 |
Image management supports building custom images |
PAI now offers a custom image building feature. You can flexibly install dependencies on an existing image or build a custom image by using a Dockerfile. The image is automatically pushed to ACR and registered on the PAI platform, which meets custom requirements and eliminates the tedious process of local building and uploading. |
|
|
2025-04-01 |
DSW is generally available in US (Virginia) and US (Silicon Valley) |
DSW is now generally available in the US (Virginia) and US (Silicon Valley) regions. |
March
|
Release date |
Feature |
Description |
Related documentation |
|
2025-03-28 |
LangStudio 1.0 is released |
This version adds the following features based on LangStudio 0.1: 1. Knowledge base creation and management: You can create and synchronize knowledge bases on the console and use them when building application flows. 2. Use PAI-DSW as an application development environment: Provides DSW-based Notebook and WebIDE environments for application flow development. 3. Application flow performance evaluation: Prebuilt evaluation templates support offline performance evaluation and online service performance evaluation for applications. 4. Conversation history for deployed applications: You can use a cloud database or local storage for conversation history, and manage and export the history. |
|
|
2025-03-28 |
DSW supports dynamic mounting of NAS datasets and storage paths |
DSW now supports dynamic storage mounting. You can mount or unmount NAS datasets without restarting the instance. |
|
|
2025-03-28 |
DLC supports mounting OSS data sources by using ossfs |
DLC now supports mounting OSS data sources by using ossfs. This provides good OSS read and write performance for compute-intensive tasks such as autonomous driving, which typically involve sequential and random reads, and sequential append writes. |
|
|
2025-03-27 |
AI compute node statuses upgraded |
The compute node statuses are optimized. A new status code that indicates scheduling is prohibited is added to improve your user experience. |
|
|
2025-03-19 |
DLC supports custom roles for Ray tasks |
When you submit Ray framework tasks in DLC, you can customize the Worker role to enable hybrid execution on heterogeneous resources. |
|
|
2025-03-19 |
Resource quotas support scaling of specified nodes |
Resource quota scaling now supports node-level operations. This provides greater flexibility for managing, reallocating, and transferring computing power between quotas. |
|
|
2025-03-07 |
PAI training service is generally available in US (Silicon Valley) |
DLC and resource quotas are now available in the US (Silicon Valley) region. You can use resource quotas and public resources (pay-as-you-go) to submit training jobs. |
February
|
Release date |
Feature |
Description |
Related documentation |
|
2025-02-28 |
DLC supports read/write permissions when mounting storage instances |
When you mount Alibaba Cloud storage instances (such as OSS, NAS, and CPFS) in DLC, you can configure read/write permissions. This supports fine-grained permission management for your storage instances. |
|
|
2025-02-21 |
AI scheduling engine v2.0 implements multi-level task preemption |
Based on resource quotas, the PAI AI scheduling engine uses task classification and a dynamic priority algorithm to trigger preemption, ensuring high-priority jobs can execute quickly. Combined with AIMaster's preemptive rollback technology, interrupted jobs automatically save their intermediate state and enter a queue. They are then prioritized for resumption after resources are released. This achieves efficient scheduling in resource-constrained scenarios. |
|
|
2025-02-10 |
PAI-DLC supports nconnect when mounting Alibaba Cloud file storage |
When mounting Alibaba Cloud file storage (such as NAS and CPFS) for PAI-DLC jobs, you can now configure multiple connections (nconnect). This provides fine-grained control over the number of mount connections, which optimizes concurrent access performance from multiple nodes and ensures the stability of large-scale training jobs. |
|
|
2025-02-07 |
EAS supports multi-node distributed inference |
With the advent of ultra-large-scale Mixture-of-Experts (MoE) models like Qwen-Max and DeepSeek, a single device can no longer handle their vast number of parameters. To address this, EAS introduces a multi-node distributed inference solution that overcomes hardware limitations and efficiently supports the deployment and operation of ultra-large-scale models. EAS distributed inference supports various parallelism methods, including pipeline parallelism, tensor parallelism, and data parallelism, and is compatible with high-performance inference engine frameworks such as BladeLLM, vLLM, and SGLang. |
January
|
Release date |
Feature |
Description |
Related documentation |
|
2025-01-21 |
DLC supports training timeout alerts |
DLC now supports the configuration of training timeout alerts. You can customize timeout alert rules for training jobs during the environment preparation, queuing, and running stages. When a rule is triggered, an alert notification is sent, making it easier for you to monitor abnormal training progress. |
|
|
2025-01-21 |
DLC supports training status notifications |
DLC now supports subscriptions to training status notifications. New status events such as queuing, preempting, preparing, and running are added. This makes it easier for you to track training progress and enhances the message notification capabilities of the training service. |
|
|
2025-01-20 |
DLC supports direct mounting of storage services during job submission |
When you submit a training job by using DLC, you can directly select different storage instances in the submission form. A variety of Alibaba Cloud storage instances are supported, including OSS, General-Purpose NAS, Extreme NAS, General-Purpose CPFS, and AI-Computing CPFS. This lowers the entry barrier and simplifies usage. |
|
|
2025-01-20 |
Ray on DLC supports idle resources |
DLC now supports submitting Ray jobs by using idle resources. This allows you to run multiple types of jobs with a single set of resources, enabling resource sharing between jobs and improving resource utilization. |
|
|
2025-01-20 |
ArtLab launches industry tools |
ArtLab has launched industry tools. The first phase includes applications such as realistic e-commerce product rendering (for home appliances and furniture), corporate-style poster generation, and creative footwear design. More prebuilt applications will be continuously added. |
|
|
2025-01-20 |
ArtLab launches the AIGC Application Zone |
PAI-ArtLab has launched the AIGC Application Zone module. This allows you to use online applications that encapsulate ComfyUI workflows for operations like text-to-image and image-to-image generation. It lowers the barrier to using AIGC production tools and reduces costs through a Serverless service model. 1. Out-of-the-box: Start applications with one click, with no environment configuration required. 2. Platform-integrated enterprise-grade AIGC applications, such as generating corporate-style posters and creating profile pictures for corporate events. 3. Serverless application mode: Billing occurs only during GPU inference, significantly reducing user costs. |
|
|
2025-01-20 |
Model Gallery supports model inference acceleration |
Pretrained models in PAI-Model Gallery can be matched with supported inference acceleration capabilities (such as vLLM and BladeLLM) based on the selected machine type. |
|
|
2025-01-16 |
EAS upgrades BladeLLM high-performance deployment service |
PAI-EAS now supports scenario-based deployment of BladeLLM, achieving faster response times and higher throughput for LLM inference. BladeLLM is a PAI-proprietary inference engine that provides an efficient runtime, high-performance operator implementation, and hybrid quantization. PAI-EAS fully integrates with BladeLLM to launch a high-performance LLM inference service. It supports the deployment of prebuilt and custom models and lets you enable advanced options like model parallelism and speculative sampling with a single click, providing an efficient LLM deployment solution. |
|
|
2025-01-02 |
Model Gallery is generally available in more regions |
PAI-Model Gallery integrates pretrained models from fields such as LLM, computer vision (CV), natural language processing (NLP), and speech, providing one-stop, no-code features for model training, compression, evaluation, and deployment. The service is now available in the following regions: China (Hong Kong), Japan (Tokyo), Indonesia (Jakarta), Germany (Frankfurt), and US (Virginia). |
2024
December
|
Release date |
Feature |
Description |
References |
|
2024-12-23 |
Billing for DLC pay-as-you-go jobs now supports filtering by job type |
DLC training jobs now support a system tag (key:acs:pai:payType) to distinguish between pay-as-you-go jobs and preemptible jobs. This tag lets you quickly identify and filter job types in your billing system, providing a clearer view of your consumption and discounts. |
|
|
2024-12-16 |
Machine Learning Designer now supports aggregated execution for LLM data preprocessing pipelines |
Machine Learning Designer now supports merging and executing multiple serial large model data preprocessing (on DLC) nodes. This avoids time-consuming disk writes and the overhead from repeatedly starting and stopping distributed tasks, which improves execution efficiency. The feature also supports automatic aggregation. |
|
|
2024-12-16 |
DLC now supports preemptible jobs on general-purpose computing resources |
DLC now supports running preemptible jobs on general-purpose computing resources. This offers a more cost-effective option for AI computing power. |
|
|
2024-12-10 |
PAI training service is now available in the Germany (Frankfurt) region |
DLC and resource quotas are now available in the Germany (Frankfurt) region. You can use resource quotas to submit training jobs in this region. |
|
|
2024-12-09 |
DLC job status lifecycle updated to v2.0 |
To simplify the job lifecycle for jobs that use a resource quota, the "Queuing" and "PreAllocation" statuses are now consolidated into a single "Queuing" status. This makes job statuses clearer and easier to track. |
|
|
2024-12-06 |
DLC SanityCheck now supports custom check items |
The DLC SanityCheck feature now supports over 15 check items, including compute performance checks, node communication checks, cross-checks for compute and communication, and simulation verification. This improves troubleshooting and fault diagnosis for issues related to computing power and networks. You can also select specific checks to run, giving you greater control. |
November
|
Release date |
Feature |
Description |
References |
|
November 20, 2024 |
DSW supports dynamic mounting of OSS datasets |
1. Dynamically mount and unmount OSS datasets without restarting your instance for rapid data access. 2. A user-friendly SDK lets you mount or unmount a dataset with simple configurations or a single line of code. 3. Mount AI asset datasets, such as PAI public or custom datasets, or directly mount an OSS storage path. |
|
|
November 20, 2024 |
DSW instances support custom service access configuration |
To support the growing use of WebUI and application frameworks in AIGC development, PAI-DSW offers a custom service access configuration. Use this feature to securely share services with collaborators for testing and verification at any point during development. |
October
|
Release date |
Feature |
Description |
Related documentation |
|
2024-10-17 |
AI general computing resource groups in international regions support L20 instances |
PAI AI general computing resource groups in international regions support L20 (gn8is series) instances. |
|
|
2024-10-12 |
DLC job statuses upgraded to v1.0 |
DLC job statuses have been updated for improved clarity and simplicity. This update applies to both jobs and instances across various resource types (such as quota and bidding resource) and billing models (like subscription and pay-as-you-go). A new EnvPreparing status is added, a Bidding status is introduced for bidding jobs, and the Created status is simplified. The Queuing and PreAllocation statuses for pay-as-you-go jobs are also simplified. |
|
|
2024-10-11 |
ArtLab adds serverless ComfyUI tool |
The PAI-ArtLab toolbox includes a serverless ComfyUI tool, allowing you to perform operations such as text-to-image and image-to-image. The serverless model reduces costs by billing you only for model inference. |
|
|
2024-10-10 |
QuickStart supports DPO and CPT for LLM training |
PAI QuickStart's Model Gallery offers enhanced LLM training capabilities. In addition to the existing Supervised Fine-Tuning (SFT), this update introduces Direct Preference Optimization (DPO) and Continued Pre-training (CPT) methods. |
September
|
Release date |
Feature |
Description |
Related documentation |
|
September 29, 2024 |
DSW now integrates Tongyi Lingma |
DSW now integrates the AI coding assistant Tongyi Lingma (Personal Edition), providing features such as real-time, line-level and function-level code completion, natural language to code, unit test generation, code optimization, comment generation, code explanation, developer Q&A, and error troubleshooting. You can use these features immediately—no installation or login required—to code more efficiently. |
|
|
October 8, 2024 |
PAI Training Service is now available in China (Hong Kong) and Indonesia (Jakarta) |
DLC and AI resource quota are now available in the China (Hong Kong) and Indonesia (Jakarta) regions. You can submit training jobs in these regions using resource quotas or pay-as-you-go public resources. |
August
|
Date |
Feature |
Description |
References |
|
November 11, 2024 |
The judge model service is now generally available (GA) |
The PAI judge model service uses an LLM fine-tuned on Qwen2 as a judge to score the outputs of evaluated models. It is designed for open-ended and complex Q&A scenarios. Its key benefits are as follows: 1. Accuracy: The judge model excels at evaluating subjective questions. It intelligently classifies questions into various scenarios, such as open-ended chats, consultations, recommendations, creative writing, code generation, and role-playing. It then applies scenario-specific evaluation criteria, significantly improving evaluation accuracy. 2. Efficiency: The judge model eliminates the need for manual data labeling. Simply provide the prompts and model responses, and the service will automatically analyze and evaluate the LLM, significantly increasing evaluation efficiency. 3. Ease of use: You can create evaluation tasks through multiple methods, including the console, API calls, and SDK calls. This supports both quick-start experiences and flexible integration for developers. 4. Cost-effectiveness: At a lower cost, the service provides evaluation performance comparable to that of ChatGPT-4 in Chinese-language evaluation scenarios. |
|
|
September 3, 2024 |
DSW launches its lightweight edition (Notebook Lab) |
1. Lightweight notebook editing: You can develop directly in your browser without pre-provisioning resources. 2. Notebooks as assets: Notebooks are decoupled from instance resources, making it easier to archive and share them as technical documents or code. |
|
|
August 26, 2024 |
EAS launches LLM Intelligent Router to improve LLM inference efficiency |
When you deploy an LLM-based EAS service, you can associate it with an LLM Intelligent Router. The router intelligently distributes requests to ensure that the compute and GPU memory usage across backend instances remain balanced, improving overall cluster resource utilization. |
|
|
August 26, 2024 |
DLC training jobs on general-purpose compute resources now support CPU affinity |
For workloads whose performance is significantly affected by CPU cache affinity and scheduling, PAI DLC general-purpose compute resource groups now support CPU core binding. This feature improves job performance. |
|
|
August 15, 2024 |
EAS launches the dedicated gateway feature |
The EAS dedicated gateway is designed to meet inference requirements for security isolation and access control and to reduce network risk in high-concurrency and high-throughput scenarios. You can configure access allowlists for both internet and intranet traffic, enabling fine-grained management. Dedicated gateway resources ensure stable service connections. You can use PrivateLink to connect the intranet to your enterprise VPC. You can also enable or disable internet access at any time, giving you full control over your network. |
|
|
August 15, 2024 |
Workspaces now support custom user roles |
A workspace is the top-level concept in PAI. It provides enterprises and teams with unified management of computing resources and user permissions, along with end-to-end development tools that support team collaboration and AI asset management. As usage scenarios become more advanced, built-in roles may not meet the specific management needs of customers. For example, you might want a role that can use DSW but not DLC. To address such needs, PAI workspaces now support custom roles, allowing you to configure permissions with greater flexibility. |
|
|
August 5, 2024 |
Discontinuation of early-version PAI-PyTorch algorithm components (PyTorch100 and PyTorch131) |
Dear PAI users, Due to a system-wide upgrade, Platform for AI (PAI) will officially discontinue early-version PyTorch algorithm components in all clusters on August 30, 2024. If you submit PyTorch jobs to MaxCompute using the PAI command If you have any questions or need assistance, please contact us through your dedicated DingTalk group or by submitting a ticket. Thank you for your cooperation. |
July
|
Release date |
Feature |
Description |
References |
|
July 3, 2024 |
EAS now supports GPU sharing |
When you deploy a model in EAS, you can partition GPU resources based on compute share and memory size. This feature reduces resource costs and improves utilization. On the deployment page, you can schedule instances by either memory or compute power, enabling multiple instances to run on a single GPU. |
|
|
July 3, 2024 |
Health checks now available for EAS service instances |
The EAS health check feature helps maintain high availability by providing rapid fault detection and automatic recovery, enabling enterprise-grade inference service deployments. By using the native health check mechanism of Kubernetes, EAS automatically detects and recovers failed containers. This ensures that traffic is routed only to healthy instances and prevents resources from being allocated to unhealthy ones. |
June
|
Release date |
Feature |
Description |
Related documentation |
|
2024-07-01 |
QuickStart now supports model evaluation for LLMs |
PAI-QuickStart offers a model evaluation feature for LLMs. You can assess models by using authoritative public datasets (such as CMMLU, C-Eval, and MMLU) or your own custom datasets. This helps you determine a model's suitability for your business scenarios and compare the performance of different models. |
|
|
2024-06-19 |
PAI general computing resources support CPFS for Lingjun (invitational preview) |
The PAI training service running on general computing resources (ECI) can now mount CPFS for Lingjun, providing a cost-effective solution for data storage and computing in large-scale model scenarios. |
|
|
2024-06-12 |
Machine Learning Designer is now available in the China (Ulanqab) region |
You can now use Machine Learning Designer in the China (Ulanqab) region from the PAI console. |
|
|
2024-06-11 |
Machine Learning Designer adds the Notebook component |
Machine Learning Designer now includes the Notebook component, which connects to DSW instances. This allows you to write, debug, and run code directly within a pipeline, preserving its context and state. |
May
|
Release date |
Feature |
Description |
Related documentation |
|
2024-07-01 |
QuickStart now supports QLoRA, LoRA, and full-parameter fine-tuning for LLMs |
For LLMs in PAI-QuickStart, you can now choose from full-parameter fine-tuning or the more cost-effective LoRA and QLoRA methods. This allows you to select the best approach for your needs and reduce training costs. |
QuickStart: Fine-tune, evaluate, and deploy Qwen2.5 series models |
|
2024-06-07 |
DSW now supports configuring a RAM role for an instance |
After you grant the default PAI RAM role to an instance, you can develop and debug in DSW without configuring an AccessKey pair (AK). Authentication is handled securely through the credential chain in the following scenarios:
If a custom role is injected, the DSW instance uses the role's temporary access credentials to access specified Alibaba Cloud services, such as OSS and RDS. This ensures secure communication between the DSW instance and other Alibaba Cloud services. |
April
|
Release date |
Feature |
Description |
References |
|
April 29, 2024 |
EAS introduces serverless model deployment for AI image generation |
EAS offers a serverless model deployment option ideal for model service workloads with intermittent or unpredictable traffic. With this serverless option, services incur no costs while idle. You are billed only for the GPU computation time used during image generation. |
Quickly deploy Stable Diffusion for text-to-image generation in EAS |
March
|
Release date |
Feature |
Description |
References |
|
March 25, 2024 |
DSW supports integrated AI and big data development |
DSW allows you to submit data analysis and preprocessing jobs to MaxCompute or E-MapReduce using Python code. After processing, you can use the data for model training on the DSW instance's GPU or in DLC. |
|
|
March 25, 2024 |
File transfer station is available in DSW |
DSW provides a file transfer station to accelerate uploading large files, such as large models, from your local computer to a DSW instance. This allows you to upload a large file once and use it across multiple DSW instances within the same RAM account. |
|
|
March 15, 2024 |
PAI-Lingjun Intelligent Computing Service is available in the Singapore region |
PAI-Lingjun Intelligent Computing Service, a next-generation intelligent computing product from Alibaba Cloud, provides deeply optimized heterogeneous computing cluster instances. Proven through extensive use in AI applications, it delivers high performance, efficiency, and resource utilization. The service meets the demands of various industries, including autonomous driving, scientific research, drug discovery, finance, and the metaverse, by providing accessible intelligent computing power to accelerate technological innovation. PAI-Lingjun Intelligent Computing Service is now available in the Singapore region, and you can activate it on demand in the console. |
|
|
June 6, 2024 |
Discontinuation of GPU servers and related algorithm components in Machine Learning Designer |
Due to the expiration of the service warranty for the V100 and P100 server clusters, PAI discontinued the TensorFlow(GPU), MXNet, and PyTorch algorithm components in Machine Learning Designer on March 1, 2024. You can continue to use the cloud-native versions of these components by submitting training jobs to PAI-DLC. We recommend using the Python component in Machine Learning Designer to submit DLC jobs, which provides full feature parity and supports more framework versions. As of June 1, 2024, jobs that use these components are no longer covered by the SLA. All related clusters will be fully decommissioned by June 30, 2024. |
February
|
Release date |
Feature |
Description |
References |
|
2024-02-28 |
Machine Learning Designer supports data preprocessing operators and common templates for LLMs |
High-quality data preprocessing is crucial for successfully applying LLMs. Machine Learning Designer provides common, high-performance data preprocessing operators, including deduplication, standardization, and sensitive information masking. It leverages the large-scale distributed computing power of MaxCompute to significantly improve data preprocessing efficiency in LLM scenarios and enhance model reliability and performance. |
|
|
2024-02-04 |
EAS serverless deployment is in invitational preview |
EAS offers a serverless deployment option, which is ideal for model services with intermittent or unpredictable traffic. With this option, services have no startup cost, and you are billed only when GPU computing occurs. For example, with an AI painting model, you are charged only for the duration of image generation. |
January
|
Release date |
Feature |
Description |
References |
|
2024-02-04 |
QuickStart is now available on the International Site |
QuickStart is now available in the Singapore region. |
|
|
2024-02-04 |
EAS now supports simple deployment |
EAS provides a simplified deployment method for common use cases, including models from ModelScope and Hugging Face, as well as Triton, TFServing, LLM, and SD web application deployments. For these deployments, you only need to provide the model's storage directory to launch the service or application with a single click. |
|
|
2024-02-01 |
One-click deployment of AI video generation applications in EAS |
EAS now lets you deploy AI video generation web applications based on ComfyUI and Stable Video Diffusion models with a single click. This feature provides a quick solution for text-to-video and image-to-video generation, and helps customers in industries such as short video and live streaming platforms, gaming and entertainment, and animation production rapidly adopt AIGC. |
Deploy an AI video generation application in EAS in 5 minutes |
2023
December
|
Release date |
Feature |
Description |
Related documentation |
|
2023-12-13 |
Machine Learning Designer is now available in the Indonesia (Jakarta) region |
You can now use Machine Learning Designer on demand in the Indonesia (Jakarta) region through the PAI console. |
|
|
2023-12-06 |
DSW instances now support direct SSH access |
You can now access your DSW instances from machines within your VPC or from your local development environment for easier development and training. |
November
|
Release date |
Feature |
Description |
Related documentation |
|
2023-11-20 |
PAI launches the AutoML platform |
PAI provides AutoML, a service that enhances the machine learning capabilities of the PAI platform. AutoML integrates a variety of algorithms and distributed computing resources and supports multiple access methods. For hyperparameter tuning, it automatically finds optimal hyperparameter values, significantly improving model tuning efficiency. |
October
|
Release date |
Feature |
Description |
Related documentation |
|
2023-10-27 |
Subscription-based AI training is now available in five international regions |
PAI's subscription-based AI training is now available in the China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), and Singapore regions. |
September
|
Release date |
Feature |
Description |
Related documentation |
|
2023-09-28 |
One-click deployment of Tongyi Qianwen models on EAS |
You can use PAI-EAS to deploy a web application based on the open-source Tongyi Qianwen model with a single click and perform model inference using the WebUI and API. Qwen-7B is a 7-billion-parameter model in the Tongyi Qianwen large language model series developed by Alibaba Cloud. Qwen-7B is a Transformer-based model trained on a large-scale and diverse pre-training dataset that includes web text, professional books, and code. Based on Qwen-7B, an AI assistant named Qwen-7B-Chat was also developed by using alignment mechanisms. |
|
|
2023-09-18 |
DLC now supports subscriptions and alerts for monitoring metrics |
PAI-DLC provides comprehensive monitoring metrics to help you understand resource load. You can use these metrics to view the resource status of your jobs and configure alert rules for them. |
|
|
2023-09-18 |
EasyCKPT, a high-performance checkpointing framework, is now available |
PAI-EasyCKPT is a high-performance checkpoint framework for large model training in PyTorch. It uses strategies such as asynchronous hierarchical saving, overlapping model copying and computation, and network-aware asynchronous storage to achieve a nearly zero-overhead checkpointing mechanism with lossless model save and restore capabilities throughout the training process. It supports mainstream large model training frameworks such as Megatron and DeepSpeed and can be integrated with minimal code changes. |
EasyCkpt: High-performance state saving and recovery for large AI models |
August
|
Release date |
Feature |
Description |
Related documentation |
|
2023-09-04 |
Support for fine-tuning and deploying Stable Diffusion models |
|
July
June
May
|
Release date |
Feature |
Description |
Related documentation |
|
2023-05-20 |
PAI launches the PAI SDK for Python |
The PAI SDK for Python provides high-level APIs to help machine learning engineers easily train models, deploy them on PAI, and connect their end-to-end machine learning workflow. |
April
|
Release date |
Feature |
Description |
Related documentation |
|
2023-04-19 |
EAS now supports an elastic resource pool |
EAS supports an elastic resource pool: for a service deployed in a dedicated resource group, if resources are insufficient during a scale-out event, EAS automatically launches new instances on pay-as-you-go public resources, which are billed as a public resource group. During a scale-in event, service instances in the public resource group are scaled in first. |
|
|
2023-04-04 |
The EAS console launches a new quick service deployment wizard |
EAS service deployment supports three methods: deploying a service from an image, deploying an AI web application from an image, and deploying a service from a model processor. You can use one-click deployment to quickly deploy AI services or applications to EAS, which simplifies the deployment process. |
March
|
Release date |
Feature |
Description |
Related documentation |
|
2023-03-23 |
Model management feature upgrade |
Machine Learning Designer can quickly register trained models with the model management service. PAI model management adds a model version admission mechanism. Changes to a model's admission status can trigger downstream events, such as sending automatic messages to DingTalk group chatbots or calling a specified HTTP or HTTPS service. |
February
|
Release date |
Feature |
Description |
Related documentation |
|
2023-02-13 |
EAS now supports preemptible instances |
When you use a public resource group to deploy an EAS service, you can use preemptible instances to reduce running costs. |
|
|
2023-02-06 |
EAS now supports specifying multiple instance specifications |
During EAS deployment, you can specify multiple instance specifications. The system then provisions resources by iterating through the list of specifications, which significantly reduces the deployment risk caused by insufficient inventory of a single specification. |
January
|
Release date |
Feature |
Description |
Related documentation |
|
2023-01-13 |
Machine Learning Designer now supports one-click deployment of end-to-end pipelines as online services |
Machine Learning Designer can package an offline data processing pipeline that includes data preprocessing, feature engineering, and model prediction into a pipeline model, then deploy it as an EAS online service with one click. |
2022
December
|
Feature |
Description |
Release date |
Region |
Related documentation |
|
EAS adds support for Yitian 710 series computing resources |
EAS now supports computing resources from the Yitian 710 series. These resources provide superior cost-effectiveness, reducing your model deployment and inference costs while improving efficiency. |
2022-12-08 |
All regions |
|
|
Designer adds new algorithm components |
Designer now provides several new algorithm components, including the Prophet time-series algorithm, MTable Expander, MTable Assembler, and Time Window SQL. These components are available in the component tree on the left side of the Designer UI. |
2022-12-05 |
All regions |
|
|
Designer adds support for custom templates |
Designer now supports creating custom templates from successfully run pipelines. Use these templates to build similar pipelines faster and more efficiently. |
2022-12-01 |
All regions |
November
|
Feature |
Description |
Release date |
Region |
Related documentation |
|
EAS adds self-service O&M for machine nodes |
EAS now provides self-service operations and maintenance (O&M) for machine instances through resource groups. Supported operations include viewing basic machine information, stopping and restarting node scheduling, and clearing service instances from a node. |
2022-11-30 |
All regions |
|
|
DSW instance updates |
DSW now displays the instance lifecycle, allowing you to track status changes. You can also view instance details and modify configurations. |
2022-11-18 |
All regions |
September
|
Feature |
Description |
Release date |
Region |
Related documentation |
|
EAS adds support for service groups and asynchronous inference |
When creating an EAS service, you can assign it to a service group. A service group uses a unified ingress to distribute traffic to its services based on a defined allocation policy. You can specify the traffic ratio for each service to optimize resource utilization. PAI provides a queue service and an asynchronous inference feature. You can perform inference by distributing requests, subscribing to pushed results, or periodically querying results. |
2022-09-30 |
All regions |
August
|
Feature |
Description |
Release date |
Region |
Related documentation |
|
Designer adds new algorithm components |
Designer has added several new training and prediction components, including XGBoost, DBSCAN, Gaussian Mixture Model, Ridge Regression, and Lasso Regression. These components are available in the component tree on the left side of the Designer UI. |
2022-08-02 |
All regions |
July
|
Feature |
Description |
Release date |
Region |
Related documentation |
|
Designer adds custom Python Script component |
Designer now includes a custom Python Script component. Use this component to develop custom algorithms and combine them with PAI's built-in algorithms for greater flexibility. |
2022-07-15 |
||
|
Designer is now available in US (Virginia) |
Designer is now available in the US (Virginia) region. Select this region in the PAI console and create a workspace to begin using Designer features. |
2022-07-05 |
US (Virginia) |
None |
|
EAS-benchmark for automatic service stress testing is now available |
EAS now provides EAS-benchmark, a distributed, general-purpose stress testing tool. This tool enables you to create stress testing tasks and perform one-click stress tests on prediction services deployed in EAS. |
2022-07-04 |
None |
June
|
Feature |
Description |
Release date |
Region |
Related documentation |
|
New visualization and analysis capabilities |
Designer's visual deep learning components now integrate with TensorBoard for visualization and analysis. A dedicated dashboard also displays feature importance evaluations, correlation analyses, and scatter plots. |
2022-06-22 |
||
|
Designer is now available in China (Hong Kong) |
Designer is now available in the China (Hong Kong) region, providing hundreds of PAI's proprietary machine learning algorithms and dozens of industry templates. You can access these resources on demand in the PAI console. |
2022-06-20 |
China (Hong Kong) |
None |
May
|
Feature |
Description |
Release date |
Region |
Related documentation |
|
Designer is now available in Singapore and US (Silicon Valley) |
Designer is now available in the Singapore and US (Silicon Valley) regions, providing hundreds of PAI's proprietary machine learning algorithms and dozens of industry templates. You can access these resources on demand in the PAI console. |
2022-05-10 |
|
None |
April
|
Feature |
Description |
Release date |
Region |
Related documentation |
|
Fully managed Flink resources are now available |
You can now purchase fully managed Flink resources and associate them with a workspace. Once associated, you can build AI pipelines for large-scale distributed model training in Designer by using a drag-and-drop interface or a PyAlink Script. |
2022-04-30 |
Germany (Frankfurt) |
|
|
Designer adds components for anomaly detection, recommendation, data sources, and custom algorithms |
Designer has introduced several new components: PyAlink Script, Read CSV File, IForest Anomaly Detection, Local Outlier Factor Anomaly Detection, One-Class SVM Anomaly Detection, and Swing Recommendation. Using the PyAlink Script component, you can call hundreds of algorithms from the Alink framework. |
2022-04-16 |
Germany (Frankfurt) |
March
|
Feature |
Description |
Release date |
Region |
Related documentation |
|
Designer is now available in Germany (Frankfurt) |
Designer is now available in the Germany (Frankfurt) region, providing hundreds of PAI's proprietary machine learning algorithms and dozens of industry templates. You can access these resources on demand in the PAI console. |
2022-03-30 |
Germany (Frankfurt) |
None |
|
PAI-Blade adds support for TensorFlow 2.7 |
PAI-Blade now supports TensorFlow 2.7. |
2022-03-27 |
All regions |
None |
|
DSW is now available in four new regions |
You can now create DSW instances and use DSW features to build and train models in these regions. |
2022-03-21 |
|
None |
|
EAS adds scheduled auto scaling and support for deploying images for gRPC or WebSocket services. |
EAS now supports scheduled auto scaling, which allows you to scale instances for deployed services on a schedule. EAS also supports deploying services by using open-source TFServing or Triton. |
2022-03-21 |