This topic describes how to quickly create a Vector Retrieval Service for Milvus (Milvus) instance.
Usage notes
Milvus is available in Standard Edition and StandalonePro Edition:
Standard Edition: A distributed vector retrieval solution designed for enterprise-grade applications and large-scale production environments. Deployed as a multi-availability zone cluster, it provides a production-level high availability SLA and supports independent horizontal scaling of compute and storage resources. This edition is ideal for production scenarios that require high reliability, high concurrency, and large-scale data processing.
StandalonePro Edition: A lightweight vector retrieval solution designed for individual developers and small teams. Deployed as a single process, it does not support horizontal scaling. This edition is recommended only for development, learning, feature validation, and initial testing, and is not recommended for production environments.
Feature | StandalonePro Edition | Standard Edition |
Primary use case | Development, testing, and feature validation | Production environments and large-scale applications |
Deployment architecture | Single-process deployment in a single availability zone | Distributed cluster with support for multi-availability zone high availability |
Service level agreement (SLA) | The StandalonePro Edition SLA is lower than that of the standard cluster and does not include a production-level availability guarantee. | Provides a production-level high availability SLA |
Scalability | Does not support horizontal scaling; only vertical scaling is supported | Supports independent horizontal scaling of compute and storage resources |
Instance upgrade | Cannot be directly upgraded to the standard cluster edition; requires data migration | Supports smooth scaling within the cluster edition |
Prerequisites
You have an Alibaba Cloud account. If you do not have an Alibaba Cloud account, complete the registration first. For more information, see the Alibaba Cloud account registration process.
When you purchase for the first time, you must grant Milvus the permissions to access the corresponding cloud resources. For more information, see Grant permissions to access cloud resources.
If you use a RAM user, you must complete RAM user authorization. For more information, see Grant permissions to a RAM user.
Procedure
Go to the Alibaba Cloud Milvus page.
Log on to the Alibaba Cloud Milvus console.
In the left-side navigation pane, click Instances.
On the Instances page, click Create Instance and configure the following parameters.
Configuration item
Description
Billing Method
Supports the subscription and pay-as-you-go billing methods.
Duration
For the subscription billing method, the default purchase duration is 1 month. The supported purchase durations are subject to the actual console page.
Region
The physical location where the instance resides.
ImportantAfter the instance is created, the region cannot be changed. Choose carefully.
Network
A virtual private cloud (VPC) is an isolated network environment that you define on Alibaba Cloud. You have full control over your own VPC.
Select an existing VPC. To create a new VPC, click Create in Console. For more information, see Create a VPC.
A vSwitch is a basic network module that makes up a VPC and is used to connect different cloud resources.
Select an existing vSwitch. To create a new vSwitch, click Create in Console. For more information, see Create a vSwitch.
Deployment Model
Single-zone: Suitable for development and testing environments. It has lower costs and simple deployment, but does not provide cross-zone disaster recovery.
Multi-zone Deployment (Basic): Provides only one copy of compute resources, so a certain recovery time is required in the event of a failure. The RTO is within 1 hour, offering a better cost-performance ratio.
Multi-zone Deployment (HA): Provides two copies of compute resources. When a service anomaly occurs, you can directly switch between the primary and secondary clusters, with an RTO within 3 minutes.
For details about the differences, applicable scenarios, and architectural characteristics of the multi-zone Basic and HA editions, see Multi-zone deployment.
Version
Supports versions 2.4, 2.5, and 2.6.
Instance Series
Standard Edition: A distributed vector retrieval solution designed for enterprise-grade applications and large-scale production environments. Deployed as a multi-availability zone cluster, it provides a production-level high availability SLA and supports independent horizontal scaling of compute and storage resources. This edition is ideal for production scenarios that require high reliability, high concurrency, and large-scale data processing.
StandalonePro Edition: A lightweight vector retrieval solution designed for individual developers and small teams. Deployed as a single process, it does not support horizontal scaling. This edition is recommended only for development, learning, feature validation, and initial testing, and is not recommended for production environments.
Service Node
This parameter must be configured when you create a Standard Edition instance.
Mainly responsible for processing client requests and managing the cluster state. It distributes query requests to the appropriate compute nodes and collects the results to return to you. It also maintains the cluster metadata to ensure that requests are correctly routed to the corresponding compute nodes.
Streaming Node: Mainly responsible for real-time writes and incremental data consumption.
When writes are frequent and the latency requirement from "write to searchable" is high, we strongly recommend that you increase the number of nodes or upgrade the specifications to improve real-time data processing capabilities.
Datanode: Mainly responsible for data writing, persistence, and management.
When the data write volume is large and imports are frequent, we strongly recommend that you increase the number of nodes or upgrade the specifications to improve overall write bandwidth and stability.
Proxy: Mainly responsible for receiving and routing client requests.
When there are many client connections and high concurrent requests, we strongly recommend that you increase the number of nodes or upgrade the specifications to improve request access and forwarding capabilities.
Metadata Service: Responsible for resource scheduling and task coordination.
When the cluster is large and has many data partitions, we strongly recommend that you upgrade the specifications to ensure scheduling stability.
Compute Node
Querynode: Responsible for vector retrieval and filtering.
When the memory usage exceeds 70%, you must increase the number of nodes or upgrade the specifications to ensure query performance and stability. For more information about nodes, see Node types.
Data Replicas
Enabling data replicas effectively ensures cluster availability. We recommend that you enable this feature for production workloads.
Data Storage
Zone-redundant storage is used by default. You are charged based on the actual amount of data stored, and fees are calculated hourly based on the storage space (GB) that you actually use.
Automatic Backup
ImportantUsing the backup feature incurs storage fees. For more information, see Billing.
The automatic backup feature is enabled by default. This feature is designed to ensure the data security of your instance and guarantee the service SLA. If data is accidentally lost, you can use this feature to recover the data.
NoteIf you need to disable this feature, after the instance is created, go to the Backup Snapshot tab to disable it. For more information, see Backup and restore.
Password
Set the password for the root (administrator) account of the Milvus instance to log on to the database.
NoteIf you forget the password, see FAQ.
OSS Data Encryption
OSS data encryption requires calling the Key Management Service (KMS). Go to the Key Management Service console to enable it.
Resource Group
Select an existing resource group. To create a new dedicated resource group, click Create Resource Group. Resource groups group your cloud resources by dimensions such as purpose, permissions, and ownership. For more information, see Resource groups.
Tag
You can bind tags when you create an instance, or add tags after the instance is created. This helps you identify and manage your instance resources. For more information about tags, see Tags.
After you confirm that the configuration is correct, read and select the service agreement, and then click Create Instance.
When the instance status is Running, the instance is created.