All Products
Search
Document Center

Vector Retrieval Service for Milvus:Quick start: Create a Milvus instance

Last Updated:Sep 17, 2026

This topic describes how to quickly create a Vector Retrieval Service for Milvus (Milvus) instance.

Usage notes

Milvus is available in Standard Edition and StandalonePro Edition:

  • Standard Edition: A distributed vector retrieval solution designed for enterprise-grade applications and large-scale production environments. Deployed as a multi-availability zone cluster, it provides a production-level high availability SLA and supports independent horizontal scaling of compute and storage resources. This edition is ideal for production scenarios that require high reliability, high concurrency, and large-scale data processing.

  • StandalonePro Edition: A lightweight vector retrieval solution designed for individual developers and small teams. Deployed as a single process, it does not support horizontal scaling. This edition is recommended only for development, learning, feature validation, and initial testing, and is not recommended for production environments.

Feature

StandalonePro Edition

Standard Edition

Primary use case

Development, testing, and feature validation

Production environments and large-scale applications

Deployment architecture

Single-process deployment in a single availability zone

Distributed cluster with support for multi-availability zone high availability

Service level agreement (SLA)

The StandalonePro Edition SLA is lower than that of the standard cluster and does not include a production-level availability guarantee.

Provides a production-level high availability SLA

Scalability

Does not support horizontal scaling; only vertical scaling is supported

Supports independent horizontal scaling of compute and storage resources

Instance upgrade

Cannot be directly upgraded to the standard cluster edition; requires data migration

Supports smooth scaling within the cluster edition

Prerequisites

  • You have an Alibaba Cloud account. If you do not have an Alibaba Cloud account, complete the registration first. For more information, see the Alibaba Cloud account registration process.

  • When you purchase for the first time, you must grant Milvus the permissions to access the corresponding cloud resources. For more information, see Grant permissions to access cloud resources.

  • If you use a RAM user, you must complete RAM user authorization. For more information, see Grant permissions to a RAM user.

Procedure

  1. Go to the Alibaba Cloud Milvus page.

    1. Log on to the Alibaba Cloud Milvus console.

    2. In the left-side navigation pane, click Instances.

  2. On the Instances page, click Create Instance and configure the following parameters.

    Configuration item

    Description

    Billing Method

    Supports the subscription and pay-as-you-go billing methods.

    Duration

    For the subscription billing method, the default purchase duration is 1 month. The supported purchase durations are subject to the actual console page.

    Region

    The physical location where the instance resides.

    Important

    After the instance is created, the region cannot be changed. Choose carefully.

    Network

    • A virtual private cloud (VPC) is an isolated network environment that you define on Alibaba Cloud. You have full control over your own VPC.

      Select an existing VPC. To create a new VPC, click Create in Console. For more information, see Create a VPC.

    • A vSwitch is a basic network module that makes up a VPC and is used to connect different cloud resources.

      Select an existing vSwitch. To create a new vSwitch, click Create in Console. For more information, see Create a vSwitch.

    Deployment Model

    • Single-zone: Suitable for development and testing environments. It has lower costs and simple deployment, but does not provide cross-zone disaster recovery.

    • Multi-zone Deployment (Basic): Provides only one copy of compute resources, so a certain recovery time is required in the event of a failure. The RTO is within 1 hour, offering a better cost-performance ratio.

    • Multi-zone Deployment (HA): Provides two copies of compute resources. When a service anomaly occurs, you can directly switch between the primary and secondary clusters, with an RTO within 3 minutes.

    For details about the differences, applicable scenarios, and architectural characteristics of the multi-zone Basic and HA editions, see Multi-zone deployment.

    Version

    Supports versions 2.4, 2.5, and 2.6.

    Instance Series

    • Standard Edition: A distributed vector retrieval solution designed for enterprise-grade applications and large-scale production environments. Deployed as a multi-availability zone cluster, it provides a production-level high availability SLA and supports independent horizontal scaling of compute and storage resources. This edition is ideal for production scenarios that require high reliability, high concurrency, and large-scale data processing.

    • StandalonePro Edition: A lightweight vector retrieval solution designed for individual developers and small teams. Deployed as a single process, it does not support horizontal scaling. This edition is recommended only for development, learning, feature validation, and initial testing, and is not recommended for production environments.

    Service Node

    This parameter must be configured when you create a Standard Edition instance.

    Mainly responsible for processing client requests and managing the cluster state. It distributes query requests to the appropriate compute nodes and collects the results to return to you. It also maintains the cluster metadata to ensure that requests are correctly routed to the corresponding compute nodes.

    • Streaming Node: Mainly responsible for real-time writes and incremental data consumption.

      When writes are frequent and the latency requirement from "write to searchable" is high, we strongly recommend that you increase the number of nodes or upgrade the specifications to improve real-time data processing capabilities.

    • Datanode: Mainly responsible for data writing, persistence, and management.

      When the data write volume is large and imports are frequent, we strongly recommend that you increase the number of nodes or upgrade the specifications to improve overall write bandwidth and stability.

    • Proxy: Mainly responsible for receiving and routing client requests.

      When there are many client connections and high concurrent requests, we strongly recommend that you increase the number of nodes or upgrade the specifications to improve request access and forwarding capabilities.

    • Metadata Service: Responsible for resource scheduling and task coordination.

      When the cluster is large and has many data partitions, we strongly recommend that you upgrade the specifications to ensure scheduling stability.

    Compute Node

    Querynode: Responsible for vector retrieval and filtering.

    When the memory usage exceeds 70%, you must increase the number of nodes or upgrade the specifications to ensure query performance and stability. For more information about nodes, see Node types.

    Data Replicas

    Enabling data replicas effectively ensures cluster availability. We recommend that you enable this feature for production workloads.

    Data Storage

    Zone-redundant storage is used by default. You are charged based on the actual amount of data stored, and fees are calculated hourly based on the storage space (GB) that you actually use.

    Automatic Backup

    Important

    Using the backup feature incurs storage fees. For more information, see Billing.

    The automatic backup feature is enabled by default. This feature is designed to ensure the data security of your instance and guarantee the service SLA. If data is accidentally lost, you can use this feature to recover the data.

    Note

    If you need to disable this feature, after the instance is created, go to the Backup Snapshot tab to disable it. For more information, see Backup and restore.

    Password

    Set the password for the root (administrator) account of the Milvus instance to log on to the database.

    Note

    If you forget the password, see FAQ.

    OSS Data Encryption

    OSS data encryption requires calling the Key Management Service (KMS). Go to the Key Management Service console to enable it.

    Resource Group

    Select an existing resource group. To create a new dedicated resource group, click Create Resource Group. Resource groups group your cloud resources by dimensions such as purpose, permissions, and ownership. For more information, see Resource groups.

    Tag

    You can bind tags when you create an instance, or add tags after the instance is created. This helps you identify and manage your instance resources. For more information about tags, see Tags.

  3. After you confirm that the configuration is correct, read and select the service agreement, and then click Create Instance.

    When the instance status is Running, the instance is created.