Model customization allows you to fine-tune a text embedding model with your business data to enhance its performance. You can also train a custom embedding dimensionality reduction model by using vector data that you provide. In a typical business scenario, you first generate embeddings from text or queries by using a text embedding model, and then use an embedding dimensionality reduction model to reduce the dimensionality of the embeddings.
Background information
In intelligent search and RAG scenarios, vector model performance is crucial to business outcomes. However, the effectiveness of general-purpose vector models is often limited by their training data in specific domains. To improve retrieval performance, you can fine-tune a general-purpose model with your business data. Additionally, higher vector dimensionality significantly increases storage and computation costs for large-scale data. To address this, AI Search Open Platform provides an embedding dimensionality reduction service. This service allows you to train a custom model that converts high-dimensional vectors into lower-dimensional ones, helping you save costs without a significant loss in performance.
Billing
You are charged for training based on the computing units (CUs) consumed. Each CU is priced at CNY 3.87. The actual number of CUs consumed depends on the volume and dimensionality of your training data. For example, training with a minimum of 100,000 records of 1,024 dimensions consumes approximately 250 CUs, for a total cost of 250 × 3.87 = CNY 967.50.
Customize embedding dimensionality reduction
-
In the AI Search Open Platform console, navigate to Model Service > Model Customization, and then click Create.
A RAM user must have the required model service permissions to create a model, modify its configuration, or view its details.
-
On the Model Customization page, configure the following parameters.
Parameter
Description
Model name
The name used to invoke the embedding dimensionality reduction service.
Model type
The type of model to train. Select Vector Dimensionality Reduction (embedding-dim-reduction).
Model service
The base model for training, such as ops-embedding-dim-reduction-001.
Training data source
MaxCompute or OSS
MaxCompute data source
Parameter
Description
Training data source
MaxCompute.
Region
The region where your MaxCompute project is located.
Project name
The name of your project in MaxCompute.
AccessKey ID
The AccessKey ID of the Alibaba Cloud account or RAM user with read and write permissions for MaxCompute.
You can obtain an AccessKey ID from the AccessKey Management page.
AccessKey secret
The AccessKey secret that corresponds to the AccessKey ID.
Table name
The name of the table in MaxCompute that stores your training data.
Table partition
The partition information of the table.
Training fields
To select the primary key field and String-type vector fields, grant the GetTableFields (get MaxCompute table schema) permission to the RAM user with read and write permissions on the MaxCompute table schema. The vector dimensionality must be between 1,024 and 4,096.
OSS data source
Parameter
Description
Training data source
OSS
Region
The region where your OSS Bucket is located.
OSS Bucket
The name of your OSS Bucket.
Doc data
The data in OSS used for training.
OSS Endpoint
This value is automatically generated after you configure the preceding parameters.
-
In the confirmation dialog, click OK. In the confirmation dialog box that appears, click Create and Train. The model enters a pre-processing state. Training starts after pre-processing is complete.
Alternatively, click Confirm Creation. You can then find the model with the Pending Training status in the model customization list and start the training job later.
In the model list, a model with the Available status is fully trained and ready for invocation. Click Try Now to test the performance of the fine-tuned embedding model.
Customize text embedding
-
In the AI Search Open Platform console, navigate to Model Service > Model Customization, and then click Create.
A RAM user must have the required model service permissions to create a model, modify its configuration, or view its details.
-
On the Model Customization page, configure the following parameters.
Parameter
Description
Model name
A custom name for the model.
Model type
The type of model to train. Select Text Embedding.
Base model
The base model for training, such as ops-text-embedding-001.
Embedding dimensionality reduction
If you enable this option, an embedding dimensionality reduction training job also runs.
Base model for dimensionality reduction
Specifies the model for dimensionality reduction. This parameter is available only if you enable Embedding Dimensionality Reduction.
Training data source
MaxCompute or OSS.
MaxCompute data source
Parameter
Description
Training data source
MaxCompute.
Region
The region where your MaxCompute project is located.
Project name
The name of your project in MaxCompute.
AccessKey ID
The AccessKey ID of the Alibaba Cloud account or RAM user with read and write permissions for MaxCompute.
You can obtain an AccessKey ID from the AccessKey Management page.
AccessKey secret
The AccessKey secret that corresponds to the AccessKey ID.
Table name
The name of the table in MaxCompute that stores your training data.
Table partition
The partition information of the table.
Training fields
To select the primary key field and String-type text data, grant the GetTableFields (get MaxCompute table schema) permission to the RAM user with read and write permissions on the MaxCompute table schema.
Query-doc pairs
For the required data format, see the sample data in the console.
OSS data source
Parameter
Description
Training data source
OSS
Region
The region where your OSS Bucket is located.
OSS Bucket
The name of your OSS Bucket.
Doc data
The data in OSS used for training.
Query-doc pairs
For the required data format, see the sample data in the console.
OSS Endpoint
This value is automatically generated after you configure the preceding parameters.
-
Click OK. In the confirmation dialog box that appears, click Create and Train. The model enters a pre-processing state. Training starts after pre-processing is complete.
Alternatively, click Confirm Creation. You can then find the model with the Pending Training status in the model customization list and start the training job later.
In the model list, a model with the Available status is ready for deployment.
Service invocation
When the model's performance is satisfactory, you can invoke the service by using an API. For more information, see Embedding Dimensionality Reduction API and Custom Deployment Service API.