All Products
Search
Document Center

Alibaba Cloud Model Studio:Fine-tune with the API or CLI

Last Updated:Jul 14, 2026

Learn how to tune Qwen models in Alibaba Cloud Model Studio using the API (HTTP) and the command line (shell). Model tuning involves three methods: supervised fine-tuning (SFT), continual pre-training (CPT), and direct preference optimization (DPO).

Prerequisites

Note

The API supports only token-based billing for training jobs. To use model training units (prepaid or postpaid), create the job in the console.

Tuning file upload

Preparing fine-tuning files

SFT training set

SFT ChatML (Chat Markup Language) format training data supports multi-turn conversations and various role settings.

The OpenAI name and weight parameters are not supported. All assistant outputs will be trained.
# A single line of training data (in JSON format) has the following typical structure when expanded:
{"messages": [
  {"role": "system", "content": "System input 1"}, 
  {"role": "user", "content": "User input 1"}, 
  {"role": "assistant", "content": "Expected model output 1"}, 
  {"role": "user", "content": "User input 2"}, 
  {"role": "assistant", "content": "Expected model output 2"}
  ...
]}

For information about the differences between system, user, and assistant, see Overview. Sample training datasets: SFT-ChatML_format_example.jsonl, SFT-ChatML_format_example.xlsx. XLS and XLSX formats support only single-turn conversations.

All assistant lines in a single training data entry support the "loss_weight" parameter, which sets the relative importance of that line during training. (Range: 0.0 to `1.0`. A larger value indicates higher importance.)

This parameter is available for invitational preview. To use it, contact your account manager.
 {"role": "assistant", "content": "Expected model output 1", "loss_weight": 1.0}, 
 {"role": "assistant", "content": "Expected model output 2", "loss_weight": 0.5}

You can also download a data template from the Model Studio console.

image

Upload fine-tuning file to Model Studio

OpenAI-compatible Files API

import os
from pathlib import Path
from openai import OpenAI


client = OpenAI(
    # If you have not configured an environment variable, replace the following line with api_key="sk-xxx" and use your Model Studio API key.
    # API keys for the Singapore and China (Beijing) regions are different. Get an API key: https://www.alibabacloud.com/help/en/model-studio/get-api-key
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    # The following is the URL for the Singapore region. If you use a service in the China (Beijing) region, replace the URL with: https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1
    base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)

# test.jsonl is a local sample file.
file_object = client.files.create(file=Path("test.jsonl"), purpose="fine-tune")

print(file_object.model_dump_json())
import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.files.*;

import java.nio.file.Path;
import java.nio.file.Paths;

public class Main {
    public static void main(String[] args) {
        // Create a client and use the API key from the environment variable.
        OpenAIClient client = OpenAIOkHttpClient.builder()
                // API keys for the Singapore and China (Beijing) regions are different. Get an API key: https://www.alibabacloud.com/help/en/model-studio/get-api-key
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                // The following is the URL for the Singapore region. If you use a service in the China (Beijing) region, replace the URL with: https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1
                .baseUrl("https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1")
                .build();
        // Set the file path. Modify the path and filename as needed.
        Path filePath = Paths.get("src/main/java/org/example/test.txt");
        // Create file upload parameters.
        FileCreateParams params = FileCreateParams.builder()
                .file(filePath)
                .purpose(FilePurpose.of("fine-tune"))
                .build();

        // Upload the file.
        FileObject fileObject = client.files().create(params);
        System.out.println(fileObject);
    }
}
# ======= Important =======
# API keys for the Singapore and China (Beijing) regions are different. Get an API key: https://www.alibabacloud.com/help/en/model-studio/get-api-key
# The following is the URL for the Singapore region. If you use a service in the China (Beijing) region, replace the URL with: https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/files
# === Delete this comment before running ===

curl -X POST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/files \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
--form 'file=@"test.jsonl"' \
--form 'purpose="fine-tune"'
Note

Limitations:

  • The maximum size of a single file is 300 MB.

  • The total size of all non-deleted files is limited to 5 GB.

  • You can store a maximum of 100 non-deleted files.

  • Files are stored indefinitely.

Model fine-tuning

Create a fine-tuning job

HTTP

For Windows CMD, replace ${DASHSCOPE_API_KEY} with %DASHSCOPE_API_KEY%. For PowerShell, use $env:DASHSCOPE_API_KEY.
curl --location "https://dashscope.aliyuncs.com/api/v1/fine-tunes" \
--header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
--header 'Content-Type: application/json' \
--data '{
    "model":"qwen3-8b",
    "training_file_ids":[
        "<your_training_file_id_1>",
        "<your_training_file_id_2>"
    ],
    "hyper_parameters":
    {
        "n_epochs": 3,
        "batch_size": 16,
        "max_length": 8192,
        "learning_rate": "1.6e-5",
        "lr_scheduler_type": "linear",
        "split": 0.9,
        "warmup_ratio": 0.05,
        "eval_steps": 50,
        "data_augmentation": true,
        "augmentation_ratio": "0.1,0.05,0.15",
        "augmentation_types": "dialogue_CN,general_purpose_CN,NLP",
        "save_strategy": "epoch",
        "save_total_limit": 10
    },
    "training_type":"sft"
}'

Parameters

Parameter

Required

Type

Location

Description

training_file_ids

Yes

Array

Body

A list of file IDs for the training set.

validation_file_ids

No

Array

Body

A list of file IDs for the validation set.

model

Yes

String

Body

The ID of the base model for fine-tuning, or the ID of a previously fine-tuned model.

hyper_parameters

No

Map

Body

Hyperparameters for the fine-tuning job. Supported parameters and their default values vary by model. To view the default values, go to the console and select the same model and fine-tuning method.

The following parameters are required because they affect the training cost: n_epochs (number of epochs), batch_size (batch size), and max_length (sequence length).

training_type

No

String

Body

Specifies the fine-tuning method. Valid values are:

cpt

sft

efficient_sft

dpo_full

dpo_lora

job_name

No

String

Body

Specifies the name of the fine-tuning job.

model_name

No

String

Body

Specifies the name of the fine-tuned model. This is not the model ID, which is generated by the system.

Response

{
    "request_id": "635f7047-003e-4be3-b1db-6f98e239f57b",
    "output":
    {
        "job_id": "ft-202511272033-8ae7",
        "job_name": "ft-202511272033-8ae7",
        "status": "PENDING",
        "finetuned_output": "qwen3-8b-ft-202511272033-8ae7",
        "model": "qwen3-8b",
        "base_model": "qwen3-8b",
        "training_file_ids":
        [
            "9e9ffdfa-c3bf-436e-9613-6f053c66aa6e"
        ],
        "validation_file_ids":
        [],
        "hyper_parameters":
        {
            "n_epochs": 3,
            "batch_size": 16,
            "max_length": 8192,
            "learning_rate": "1.6e-5",
            "lr_scheduler_type": "linear",
            "split": 0.9,
            "warmup_ratio": 0.05,
            "eval_steps": 50,
            "data_augmentation": true,
            "augmentation_ratio": "0.1,0.05,0.15",
            "augmentation_types": "dialogue_CN,general_purpose_CN,NLP",
            "save_strategy": "epoch",
            "save_total_limit": 10
        },
        "training_type": "sft",
        "create_time": "2025-11-27 20:33:15",
        "workspace_id": "llm-8v53etv3hwb8orx1",
        "user_identity": "1654290265984853",
        "modifier": "1654290265984853",
        "creator": "1654290265984853",
        "group": "llm",
        "max_output_cnt": 10
    }
}

Base models (model) and training types (training_type)

Supported models
Singapore
Text generation

Model name

Model code

SFT full-parameter training (sft)

SFT efficient training (efficient_sft)

Qwen3-14B

qwen3-14b

×

Supported

Visual understanding (Qwen-VL)

Model name

Model code

SFT full-parameter training (sft)

SFT efficient training (efficient_sft)

-

-

-

-

North China 2 (Beijing)
Text generation

Model service

Model code

CPT full-parameter training (cpt)

SFT full-parameter training (sft)

SFT efficient training (efficient_sft)

DPO full-parameter training (dpo_full)

DPO efficient training (dpo_lora)

Qwen3.6-Flash-2026-04-16

qwen3.6-flash-2026-04-16

×

Supported

×

×

×

Qwen3.5-27B

qwen3.5-27b

×

Supported

Supported

×

×

Qwen3.5-9B

qwen3.5-9b

×

Supported

Supported

×

×

Qwen3.5-Flash-2026-02-23

qwen3.5-flash-2026-02-23

×

Supported

×

×

×

Qwen3-32B

qwen3-32b

Supported

Supported

Supported

Supported

Supported

Qwen3-30B-A3B-Instruct-2507

qwen3-30b-a3b-instruct-2507

Supported

Supported

Supported

×

×

Qwen3-14B

qwen3-14b

×

Supported

Supported

Supported

Supported

Qwen3-8B

qwen3-8b

×

Supported

Supported

Supported

Supported

Qwen3-4B-Instruct-2507

qwen3-4b-instruct-2507

Supported

Supported

Supported

Supported

Supported

Qwen3-1.7B

qwen3-1.7b

Supported

Supported

Supported

Supported

Supported

Qwen3-0.6B

qwen3-0.6b

Supported

Supported

Supported

Supported

Supported

Qwen2.5-72B-Instruct

qwen2.5-72b-instruct

Supported

Supported

Supported

Supported

Supported

Qwen2.5-32B-Instruct

qwen2.5-32b-instruct

Supported

Supported

Supported

Supported

Supported

Qwen2.5-14B-Instruct

qwen2.5-14b-instruct

Supported

Supported

Supported

Supported

Supported

Qwen2.5-7B-Instruct

qwen2.5-7b-instruct

Supported

Supported

Supported

Supported

Supported

Qwen-Plus-Character-2025-11-06

qwen-plus-character-2025-11-06

×

Supported

Supported

Supported

Supported

Visual understanding (Qwen-VL)

Model service

Model code

CPT full-parameter training (cpt)

SFT full-parameter training (sft)

SFT efficient training (efficient_sft)

DPO full-parameter training (dpo_full)

DPO efficient training (dpo_lora)

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

×

Supported

Supported

×

×

Qwen3-VL-8B-Thinking

qwen3-vl-8b-thinking

×

Supported

Supported

×

×

Qwen3-VL-4B-Instruct

qwen3-vl-4b-instruct

×

Supported

Supported

×

×

Qwen2.5-VL-72B-Instruct

qwen2.5-vl-72b-instruct

×

Supported

Supported

×

×

Qwen2.5-VL-32B-Instruct

qwen2.5-vl-32b-instruct

×

Supported

Supported

×

×

Qwen2.5-VL-7B-Instruct

qwen2.5-vl-7b-instruct

×

Supported

Supported

×

×

Comparison of tuning methods

Feature

CPT (Continual Pre-training)

SFT (Supervised Fine-tuning)

DPO (Direct Preference Optimization)

Summary

Supplements knowledge (Injects domain knowledge)

Learns to perform tasks (Follows instructions)

Performs tasks better (Aligns with human preferences)

Input data

10 million+ tokens

Unlabeled domain text

Over 1,000 entries

High-quality "question-answer" pairs

100+ sets

"Better-worse" response pairs for the same instruction

Core objective

Domain adaptation. Learns specialized vocabulary and facts.

Teaches the model conversation formats and task execution capabilities.

Makes model outputs better align with human values and preferences.

Learning method

Self-supervised learning (Predicts the next word)

Supervised learning (Imitates the ground truth)

Direct preference learning (Increases the probability of good responses and decreases the probability of bad responses)

Model stage

Typically before SFT

After CPT and before DPO

Typically after SFT, as the final step for alignment.

Comparison of training patterns

Full-parameter training

Efficient training (LoRA, recommended)

Scenarios

• The model needs to acquire new capabilities

• Achieving optimal global performance.

• Optimizing model performance for specific scenarios.

• For cost-sensitive and time-sensitive scenarios.

Training time

Longer, with slower convergence.

Shorter, with faster convergence.

hyper_parameters: Supported settings

Supported parameters and their default values vary by model. To view specific default values, go to the console and select the model and training method.

Parameter

Recommended setting

Type

Description

n_epochs

(Number of epochs) [Required]

Data size < 10,000: 3–5

Data size > 10,000: 1–2

Integer

The number of times the model iterates over the entire training set. Adjust this value based on your fine-tuning experience.

More epochs increase training time and cost.

learning_rate

(Learning rate)

Use the recommended default value provided by Model Studio.

Float

Controls the step-size for weight updates during training.

  • A high learning rate can cause training to diverge, yielding poor results.

  • A low learning rate can result in slow training and minimal performance improvement.

freeze_vit

(Freeze visual backbone)

Adjust as needed

Boolean

Whether to freeze the visual backbone parameters, which prevents their weights from being updated during training. This parameter applies only to Qwen-VL (visual understanding) models.

Warning

Token-based billing is available only when freeze_vit is set to “true”.

batch_size

(Batch size) [Required]

Use the recommended default value provided by Model Studio.

Integer

Specifies the number of training examples processed in one iteration. A small value can significantly increase training time. Default values vary by model; check the console for details.

eval_steps

(Evaluation steps)

Adjust as needed

Integer

The step interval for evaluating the model's training accuracy and loss.

This parameter affects how often Validation Loss and Validation Token Accuracy are displayed during fine-tuning.

logging_steps

(Logging steps)

Adjust as needed

Integer

The step interval for logging fine-tuning progress.

lr_scheduler_type

(Learning rate scheduler)

Recommended linear/Inverse_sqrt

String

The strategy for dynamically adjusting the learning rate during training.

For details about each strategy, see Fine-tune a model in the console.

max_length

(Sequence length) [Required]

8192

Integer

The maximum sequence length (in tokens) for a single training example. Examples that exceed this length are discarded.

For the relationship between characters and tokens, see How to convert between tokens and characters.

max_split_val_dataset_sample

(Max validation set samples)

Use the recommended default value provided by Model Studio.

Integer

When "validation_file_ids" is not set, the validation set that is automatically split by Model Studio contains a maximum of 1,000 entries.

This parameter has no effect when "validation_file_ids" is set.

split

(Training set ratio)

Use the recommended default value provided by Model Studio.

Float

If you do not set "validation_file_ids", Alibaba Cloud Model Studio automatically uses 80% of the training file as the training set and 20% as the validation set.

When "validation_file_ids" is set, this parameter has no effect.

warmup_ratio

(Warm-up ratio)

Use the recommended default value provided by Model Studio.

Float

The proportion of the total training process used for learning rate warm-up. During warm-up, the learning rate linearly increases from a small initial value to the specified learning rate.

This parameter helps stabilize training by limiting large parameter changes at the beginning.

A ratio that is too high has an effect similar to a low learning rate, causing minimal performance changes.

A ratio that is too low has an effect similar to a high learning rate and may degrade model performance.

This parameter has no effect if the learning rate scheduler is set to constant.

weight_decay

(Weight decay)

Use the recommended default value provided by Model Studio.

Float

The strength of L2 regularization. Regularization helps maintain the model's generalization ability. An excessively high value can reduce fine-tuning effectiveness.

Parameters for efficient fine-tuning (supports efficient_sft and dpo_lora)

Note

When you perform a second round of efficient fine-tuning on a model that has already been efficiently fine-tuned, the lora_rank, lora_alpha, and lora_dropout parameters must remain consistent.

lora_rank

(LoRA rank)

64

Integer

The rank of the low-rank matrices in LoRA. A higher rank can improve fine-tuning results but may slightly increase training time.

lora_alpha

(LoRA alpha)

Use the recommended default value provided by Model Studio.

Integer

The scaling factor that controls the combination of original model weights and the LoRA low-rank correction term.

A larger alpha gives more weight to the LoRA correction, making the model rely more on task-specific information.

A smaller alpha makes the model retain more knowledge from the base model.

lora_dropout

(LoRA dropout)

Use the recommended default value provided by Model Studio.

Float

The dropout rate for the values in the low-rank matrices during LoRA training.

Using the recommended value enhances the model's generalization capabilities.

An overly large value can diminish the fine-tuning effect.

Parameters for publishing model parameter snapshots (for efficient_sft and sft only)

save_strategy

(Snapshot save strategy)

It can be set to epoch or steps.

  • When set to steps, you can adjust the save interval by setting the save_steps parameter.

String

The strategy for saving model parameter snapshots (checkpoints). Valid options are epoch (save after each epoch) and steps (save at a specified step interval).

save_steps

(Save steps)

If you need to modify it manually, set it to an integer multiple of the eval_steps parameter.

Integer

The interval, in training steps, between each model parameter snapshot save.

save_total_limit

(Snapshot save limit)

10

Integer

The maximum number of model parameter snapshots to retain. Once the limit is reached, older snapshots are automatically deleted.

Retrieve a fine-tuning job

To retrieve the details of a fine-tuning job, use the job_id returned when you create the job.

HTTP

curl 'https://dashscope.aliyuncs.com/api/v1/fine-tunes/<job_id>' \
--header 'Authorization: Bearer '${DASHSCOPE_API_KEY} \
--header 'Content-Type: application/json'

Request parameters

Parameter

Type

Location

Required

Description

job_id

String

Path

Yes

The ID of the fine-tuning job.

Successful response

{
    "request_id": "d100cddb-ac85-4c82-bd5c-9b5421c5e94d",
    "output":
    {
        "job_id": "ft-202511272033-8ae7",
        "job_name": "ft-202511272033-8ae7",
        "status": "RUNNING",
        "finetuned_output": "qwen3-8b-ft-202511272033-8ae7",
        "model": "qwen3-8b",
        "base_model": "qwen3-8b",
        "training_file_ids":
        [
            "9e9ffdfa-c3bf-436e-9613-6f053c66aa6e"
        ],
        "validation_file_ids":
        [],
        "hyper_parameters":
        {
            "n_epochs": 3,
            "batch_size": 16,
            "max_length": 8192,
            "learning_rate": "1.6e-5",
            "lr_scheduler_type": "linear",
            "split": 0.9,
            "warmup_ratio": 0.05,
            "eval_steps": 50,
            "data_augmentation": true,
            "augmentation_ratio": "0.1,0.05,0.15",
            "augmentation_types": "dialogue_CN,general_purpose_CN,NLP",
            "save_strategy": "epoch",
            "save_total_limit": 10
        },
        "training_type": "sft",
        "create_time": "2025-11-27 20:33:15",
        "workspace_id": "llm-8v53etv3hwb8orx1",
        "user_identity": "1654290265984853",
        "modifier": "1654290265984853",
        "creator": "1654290265984853",
        "group": "llm",
        "max_output_cnt": 10
    }
}

Job status

Description

PENDING

The job is waiting to start.

QUEUING

The job is queued. Only one fine-tuning job runs at a time.

RUNNING

The job is running.

CANCELING

The job is being canceled.

SUCCEEDED

The job succeeded.

FAILED

The job failed.

CANCELED

The job was canceled.

Note

After a fine-tuning job succeeds, the finetuned_output field provides the resulting model ID. Use this ID for model deployment.

Get fine-tuning job logs

HTTP

curl 'https://dashscope.aliyuncs.com/api/v1/fine-tunes/<job_id>/logs?offset=0&line=1000' \
--header 'Authorization: Bearer '${DASHSCOPE_API_KEY} \
--header 'Content-Type: application/json' 
Use the offset and line parameters to retrieve a range of log lines. The offset parameter specifies the starting line, and the line parameter specifies the maximum number of lines to return.

Sample response:

{
    "request_id":"1100d073-4673-47df-aed8-c35b3108e968",
    "output":{
        "total":57,
        "logs":[
            "{Fine-tuning log 1}",
            "{Fine-tuning log 2}",
            ...
            ...
            ...
        ]
    }
}

Query and publish model checkpoints

Only SFT fine-tuning (efficient_sft and sft) supports saving and publishing checkpoints from intermediate training states.

List checkpoints for a fine-tuning job

curl 'https://dashscope.aliyuncs.com/api/v1/fine-tunes/<job_id>/checkpoints' \
--header 'Authorization: Bearer '${DASHSCOPE_API_KEY} \
--header 'Content-Type: application/json'

Request parameters

Parameter

Type

Parameter location

Required

Description

job_id

String

Path Parameter

Yes

The ID of the fine-tuning job.

Sample response

Note

The checkpoint field contains the checkpoint ID, which specifies the checkpoint to publish in the Model publishing (optional) API. The model_name field contains the model ID used for model deployment. The finetuned_output field in the original fine-tuning job response is the model_name of the final checkpoint.

{
    "request_id": "c11939b5-efa6-4639-97ae-ed4597984647",
    "output":
    [
        {
            "create_time": "2025-11-11T16:25:42",
            "full_name": "ft-202511272033-8ae7-checkpoint-20",
            "job_id": "ft-202511272033-8ae7",
            "checkpoint": "checkpoint-20",
            "model_name": "qwen3-8b-instruct-ft-202511272033-8ae7",
            "status": "SUCCEEDED"
        }
    ]
}

Status

Description

PENDING

The checkpoint is pending publication. You must publish it using the Model publishing API before you can use it for model deployment and invocation.

PROCESSING

The checkpoint is being published.

SUCCEEDED

The checkpoint has been published successfully. You can now use it for model deployment and invocation.

FAILED

The checkpoint failed to publish.

Model publishing (optional)

Note

In Model Studio, after a fine-tuning job completes, you must export a checkpoint before you can deploy the model.

Exported checkpoints are stored in cloud storage. You cannot access or download them at this time.

curl --request GET 'https://dashscope.aliyuncs.com/api/v1/fine-tunes/<job_id>/export/<checkpoint_id>?model_name=<model_name>' \
--header 'Authorization: Bearer '${DASHSCOPE_API_KEY} \
--header 'Content-Type: application/json'

Request parameters

Parameter

Type

Parameter location

Required

Description

job_id

String

Path Parameter

Yes

The ID of the fine-tuning job.

checkpoint_id

String

Path Parameter

Yes

The ID of the checkpoint to publish.

model_name

String

Path Parameter

Yes

The custom model ID to assign to the published model.

Sample response

{
    "request_id": "ed3faa41-6be3-4271-9b83-941b23680537",
    "output": true
}

The publishing task is asynchronous. Use the List checkpoints for a fine-tuning job API to monitor the publishing status of the checkpoint.

More fine-tuning operations

List fine-tuning jobs

curl 'https://dashscope.aliyuncs.com/api/v1/fine-tunes' \
--header 'Authorization: Bearer '${DASHSCOPE_API_KEY} \
--header 'Content-Type: application/json' 

Cancel a fine-tuning job

Cancels a running fine-tuning job.
curl --request POST 'https://dashscope.aliyuncs.com/api/v1/fine-tunes/<job_id>/cancel' \
--header 'Authorization: Bearer '${DASHSCOPE_API_KEY} \
--header 'Content-Type: application/json' 

Delete a fine-tuning job

You cannot delete a running fine-tuning job.
curl --request DELETE 'https://dashscope.aliyuncs.com/api/v1/fine-tunes/<job_id>' \
--header 'Authorization: Bearer '${DASHSCOPE_API_KEY} \
--header 'Content-Type: application/json' 

Model deployment and invocation

Model deployment

To deploy the model, go to the model deployment console.

Model invocation

Once the model deployment status is RUNNING, you can invoke the fine-tuned model just like any other model.

You can also get the Model Code from the model deployment console.

For details on usage and parameters, see the DashScope API Reference.

curl 'https://dashscope.aliyuncs.com/api/v1/services/aigc/text-generation/generation' \
--header 'Authorization: Bearer '${DASHSCOPE_API_KEY}  \
--header 'Content-Type: application/json' \
--data '{
    "model": "<your_model_instance_id>",
    "input":{
        "messages":[
            {
                "role": "user",
                "content": "Who are you?"
            }
        ]
    },
    "parameters": {
        "result_format": "message"
    }
}'