All Products
Search
Document Center

SuperApp:Quick Deploy Platform Models

Last Updated:Jun 08, 2026

Deploy models from Model Gallery, upload your own models, or use existing fine-tuned models with full configuration control.

Step 1: Select a model

Select a base model or fine-tuned model from one of the available sources.

Model sources

Choose from pre-trained models in the model gallery or models you have already uploaded or fine-tuned on the platform.

Available gallery models: Qwen, GPT, DeepSeek, GLM, internlm, and more.

Browse Model Gallery

Model and display name configuration

Model type

  • Base models: Choose a base model that fits your use case. Review its capabilities, limitations, and specifications — including parameter count and context window size — before deploying.

  • Fine-tuned models: Custom models trained on the platform on top of a base model. Use these when you need task-specific behavior that a base model does not provide out of the box.

Display name

A label to identify this deployment on the dashboard. Must be 64 characters or fewer.

image

Step 2: Configure resources

Choose the compute resources and deployment settings for your model.

Contact for a discount >

Resource configuration

Region

Choose a region close to your users to minimize latency.

Available regions: Singapore, Japan, USA-East, Indonesia, Frankfurt, Hong Kong, Malaysia

GPU type

Select a GPU type that meets your model's compute requirements. Different GPU types vary in memory capacity and throughput — larger models typically require GPUs with more VRAM.

Options: NVIDIA A10, L20, and others.

If the GPU type you need is unavailable in your preferred region, try a different region or select a comparable GPU type with equivalent memory and compute performance.

Replicas

Set the number of replicas to run for this deployment.

  • Minimum replicas: Keep at least one replica running at all times to avoid cold-start latency on incoming requests. Setting this to 0 reduces idle costs but causes a startup delay when the first request arrives.

  • Maximum replicas: Controls how many replicas the platform can scale up to under load, which caps your maximum cost.

Use multiple replicas to distribute concurrent requests across instances for load balancing and high availability.

image

Step 3: Review and deploy

Review your configuration and estimated costs before creating the deployment.

Cost summary

  • GPU compute cost

  • System overhead (currently $0)

Note: Model download and storage costs are not included in the cost summary.

Troubleshooting

1. Error 403: Sales of this resource are temporarily suspended

403: Sales of this resource are temporarily suspended.
  • Reason: The selected GPU type is temporarily unavailable in the current region due to high demand and insufficient resources.

  • Solution:

    1. Switch to a different region where the same GPU type is available.

    2. If the GPU type is unavailable across all regions, select a comparable GPU type with similar memory and compute specifications and retry. If the issue persists, contact our support team for assistance.

2. Error: Account has an outstanding balance

Account has an outstanding balance.
  • Reason: Your account balance is insufficient to cover the cost of this deployment.

  • Solution:

    Go to the Billing section of your account dashboard and add funds to your balance.

3. Error: Your account information is incomplete

Your account information is incomplete.
  • Reason: Your account has not completed the required identity or information verification process.

  • Solution:

    Go to the Alibaba Cloud Account Settings page and complete identity verification.