All Products
Search
Document Center

Platform For AI:Submit commands

Last Updated:Aug 24, 2026

You can use the client tool to submit multiple types of training jobs. This topic describes the commands used to submit jobs, including invocation syntax, parameter descriptions, and usage examples.

Common parameters for job submission

When you use the DLC command line to submit TensorFlow (tfjob), PyTorch (pytorchjob), or XGBoost (xgboostjob) jobs, a set of common parameters applies. The common parameters are listed below.

Table 1. Common parameters for job submission

Parameter

Required

Description

Type

Supported in job parameter file

name

Yes

The name of the job. Multiple jobs can share the same name.

STRING

Yes

command

Yes

The startup command for each node.

STRING

Yes

data_sources

No

Mount a dataset. You can specify the dataset version and mount path in the format --data_sources=<datasets_id>:<version>:<mount_path>, where:

  • <datasets_id>: Replace with the dataset ID. You can view it on the Datasets page.

  • <version>: Replace with the dataset version, for example, v1.

  • <mount_path>: Replace with the mount path, for example, /mnt/data/.

You can bind multiple datasets at the same time, separated by commas (,). The default value is empty.

STRING

Yes

spot_spec

No

Use spot resources. The format is <limit_type>:<limit_value>, where:

  • <limit_type>: The bid cap type. Two values are supported:

    • discount: Maximum discount

    • price: Maximum price

  • <limit_value>: The maximum price or maximum discount.

For example, to set a maximum bid at 50% off, configure:

--spot_spec=discount:0.5

Spot instance types are configured through the role-specific parameters of the job framework (for example, for PyTorchJob, use --master_spec and --worker_spec).

STRING

Yes

elastic_spot_specs

No

Use elastic spot resources with multiple instance types. The format is

<instance_type>:<limit_type>:<limit_value>, where:

  • <instance_type>: The instance type. You can view available types in the console. Supported spot instance types vary by region.

  • <limit_type>: The bid cap type. Two values are supported:

    • discount: Maximum discount

    • price: Maximum price

  • <limit_value>: The maximum price or maximum discount.

You can specify multiple candidate instance types separated by ,. The job runs on the first instance type that is successfully bid for. For example:

--elastic_spot_specs=ml.gp7vf.16.40xlarge:discount:1,ml.gu8tef.8.46xlarge:discount:0.8,ml.gu8tf.8.40xlarge:discount:0.5

STRING

Yes

data_source_uris

No

Directly mount a data source path. The format is --data_source_uris=<uri>::<mount_path>, where:

  • <uri>: Replace with the data source storage path, for example, oss://examplebucket.oss-cn-hangzhou-internal.aliyuncs.com/path/.

  • <mount_path>: Replace with the mount path, for example, /mnt/data/.

You can bind multiple data source paths at the same time, separated by commas (,). The default value is empty.

STRING

Yes

code_source

No

The code source ID. You can view it on the Code configuration page. Only a single value can be specified. The default value is empty.

STRING

Yes

code_branch

No

Specifies the branch of the code repository. Use this parameter together with the code_source parameter.

STRING

Yes

code_commit

No

Specifies the commit ID of the code repository. Use this parameter together with the code_source parameter.

STRING

Yes

thirdparty_libs

No

Python third-party libraries. Separate multiple libraries with commas (,). The default value is empty.

STRING

Yes

thirdparty_lib_dir

No

The directory that contains the requirements.txt file used for installing Python third-party libraries. The default value is empty.

STRING

No

vpc_id

No

The ID of the VPC (Virtual Private Cloud) that the job can access. The default value is empty.

STRING

Yes

switch_id

No (required if vpc_id is specified)

The ID of the vSwitch in the VPC that the job accesses. The default value is empty.

STRING

Yes

security_group_id

No (required if vpc_id is specified)

The ID of the security group in the VPC that the job accesses. The default value is empty.

STRING

Yes

job_file

No

The job parameter file. Parameters in the job_file take precedence. The format is key=value, where the key name matches the command line parameter name. For configuration details, see Advanced parameters for job submission.

STRING

No

interactive

No

Whether to start the job in interactive mode.

BOOL

Yes

job_max_running_time_minutes

No

The maximum running time of the job. The default value is 0, which means no maximum running time is set.

INT64

Yes

success_policy

No

Supported only by TFJob. Valid values:

  • ChiefWorker: The job succeeds as soon as the Pod of the Chief node completes successfully.

  • AllWorkers: The job succeeds only when all nodes complete successfully.

If left empty, AllWorkers is used by default.

STRING

Yes

envs

No

Configure environment variables for the worker nodes. The format is key1=value1,key2=value2. Multiple environment variables are separated by commas (,). The key and value of each environment variable are separated by an equals sign (=).

StringToString

Yes

tags

No

Configure tags for the job. The format is key1=value1,key2=value2. Multiple tags are separated by commas (,). The key and value of each tag are separated by an equals sign (=).

StringToString

Yes

oversold_type

No

Configure how the job uses idle compute resources. Valid values:

  • AcceptQuotaOverSold (acceptable): The job can use idle compute resources.

  • ForceQuotaOverSold (idle only): The job uses only idle compute resources.

  • ForbiddenQuotaOverSold (not acceptable): The job uses only resources within the associated quota and does not use idle compute resources.

STRING

Yes

driver

No

Specifies the GPU driver version used by the job.

STRING

Yes

default_route

No

When a VPC is selected, configure the method for accessing the public network. Valid values:

  • eth0 (default): Access the public network through the public gateway.

  • eth1: Access the public network through the private gateway of the selected VPC.

STRING

Yes

priority

No

Configure the priority of the job. The default value is 1. The value range is 1 to 9, where:

  • 1 is the lowest priority.

  • 9 is the highest priority.

INT32

Yes

exit_code_on_stopped

No

When a job running in interactive mode is stopped, specifies the exit code of the command line tool. The default value is 0.

INT32

Yes

job_reserved_minutes

No

Set the retention period (in minutes) after the job ends. The default value is 0.

INT32

Yes

job_reserved_policy

No

Set the retention policy for the job. Valid values:

  • Always (default): Retain the job whether it succeeds or fails.

  • OnFailure: Retain the job only when it fails.

  • OnSucceed: Retain the job only when it succeeds.

STRING

Yes

Submit TensorFlow training jobs (submit tfjob)

  • Function

    Used to submit TensorFlow training jobs.

  • Syntax

    You can submit TensorFlow jobs using either command line parameters or a job parameter file.

    ./dlc submit tfjob [flags]
  • Parameters

    When submitting a TensorFlow job using command line parameters, replace the parameters in the command with actual values. When submitting using a job parameter file, write the parameters in the file in the format <parameterName>=<parameterValue>. The common parameters for submitting TensorFlow jobs are listed at the beginning of this topic. The following are parameters specific to TensorFlow jobs:

    Table 2. Parameters specific to TensorFlow jobs

    Parameter

    Required

    Description

    Type

    Supported in job parameter file

    workspace_id

    Yes

    The workspace ID. The default value is empty. For information about how to create a workspace, see Create and manage a workspace.

    STRING

    Yes

    chief

    No

    Whether to enable the TensorFlow Chief node. Valid values:

    • false: Default. Disables the TensorFlow Chief node.

    • true: Enables the TensorFlow Chief node.

    BOOL

    Yes

    chief_image

    No

    The image of the TensorFlow Chief node. The default value is empty.

    STRING

    Yes

    chief_spec

    No

    The instance type used by the TensorFlow Chief node. The default value is empty.

    STRING

    Yes

    master_image

    No

    The image of the TensorFlow Master node. The default value is empty.

    STRING

    Yes

    master_spec

    No

    The instance type used by the TensorFlow Master node.

    STRING

    Yes

    masters

    No

    The number of TensorFlow Master nodes. The default value is 0.

    INT

    Yes

    ps

    No

    The number of TensorFlow Parameter Server nodes. The default value is 0.

    INT

    Yes

    ps_image

    No

    The image of the TensorFlow Parameter Server node. The default value is empty.

    STRING

    Yes

    ps_spec

    No

    The instance type used by the TensorFlow Parameter Server node. The default value is empty.

    STRING

    Yes

    worker_image

    No

    The image of the TensorFlow Worker node. The default value is empty.

    STRING

    Yes

    worker_spec

    No

    The instance type used by the TensorFlow Worker node. The default value is empty.

    STRING

    Yes

    workers

    No

    The number of TensorFlow Worker nodes. The default value is 0.

    INT

    Yes

    evaluator_image

    No

    The image of the TensorFlow Evaluator node. The default value is empty.

    STRING

    Yes

    evaluator_spec

    No

    The instance type used by the TensorFlow Evaluator node. The default value is empty.

    STRING

    Yes

    evaluators

    No

    The number of TensorFlow Evaluator nodes. The default value is 0.

    INT

    Yes

    graphlearn_image

    No

    The image of the TensorFlow GraphLearn node. The default value is empty.

    STRING

    Yes

    graphlearn_spec

    No

    The instance type used by the TensorFlow GraphLearn node. The default value is empty.

    STRING

    Yes

    graphlearns

    No

    The number of TensorFlow GraphLearn nodes. The default value is 0.

    INT

    Yes

    Table 3. Parameters specific to submitting TensorFlow jobs to a dedicated resource group

    Parameter

    Required

    Description

    Type

    Supported in job parameter file

    resource_id

    No (required when submitting jobs to a dedicated resource group)

    The ID of the dedicated resource quota. The default value is empty. For information about how to create a dedicated resource quota, see General computing resource quotas.

    STRING

    Yes

    priority

    No

    The job priority. The default value is 1.

    INT

    Yes

    chief_cpu

    No

    The number of CPUs used by the TensorFlow Chief node. The default value is empty.

    STRING

    Yes

    chief_gpu

    No

    The number of GPUs used by the TensorFlow Chief node. The default value is empty.

    STRING

    Yes

    chief_gpu_type

    No

    The type of GPU used by the TensorFlow Chief node. The default value is empty. Example: GU50.

    STRING

    Yes

    chief_memory

    No

    The memory resources used by the TensorFlow Chief node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    chief_shared_memory

    No

    The shared memory resources for the TensorFlow Chief node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    master_cpu

    No

    The number of CPUs used by the TensorFlow Master node. The default value is empty.

    STRING

    Yes

    master_gpu

    No

    The number of GPUs used by the TensorFlow Master node. The default value is empty.

    STRING

    Yes

    master_gpu_type

    No

    The type of GPU used by the TensorFlow Master node. The default value is empty. Example: GU50.

    STRING

    Yes

    master_memory

    No

    The memory resources used by the TensorFlow Master node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    master_shared_memory

    No

    The shared memory resources for the TensorFlow Master node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    *_cpu

    No

    The number of CPUs used by TensorFlow nodes. The default value is empty. Replace with (ps, worker, evaluator, graphlearn).

    STRING

    Yes

    *_gpu

    No

    The number of GPUs used by TensorFlow nodes. The default value is empty. Replace with (ps, worker, evaluator, graphlearn).

    STRING

    Yes

    *_gpu_type

    No

    The type of GPU used by TensorFlow nodes. The default value is empty. Example: GU50. Replace with (ps, worker, evaluator, graphlearn).

    STRING

    Yes

    *_memory

    No

    The memory resources used by TensorFlow nodes. The default value is empty. Example: 500 Mi, 1 Gi. Replace with (ps, worker, evaluator, graphlearn).

    STRING

    Yes

    *_shared_memory

    No

    The shared memory resources for TensorFlow nodes. The default value is empty. Example: 500 Mi, 1 Gi. Replace with (ps, worker, evaluator, graphlearn).

    STRING

    Yes

  • Examples

    • Submit a distributed job with 2 Workers and 1 PS using command line parameters. Example:

      ./dlc submit tfjob --name=test_2021 --ps=1 \
        --ps_spec=ecs.g6.8xlarge \
        --ps_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/tensorflow-training:1.12.2PAI-cpu-py27-ubuntu16.04 \
        --workers=2 \
        --worker_spec=ecs.g6.4xlarge \
        --worker_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/tensorflow-training:1.12.2PAI-cpu-py27-ubuntu16.04 \
        --command="python /root/data/dist_mnist/code/dist-main.py --max_steps=10000 --data_dir=/root/data/dist_mnist/data/" \
        --workspace_id=***** \
        --data_sources=d-pcsah1t86b********:v1:/mnt/data/

      The system returns a result similar to the following:

      +----------------------------------+--------------------------------------+
      |              JobId               |              RequestId               |
      +----------------------------------+--------------------------------------+
      | dlcmp6vwljkz****                 | xxxxxxxx-79AF-4EFC-9CE9-xxxxxxxxxxxx |
      +----------------------------------+--------------------------------------+
    • Submit a distributed job with 2 Workers and 1 PS using a job parameter file. Example:

      ./dlc submit tfjob --job_file=job_file.dist_mnist.1ps2w

      Here, job_file.dist_mnist.1ps2w is the job parameter file. Parameters are specified in the format <parameterName>=<parameterValue>. The content of job_file.dist_mnist.1ps2w is as follows:

      name=test_2021
      workers=2
      worker_spec=ecs.g6.4xlarge
      worker_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/tensorflow-training:1.12.2PAI-cpu-py27-ubuntu16.04
      ps=1
      ps_spec=ecs.g6.8xlarge
      ps_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/tensorflow-training:1.12.2PAI-cpu-py27-ubuntu16.04
      command=python /root/data/dist_mnist/code/dist-main.py --max_steps=10000 --data_dir=/root/data/dist_mnist/data/
      workspace_id=*****
      data_sources=d-pcsah1t86b********:v1:/mnt/data/

Submit PyTorch training jobs (submit pytorchjob)

  • Function

    Used to submit PyTorch training jobs.

  • Syntax

    You can submit PyTorch jobs using either command line parameters or a job parameter file.

    ./dlc submit pytorchjob [flags]
  • Parameters

    When submitting a PyTorch job using command line parameters, replace the parameters in the command with actual values. When submitting using a job parameter file, write the parameters in the file in the format <parameterName>=<parameterValue>. The common parameters for submitting PyTorch jobs are listed at the beginning of this topic. The following are parameters specific to PyTorch jobs:

    Table 4. Parameters specific to PyTorch jobs

    Parameter

    Required

    Description

    Type

    Supported in job parameter file

    workspace_id

    Yes

    The workspace ID. The default value is empty. For information about how to create a workspace, see Create and manage a workspace.

    STRING

    Yes

    master_image

    No

    The image of the PyTorch Master node. The default value is empty.

    STRING

    Yes

    master_spec

    No

    The instance type used by the PyTorch Master node. The default value is empty.

    STRING

    Yes

    masters

    No

    The number of PyTorch Master nodes. The default value is 0.

    INT

    Yes

    worker_image

    No

    The image of the PyTorch Worker node. The default value is empty.

    STRING

    Yes

    worker_spec

    No

    The instance type used by the PyTorch Worker node. The default value is empty.

    STRING

    Yes

    workers

    No

    The number of PyTorch Worker nodes. The default value is 0.

    INT

    Yes

    Table 5. Parameters specific to submitting PyTorch jobs to a dedicated resource group

    Parameter

    Required

    Description

    Type

    Supported in job parameter file

    resource_id

    No (required when submitting jobs to a dedicated resource group)

    The ID of the dedicated resource quota. The default value is empty. For information about how to create a dedicated resource quota, see General computing resource quotas.

    STRING

    Yes

    priority

    No

    The job priority. The default value is 1.

    INT

    Yes

    master_cpu

    No

    The number of CPUs used by the PyTorch Master node. The default value is empty.

    STRING

    Yes

    master_gpu

    No

    The number of GPUs used by the PyTorch Master node. The default value is empty.

    STRING

    Yes

    master_gpu_type

    No

    The type of GPU used by the PyTorch Master node. The default value is empty. Example: GU50.

    STRING

    Yes

    master_memory

    No

    The memory resources used by the PyTorch Master node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    master_shared_memory

    No

    The shared memory resources for the PyTorch Master node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    worker_cpu

    No

    The number of CPUs used by the PyTorch Worker node. The default value is empty.

    STRING

    Yes

    worker_gpu

    No

    The number of GPUs used by the PyTorch Worker node. The default value is empty.

    STRING

    Yes

    worker_gpu_type

    No

    The type of GPU used by the PyTorch Worker node. The default value is empty. Example: GU50.

    STRING

    Yes

    worker_memory

    No

    The memory resources used by the PyTorch Worker node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    worker_shared_memory

    No

    The shared memory resources for the PyTorch Worker node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

  • Examples

    Submit a GPU model training job using command line parameters. Example:

    ./dlc submit pytorchjob --name=test_pt_face \
      --workers=1 \
      --worker_spec=ecs.gn6e-c12g1.3xlarge \
      --worker_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/pytorch-training:1.7.1-gpu-py37-cu110-ubuntu18.04 \
      --command="apt-get update; apt-get -y --allow-downgrades install libpcre3=2:8.38-3.1 libpcre3-dev libgl1-mesa-glx libglib2.0-dev; cd /root/data/face; python train.py --num_workers 0 --save_folder outputs" \
      --data_sources=d-pcsah1t86b********:v1:/mnt/data/ \
      --workspace_id=*****

    Submit a model training job that uses spot resources using command line parameters. Example:

    ./dlc submit pytorchjob --name=test_pt_face \
      --workers=1 \
      --elastic_spot_specs=ml.gp7vf.16.40xlarge:discount:1,ml.gu8tef.8.46xlarge:discount:1,ml.gu8tf.8.40xlarge:discount:0.5 \
      --worker_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/pytorch-training:1.7.1-gpu-py37-cu110-ubuntu18.04 \
      --command="apt-get update; apt-get -y --allow-downgrades install libpcre3=2:8.38-3.1 libpcre3-dev libgl1-mesa-glx libglib2.0-dev; cd /root/data/face; python train.py --num_workers 0 --save_folder outputs" \
      --workspace_id=*****

    The system returns a result similar to the following:

    +----------------------------------+--------------------------------------+
    |              JobId               |              RequestId               |
    +----------------------------------+--------------------------------------+
    | dlcu704xxuxk****                 | xxxxxxxx-79AF-4EFC-9CE9-xxxxxxxxxxxx |
    +----------------------------------+--------------------------------------+

Submit XGBoost training jobs (submit xgboostjob)

  • Function

    Used to submit XGBoost training jobs.

  • Syntax

    You can submit XGBoost jobs using either command line parameters or a job parameter file.

    ./dlc submit xgboostjob [flags]
  • Parameters

    When submitting an XGBoost job using command line parameters, replace the parameters in the command with actual values. When submitting using a job parameter file, write the parameters in the file in the format <parameterName>=<parameterValue>. The common parameters for submitting XGBoost jobs are listed at the beginning of this topic. The following are parameters specific to XGBoost jobs:

    Table 6. Parameters specific to XGBoost jobs

    Parameter

    Required

    Description

    Type

    Supported in job parameter file

    workspace_id

    Yes

    The workspace ID. The default value is empty. For information about how to create a workspace, see Create and manage a workspace.

    STRING

    Yes

    master_image

    No

    The image of the XGBoost Master node. The default value is empty.

    STRING

    Yes

    master_spec

    No

    The instance type used by the XGBoost Master node. The default value is empty.

    STRING

    Yes

    masters

    No

    The number of XGBoost Master nodes. The default value is 0.

    INT

    Yes

    worker_image

    No

    The image of the XGBoost Worker node. The default value is empty.

    STRING

    Yes

    worker_spec

    No

    The instance type used by the XGBoost Worker node. The default value is empty.

    STRING

    Yes

    workers

    No

    The number of XGBoost Worker nodes. The default value is 0.

    INT

    Yes

    Table 7. Parameters specific to submitting XGBoost jobs to a dedicated resource group

    Parameter

    Required

    Description

    Type

    Supported in job parameter file

    resource_id

    No (required when submitting jobs to a dedicated resource group)

    The ID of the dedicated resource quota. The default value is empty. For information about how to create a dedicated resource quota, see General computing resource quotas.

    STRING

    Yes

    priority

    No

    The job priority. The default value is 1.

    INT

    Yes

    master_cpu

    No

    The number of CPUs used by the XGBoost Master node. The default value is empty.

    STRING

    Yes

    master_gpu

    No

    The number of GPUs used by the XGBoost Master node. The default value is empty.

    STRING

    Yes

    master_gpu_type

    No

    The type of GPU used by the XGBoost Master node. The default value is empty. Example: GU50.

    STRING

    Yes

    master_memory

    No

    The memory resources used by the XGBoost Master node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    master_shared_memory

    No

    The shared memory resources for the XGBoost Master node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    worker_cpu

    No

    The number of CPUs used by the XGBoost Worker node. The default value is empty.

    STRING

    Yes

    worker_gpu

    No

    The number of GPUs used by the XGBoost Worker node. The default value is empty.

    STRING

    Yes

    worker_gpu_type

    No

    The type of GPU used by the XGBoost Worker node. The default value is empty. Example: GU50.

    STRING

    Yes

    worker_memory

    No

    The memory resources used by the XGBoost Worker node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

    worker_shared_memory

    No

    The shared memory resources for the XGBoost Worker node. The default value is empty. Example: 500 Mi, 1 Gi.

    STRING

    Yes

  • Examples

    Submit an XGBoost job using command line parameters. Example:

    ./dlc submit xgboostjob --name=test_xgboost \
      --workers=1 \
      --worker_spec=ecs.gn6e-c12g1.3xlarge \
      --worker_image=xgboost-training:1.6.0-cpu-py36-ubuntu18.04 \
      --command="python /root/code/horovod/xgboost/main.py --job_type=Train --xgboost_parameter=objective:multi:softprob,num_class:3 --n_estimators=50 --model_path=autoAI/xgb-opt/2" \
      --data_sources=d-pcsah1t86b********:v1:/mnt/data/ \
      --workspace_id=*****

    The system returns a result similar to the following:

    +----------------------------------+--------------------------------------+
    |              JobId               |              RequestId               |
    +----------------------------------+--------------------------------------+
    | dlc1nvu3gli0****                 | xxxxxxxx-79AF-4EFC-9CE9-xxxxxxxxxxxx |
    +----------------------------------+--------------------------------------+

Advanced parameters for job submission

Schedule jobs to specific nodes

When you submit a job using a Lingjun intelligent computing or general computing resource quota, you can schedule the job to specific nodes by configuring parameters in the DLC command line.

Note

This feature is currently available only to whitelisted users. If you need access, contact your account manager to be added to the whitelist.

  • Parameters

    Parameter

    Description

    Example value

    --allow_nodes="${allow_nodes}"

    Specifies a list of node names. Separate multiple node names with commas (,). Do not include spaces between node names.

    lingjuc47iextvg9-,lingjuc47iextvg9-

    --deny_nodes="${deny_nodes}"

    Specifies a list of nodes to exclude. Separate multiple node names with commas (,). Do not include spaces between node names.

    lingjuc47iextvg9-,lingjuc47iextvg9-

  • Examples

    Command line parameters

    Submit a job using command line parameters. Example:

    • No scheduling node specified

      ./dlc submit pytorchjob --name=assign_node_test_no_node  \--workers=1 \
          --worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 \
          --command="sleep 1000" \
          --workspace_id='****' \
          --resource_id='quotau2h98mt****' \
          --worker_cpu="1" \
          --worker_memory='2Gi'  
    • Specify scheduling nodes

      ./dlc submit pytorchjob --name=assign_node_test_2_allow_nodes  \--workers=1 \
          --worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 \
          --command="sleep 1000" \
          --workspace_id='****' \
          --resource_id='quotau2h98mt****' \
          --worker_cpu="1" \
          --worker_memory='2Gi' \
          --allow_nodes="lingjuc47iextvg9-****,lingjuc47iextvg9-****" 
    • Exclude specific nodes

       ./dlc submit pytorchjob --name=assign_node_test_two_deny_nodes  \--workers=1 \
          --worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 \
          --command="sleep 1000" \
          --workspace_id='****' \
          --resource_id='quotau2h98mt****' \
          --worker_cpu="1" \
          --worker_memory='2Gi' \
          --deny_nodes="lingjuc47iextvg9-****,lingjuc47iextvg9-****"
    • Specify scheduling nodes & exclude specific nodes

      ./dlc submit pytorchjob --name=assign_node_test_two_allow_two_deny  \--workers=1 \
          --worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 \
          --command="sleep 1000" \
          --workspace_id='****' \
          --resource_id='quotau2h98mt****' \
          --worker_cpu="1" \
          --worker_memory='2Gi' \
          --allow_nodes="lingjuc47iextvg9-****,lingjuc47iextvg9-****" \
          --deny_nodes="lingjuc47iextvg9-****,lingjuc47iextvg9-****"

    File-based configuration

    Submit a job using a file-based configuration. Example:

    ./dlc submit pytorchjob -f job_file

    Here, job_file is the job parameter file. The content is as follows:

    • No scheduling node specified

      name=assign_node_test_no_node
      workers=1
      worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04
      command=sleep 1000
      workspace_id=****
      resource_id=quotau2h98mt****
      worker_cpu=1
      worker_memory=2Gi
      
    • Specify scheduling nodes

      name=assign_node_test_2_allow_nodes
      workers=1
      worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04
      command=sleep 1000
      workspace_id=****
      resource_id=quotau2h98mt****
      worker_cpu=1
      worker_memory=2Gi
      allow_nodes=lingjuc47iextvg9-****,lingjuc47iextvg9-****
      
    • Exclude specific nodes

      name=assign_node_test_two_allow_two_deny
      workers=1
      worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04
      command=sleep 1000
      workspace_id=****
      resource_id=quotau2h98mt****
      worker_cpu=1
      worker_memory=2Gi
      deny_nodes=lingjuc47iextvg9-****,lingjuc47iextvg9-****
      
    • Specify scheduling nodes & exclude specific nodes

      name=assign_node_test_two_allow_two_deny
      workers=1
      worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04
      command=sleep 1000
      workspace_id=****
      resource_id=quotau2h98mt****
      worker_cpu=1
      worker_memory=2Gi
      allow_nodes=lingjuc47iextvg9-****,lingjuc47iextvg9-****
      deny_nodes=lingjuc47iextvg9-****,lingjuc47iextvg9-****
      

Disable pay-as-you-go stock check for job submission

You can configure the disable_ecs_stock_check parameter in the DLC command line to disable the pay-as-you-go stock check.

  • Parameters

    Parameter

    Description

    Example value

    disable_ecs_stock_check

    Disables the pay-as-you-go stock check. Valid values:

    • false (default): Enables the pay-as-you-go stock check.

    • true: Disables the pay-as-you-go stock check.

    true or false

  • Examples

    Command line parameters

    Submit a job using command line parameters. Example:

    • Enable pay-as-you-go stock check

      ./dlc submit pytorchjob \
          --name=test_skip_checking3 \
          --command='sleep 1000' \
          --workspace_id=**** \
          --priority=1 \
          --workers=1 \
          --worker_image=registry.cn-hangzhou.aliyuncs.com/pai-dlc/tensorflow-training:1.12PAI-gpu-py36-cu101-ubuntu18.04 \
          --worker_spec=ecs.g6.xlarge  
    • Disable pay-as-you-go stock check

      ./dlc submit pytorchjob \
          --name=test_skip_checking3 \
          --command='sleep 1000' \
          --workspace_id=**** \
          --priority=1 \
          --workers=1 \
          --worker_image=registry.cn-hangzhou.aliyuncs.com/pai-dlc/tensorflow-training:1.12PAI-gpu-py36-cu101-ubuntu18.04 \
          --worker_spec=ecs.g6.xlarge \
          --disable_ecs_stock_check=true
       

    File-based configuration

    Submit a job using a file-based configuration. Example:

    ./dlc submit pytorchjob -f job_file

    Here, job_file is the job parameter file. The content is as follows:

    • Enable pay-as-you-go stock check

      name=test_skip_checking3
      workers=1
      worker_image=registry.cn-hangzhou.aliyuncs.com/pai-dlc/tensorflow-training:1.12PAI-gpu-py36-cu101-ubuntu18.04
      command=sleep 1000
      workspace_id=****
      worker_spec=ecs.g6.xlarge
      
    • Disable pay-as-you-go stock check

      name=test_skip_checking3
      workers=1
      worker_image=registry.cn-hangzhou.aliyuncs.com/pai-dlc/tensorflow-training:1.12PAI-gpu-py36-cu101-ubuntu18.04
      command=sleep 1000
      workspace_id=****
      worker_spec=ecs.g6.xlarge
      disable_ecs_stock_check=true
      

Related documents