You can use the client tool to submit multiple types of training jobs. This topic describes the commands used to submit jobs, including invocation syntax, parameter descriptions, and usage examples.
Common parameters for job submission
When you use the DLC command line to submit TensorFlow (tfjob), PyTorch (pytorchjob), or XGBoost (xgboostjob) jobs, a set of common parameters applies. The common parameters are listed below.
Table 1. Common parameters for job submission
Parameter | Required | Description | Type | Supported in job parameter file |
name | Yes | The name of the job. Multiple jobs can share the same name. | STRING | Yes |
command | Yes | The startup command for each node. | STRING | Yes |
data_sources | No | Mount a dataset. You can specify the dataset version and mount path in the format
You can bind multiple datasets at the same time, separated by commas (,). The default value is empty. | STRING | Yes |
spot_spec | No | Use spot resources. The format is
For example, to set a maximum bid at 50% off, configure:
Spot instance types are configured through the role-specific parameters of the job framework (for example, for PyTorchJob, use --master_spec and --worker_spec). | STRING | Yes |
elastic_spot_specs | No | Use elastic spot resources with multiple instance types. The format is
You can specify multiple candidate instance types separated by
| STRING | Yes |
data_source_uris | No | Directly mount a data source path. The format is
You can bind multiple data source paths at the same time, separated by commas (,). The default value is empty. | STRING | Yes |
code_source | No | The code source ID. You can view it on the Code configuration page. Only a single value can be specified. The default value is empty. | STRING | Yes |
code_branch | No | Specifies the branch of the code repository. Use this parameter together with the code_source parameter. | STRING | Yes |
code_commit | No | Specifies the commit ID of the code repository. Use this parameter together with the code_source parameter. | STRING | Yes |
thirdparty_libs | No | Python third-party libraries. Separate multiple libraries with commas (,). The default value is empty. | STRING | Yes |
thirdparty_lib_dir | No | The directory that contains the requirements.txt file used for installing Python third-party libraries. The default value is empty. | STRING | No |
vpc_id | No | The ID of the VPC (Virtual Private Cloud) that the job can access. The default value is empty. | STRING | Yes |
switch_id | No (required if vpc_id is specified) | The ID of the vSwitch in the VPC that the job accesses. The default value is empty. | STRING | Yes |
security_group_id | No (required if vpc_id is specified) | The ID of the security group in the VPC that the job accesses. The default value is empty. | STRING | Yes |
job_file | No | The job parameter file. Parameters in the job_file take precedence. The format is | STRING | No |
interactive | No | Whether to start the job in interactive mode. | BOOL | Yes |
job_max_running_time_minutes | No | The maximum running time of the job. The default value is 0, which means no maximum running time is set. | INT64 | Yes |
success_policy | No | Supported only by TFJob. Valid values:
If left empty, AllWorkers is used by default. | STRING | Yes |
envs | No | Configure environment variables for the worker nodes. The format is | StringToString | Yes |
tags | No | Configure tags for the job. The format is | StringToString | Yes |
oversold_type | No | Configure how the job uses idle compute resources. Valid values:
| STRING | Yes |
driver | No | Specifies the GPU driver version used by the job. | STRING | Yes |
default_route | No | When a VPC is selected, configure the method for accessing the public network. Valid values:
| STRING | Yes |
priority | No | Configure the priority of the job. The default value is 1. The value range is 1 to 9, where:
| INT32 | Yes |
exit_code_on_stopped | No | When a job running in interactive mode is stopped, specifies the exit code of the command line tool. The default value is 0. | INT32 | Yes |
job_reserved_minutes | No | Set the retention period (in minutes) after the job ends. The default value is 0. | INT32 | Yes |
job_reserved_policy | No | Set the retention policy for the job. Valid values:
| STRING | Yes |
Submit TensorFlow training jobs (submit tfjob)
-
Function
Used to submit TensorFlow training jobs.
-
Syntax
You can submit TensorFlow jobs using either command line parameters or a job parameter file.
./dlc submit tfjob [flags] -
Parameters
When submitting a TensorFlow job using command line parameters, replace the parameters in the command with actual values. When submitting using a job parameter file, write the parameters in the file in the format
<parameterName>=<parameterValue>. The common parameters for submitting TensorFlow jobs are listed at the beginning of this topic. The following are parameters specific to TensorFlow jobs:Table 2. Parameters specific to TensorFlow jobs
Parameter
Required
Description
Type
Supported in job parameter file
workspace_id
Yes
The workspace ID. The default value is empty. For information about how to create a workspace, see Create and manage a workspace.
STRING
Yes
chief
No
Whether to enable the TensorFlow Chief node. Valid values:
false: Default. Disables the TensorFlow Chief node.
true: Enables the TensorFlow Chief node.
BOOL
Yes
chief_image
No
The image of the TensorFlow Chief node. The default value is empty.
STRING
Yes
chief_spec
No
The instance type used by the TensorFlow Chief node. The default value is empty.
STRING
Yes
master_image
No
The image of the TensorFlow Master node. The default value is empty.
STRING
Yes
master_spec
No
The instance type used by the TensorFlow Master node.
STRING
Yes
masters
No
The number of TensorFlow Master nodes. The default value is 0.
INT
Yes
ps
No
The number of TensorFlow Parameter Server nodes. The default value is 0.
INT
Yes
ps_image
No
The image of the TensorFlow Parameter Server node. The default value is empty.
STRING
Yes
ps_spec
No
The instance type used by the TensorFlow Parameter Server node. The default value is empty.
STRING
Yes
worker_image
No
The image of the TensorFlow Worker node. The default value is empty.
STRING
Yes
worker_spec
No
The instance type used by the TensorFlow Worker node. The default value is empty.
STRING
Yes
workers
No
The number of TensorFlow Worker nodes. The default value is 0.
INT
Yes
evaluator_image
No
The image of the TensorFlow Evaluator node. The default value is empty.
STRING
Yes
evaluator_spec
No
The instance type used by the TensorFlow Evaluator node. The default value is empty.
STRING
Yes
evaluators
No
The number of TensorFlow Evaluator nodes. The default value is 0.
INT
Yes
graphlearn_image
No
The image of the TensorFlow GraphLearn node. The default value is empty.
STRING
Yes
graphlearn_spec
No
The instance type used by the TensorFlow GraphLearn node. The default value is empty.
STRING
Yes
graphlearns
No
The number of TensorFlow GraphLearn nodes. The default value is 0.
INT
Yes
Table 3. Parameters specific to submitting TensorFlow jobs to a dedicated resource group
Parameter
Required
Description
Type
Supported in job parameter file
resource_id
No (required when submitting jobs to a dedicated resource group)
The ID of the dedicated resource quota. The default value is empty. For information about how to create a dedicated resource quota, see General computing resource quotas.
STRING
Yes
priority
No
The job priority. The default value is 1.
INT
Yes
chief_cpu
No
The number of CPUs used by the TensorFlow Chief node. The default value is empty.
STRING
Yes
chief_gpu
No
The number of GPUs used by the TensorFlow Chief node. The default value is empty.
STRING
Yes
chief_gpu_type
No
The type of GPU used by the TensorFlow Chief node. The default value is empty. Example: GU50.
STRING
Yes
chief_memory
No
The memory resources used by the TensorFlow Chief node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
chief_shared_memory
No
The shared memory resources for the TensorFlow Chief node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
master_cpu
No
The number of CPUs used by the TensorFlow Master node. The default value is empty.
STRING
Yes
master_gpu
No
The number of GPUs used by the TensorFlow Master node. The default value is empty.
STRING
Yes
master_gpu_type
No
The type of GPU used by the TensorFlow Master node. The default value is empty. Example: GU50.
STRING
Yes
master_memory
No
The memory resources used by the TensorFlow Master node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
master_shared_memory
No
The shared memory resources for the TensorFlow Master node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
*_cpu
No
The number of CPUs used by TensorFlow nodes. The default value is empty. Replace with (ps, worker, evaluator, graphlearn).
STRING
Yes
*_gpu
No
The number of GPUs used by TensorFlow nodes. The default value is empty. Replace with (ps, worker, evaluator, graphlearn).
STRING
Yes
*_gpu_type
No
The type of GPU used by TensorFlow nodes. The default value is empty. Example: GU50. Replace with (ps, worker, evaluator, graphlearn).
STRING
Yes
*_memory
No
The memory resources used by TensorFlow nodes. The default value is empty. Example: 500 Mi, 1 Gi. Replace with (ps, worker, evaluator, graphlearn).
STRING
Yes
*_shared_memory
No
The shared memory resources for TensorFlow nodes. The default value is empty. Example: 500 Mi, 1 Gi. Replace with (ps, worker, evaluator, graphlearn).
STRING
Yes
-
Examples
-
Submit a distributed job with 2 Workers and 1 PS using command line parameters. Example:
./dlc submit tfjob --name=test_2021 --ps=1 \ --ps_spec=ecs.g6.8xlarge \ --ps_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/tensorflow-training:1.12.2PAI-cpu-py27-ubuntu16.04 \ --workers=2 \ --worker_spec=ecs.g6.4xlarge \ --worker_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/tensorflow-training:1.12.2PAI-cpu-py27-ubuntu16.04 \ --command="python /root/data/dist_mnist/code/dist-main.py --max_steps=10000 --data_dir=/root/data/dist_mnist/data/" \ --workspace_id=***** \ --data_sources=d-pcsah1t86b********:v1:/mnt/data/The system returns a result similar to the following:
+----------------------------------+--------------------------------------+ | JobId | RequestId | +----------------------------------+--------------------------------------+ | dlcmp6vwljkz**** | xxxxxxxx-79AF-4EFC-9CE9-xxxxxxxxxxxx | +----------------------------------+--------------------------------------+ -
Submit a distributed job with 2 Workers and 1 PS using a job parameter file. Example:
./dlc submit tfjob --job_file=job_file.dist_mnist.1ps2wHere, job_file.dist_mnist.1ps2w is the job parameter file. Parameters are specified in the format
<parameterName>=<parameterValue>. The content of job_file.dist_mnist.1ps2w is as follows:name=test_2021 workers=2 worker_spec=ecs.g6.4xlarge worker_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/tensorflow-training:1.12.2PAI-cpu-py27-ubuntu16.04 ps=1 ps_spec=ecs.g6.8xlarge ps_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/tensorflow-training:1.12.2PAI-cpu-py27-ubuntu16.04 command=python /root/data/dist_mnist/code/dist-main.py --max_steps=10000 --data_dir=/root/data/dist_mnist/data/ workspace_id=***** data_sources=d-pcsah1t86b********:v1:/mnt/data/
-
Submit PyTorch training jobs (submit pytorchjob)
-
Function
Used to submit PyTorch training jobs.
-
Syntax
You can submit PyTorch jobs using either command line parameters or a job parameter file.
./dlc submit pytorchjob [flags] -
Parameters
When submitting a PyTorch job using command line parameters, replace the parameters in the command with actual values. When submitting using a job parameter file, write the parameters in the file in the format
<parameterName>=<parameterValue>. The common parameters for submitting PyTorch jobs are listed at the beginning of this topic. The following are parameters specific to PyTorch jobs:Table 4. Parameters specific to PyTorch jobs
Parameter
Required
Description
Type
Supported in job parameter file
workspace_id
Yes
The workspace ID. The default value is empty. For information about how to create a workspace, see Create and manage a workspace.
STRING
Yes
master_image
No
The image of the PyTorch Master node. The default value is empty.
STRING
Yes
master_spec
No
The instance type used by the PyTorch Master node. The default value is empty.
STRING
Yes
masters
No
The number of PyTorch Master nodes. The default value is 0.
INT
Yes
worker_image
No
The image of the PyTorch Worker node. The default value is empty.
STRING
Yes
worker_spec
No
The instance type used by the PyTorch Worker node. The default value is empty.
STRING
Yes
workers
No
The number of PyTorch Worker nodes. The default value is 0.
INT
Yes
Table 5. Parameters specific to submitting PyTorch jobs to a dedicated resource group
Parameter
Required
Description
Type
Supported in job parameter file
resource_id
No (required when submitting jobs to a dedicated resource group)
The ID of the dedicated resource quota. The default value is empty. For information about how to create a dedicated resource quota, see General computing resource quotas.
STRING
Yes
priority
No
The job priority. The default value is 1.
INT
Yes
master_cpu
No
The number of CPUs used by the PyTorch Master node. The default value is empty.
STRING
Yes
master_gpu
No
The number of GPUs used by the PyTorch Master node. The default value is empty.
STRING
Yes
master_gpu_type
No
The type of GPU used by the PyTorch Master node. The default value is empty. Example: GU50.
STRING
Yes
master_memory
No
The memory resources used by the PyTorch Master node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
master_shared_memory
No
The shared memory resources for the PyTorch Master node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
worker_cpu
No
The number of CPUs used by the PyTorch Worker node. The default value is empty.
STRING
Yes
worker_gpu
No
The number of GPUs used by the PyTorch Worker node. The default value is empty.
STRING
Yes
worker_gpu_type
No
The type of GPU used by the PyTorch Worker node. The default value is empty. Example: GU50.
STRING
Yes
worker_memory
No
The memory resources used by the PyTorch Worker node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
worker_shared_memory
No
The shared memory resources for the PyTorch Worker node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
-
Examples
Submit a GPU model training job using command line parameters. Example:
./dlc submit pytorchjob --name=test_pt_face \ --workers=1 \ --worker_spec=ecs.gn6e-c12g1.3xlarge \ --worker_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/pytorch-training:1.7.1-gpu-py37-cu110-ubuntu18.04 \ --command="apt-get update; apt-get -y --allow-downgrades install libpcre3=2:8.38-3.1 libpcre3-dev libgl1-mesa-glx libglib2.0-dev; cd /root/data/face; python train.py --num_workers 0 --save_folder outputs" \ --data_sources=d-pcsah1t86b********:v1:/mnt/data/ \ --workspace_id=*****Submit a model training job that uses spot resources using command line parameters. Example:
./dlc submit pytorchjob --name=test_pt_face \ --workers=1 \ --elastic_spot_specs=ml.gp7vf.16.40xlarge:discount:1,ml.gu8tef.8.46xlarge:discount:1,ml.gu8tf.8.40xlarge:discount:0.5 \ --worker_image=registry-vpc.cn-beijing.aliyuncs.com/pai-dlc/pytorch-training:1.7.1-gpu-py37-cu110-ubuntu18.04 \ --command="apt-get update; apt-get -y --allow-downgrades install libpcre3=2:8.38-3.1 libpcre3-dev libgl1-mesa-glx libglib2.0-dev; cd /root/data/face; python train.py --num_workers 0 --save_folder outputs" \ --workspace_id=*****The system returns a result similar to the following:
+----------------------------------+--------------------------------------+ | JobId | RequestId | +----------------------------------+--------------------------------------+ | dlcu704xxuxk**** | xxxxxxxx-79AF-4EFC-9CE9-xxxxxxxxxxxx | +----------------------------------+--------------------------------------+
Submit XGBoost training jobs (submit xgboostjob)
-
Function
Used to submit XGBoost training jobs.
-
Syntax
You can submit XGBoost jobs using either command line parameters or a job parameter file.
./dlc submit xgboostjob [flags] -
Parameters
When submitting an XGBoost job using command line parameters, replace the parameters in the command with actual values. When submitting using a job parameter file, write the parameters in the file in the format
<parameterName>=<parameterValue>. The common parameters for submitting XGBoost jobs are listed at the beginning of this topic. The following are parameters specific to XGBoost jobs:Table 6. Parameters specific to XGBoost jobs
Parameter
Required
Description
Type
Supported in job parameter file
workspace_id
Yes
The workspace ID. The default value is empty. For information about how to create a workspace, see Create and manage a workspace.
STRING
Yes
master_image
No
The image of the XGBoost Master node. The default value is empty.
STRING
Yes
master_spec
No
The instance type used by the XGBoost Master node. The default value is empty.
STRING
Yes
masters
No
The number of XGBoost Master nodes. The default value is 0.
INT
Yes
worker_image
No
The image of the XGBoost Worker node. The default value is empty.
STRING
Yes
worker_spec
No
The instance type used by the XGBoost Worker node. The default value is empty.
STRING
Yes
workers
No
The number of XGBoost Worker nodes. The default value is 0.
INT
Yes
Table 7. Parameters specific to submitting XGBoost jobs to a dedicated resource group
Parameter
Required
Description
Type
Supported in job parameter file
resource_id
No (required when submitting jobs to a dedicated resource group)
The ID of the dedicated resource quota. The default value is empty. For information about how to create a dedicated resource quota, see General computing resource quotas.
STRING
Yes
priority
No
The job priority. The default value is 1.
INT
Yes
master_cpu
No
The number of CPUs used by the XGBoost Master node. The default value is empty.
STRING
Yes
master_gpu
No
The number of GPUs used by the XGBoost Master node. The default value is empty.
STRING
Yes
master_gpu_type
No
The type of GPU used by the XGBoost Master node. The default value is empty. Example: GU50.
STRING
Yes
master_memory
No
The memory resources used by the XGBoost Master node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
master_shared_memory
No
The shared memory resources for the XGBoost Master node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
worker_cpu
No
The number of CPUs used by the XGBoost Worker node. The default value is empty.
STRING
Yes
worker_gpu
No
The number of GPUs used by the XGBoost Worker node. The default value is empty.
STRING
Yes
worker_gpu_type
No
The type of GPU used by the XGBoost Worker node. The default value is empty. Example: GU50.
STRING
Yes
worker_memory
No
The memory resources used by the XGBoost Worker node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
worker_shared_memory
No
The shared memory resources for the XGBoost Worker node. The default value is empty. Example: 500 Mi, 1 Gi.
STRING
Yes
-
Examples
Submit an XGBoost job using command line parameters. Example:
./dlc submit xgboostjob --name=test_xgboost \ --workers=1 \ --worker_spec=ecs.gn6e-c12g1.3xlarge \ --worker_image=xgboost-training:1.6.0-cpu-py36-ubuntu18.04 \ --command="python /root/code/horovod/xgboost/main.py --job_type=Train --xgboost_parameter=objective:multi:softprob,num_class:3 --n_estimators=50 --model_path=autoAI/xgb-opt/2" \ --data_sources=d-pcsah1t86b********:v1:/mnt/data/ \ --workspace_id=*****The system returns a result similar to the following:
+----------------------------------+--------------------------------------+ | JobId | RequestId | +----------------------------------+--------------------------------------+ | dlc1nvu3gli0**** | xxxxxxxx-79AF-4EFC-9CE9-xxxxxxxxxxxx | +----------------------------------+--------------------------------------+
Advanced parameters for job submission
Schedule jobs to specific nodes
When you submit a job using a Lingjun intelligent computing or general computing resource quota, you can schedule the job to specific nodes by configuring parameters in the DLC command line.
This feature is currently available only to whitelisted users. If you need access, contact your account manager to be added to the whitelist.
-
Parameters
Parameter
Description
Example value
--allow_nodes="${allow_nodes}"
Specifies a list of node names. Separate multiple node names with commas (,). Do not include spaces between node names.
lingjuc47iextvg9-,lingjuc47iextvg9-
--deny_nodes="${deny_nodes}"
Specifies a list of nodes to exclude. Separate multiple node names with commas (,). Do not include spaces between node names.
lingjuc47iextvg9-,lingjuc47iextvg9-
-
Examples
Command line parameters
Submit a job using command line parameters. Example:
-
No scheduling node specified
./dlc submit pytorchjob --name=assign_node_test_no_node \--workers=1 \ --worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 \ --command="sleep 1000" \ --workspace_id='****' \ --resource_id='quotau2h98mt****' \ --worker_cpu="1" \ --worker_memory='2Gi' -
Specify scheduling nodes
./dlc submit pytorchjob --name=assign_node_test_2_allow_nodes \--workers=1 \ --worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 \ --command="sleep 1000" \ --workspace_id='****' \ --resource_id='quotau2h98mt****' \ --worker_cpu="1" \ --worker_memory='2Gi' \ --allow_nodes="lingjuc47iextvg9-****,lingjuc47iextvg9-****" -
Exclude specific nodes
./dlc submit pytorchjob --name=assign_node_test_two_deny_nodes \--workers=1 \ --worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 \ --command="sleep 1000" \ --workspace_id='****' \ --resource_id='quotau2h98mt****' \ --worker_cpu="1" \ --worker_memory='2Gi' \ --deny_nodes="lingjuc47iextvg9-****,lingjuc47iextvg9-****" -
Specify scheduling nodes & exclude specific nodes
./dlc submit pytorchjob --name=assign_node_test_two_allow_two_deny \--workers=1 \ --worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 \ --command="sleep 1000" \ --workspace_id='****' \ --resource_id='quotau2h98mt****' \ --worker_cpu="1" \ --worker_memory='2Gi' \ --allow_nodes="lingjuc47iextvg9-****,lingjuc47iextvg9-****" \ --deny_nodes="lingjuc47iextvg9-****,lingjuc47iextvg9-****"
File-based configuration
Submit a job using a file-based configuration. Example:
./dlc submit pytorchjob -f job_fileHere, job_file is the job parameter file. The content is as follows:
-
No scheduling node specified
name=assign_node_test_no_node workers=1 worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 command=sleep 1000 workspace_id=**** resource_id=quotau2h98mt**** worker_cpu=1 worker_memory=2Gi -
Specify scheduling nodes
name=assign_node_test_2_allow_nodes workers=1 worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 command=sleep 1000 workspace_id=**** resource_id=quotau2h98mt**** worker_cpu=1 worker_memory=2Gi allow_nodes=lingjuc47iextvg9-****,lingjuc47iextvg9-**** -
Exclude specific nodes
name=assign_node_test_two_allow_two_deny workers=1 worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 command=sleep 1000 workspace_id=**** resource_id=quotau2h98mt**** worker_cpu=1 worker_memory=2Gi deny_nodes=lingjuc47iextvg9-****,lingjuc47iextvg9-**** -
Specify scheduling nodes & exclude specific nodes
name=assign_node_test_two_allow_two_deny workers=1 worker_image=dsw-registry-vpc.****.cr.aliyuncs.com/pai/easyanimate:1.1.5-pytorch2.2.0-gpu-py310-cu118-ubuntu22.04 command=sleep 1000 workspace_id=**** resource_id=quotau2h98mt**** worker_cpu=1 worker_memory=2Gi allow_nodes=lingjuc47iextvg9-****,lingjuc47iextvg9-**** deny_nodes=lingjuc47iextvg9-****,lingjuc47iextvg9-****
-
Disable pay-as-you-go stock check for job submission
You can configure the disable_ecs_stock_check parameter in the DLC command line to disable the pay-as-you-go stock check.
-
Parameters
Parameter
Description
Example value
disable_ecs_stock_check
Disables the pay-as-you-go stock check. Valid values:
false (default): Enables the pay-as-you-go stock check.
true: Disables the pay-as-you-go stock check.
true or false
-
Examples
Command line parameters
Submit a job using command line parameters. Example:
-
Enable pay-as-you-go stock check
./dlc submit pytorchjob \ --name=test_skip_checking3 \ --command='sleep 1000' \ --workspace_id=**** \ --priority=1 \ --workers=1 \ --worker_image=registry.cn-hangzhou.aliyuncs.com/pai-dlc/tensorflow-training:1.12PAI-gpu-py36-cu101-ubuntu18.04 \ --worker_spec=ecs.g6.xlarge -
Disable pay-as-you-go stock check
./dlc submit pytorchjob \ --name=test_skip_checking3 \ --command='sleep 1000' \ --workspace_id=**** \ --priority=1 \ --workers=1 \ --worker_image=registry.cn-hangzhou.aliyuncs.com/pai-dlc/tensorflow-training:1.12PAI-gpu-py36-cu101-ubuntu18.04 \ --worker_spec=ecs.g6.xlarge \ --disable_ecs_stock_check=true
File-based configuration
Submit a job using a file-based configuration. Example:
./dlc submit pytorchjob -f job_fileHere, job_file is the job parameter file. The content is as follows:
-
Enable pay-as-you-go stock check
name=test_skip_checking3 workers=1 worker_image=registry.cn-hangzhou.aliyuncs.com/pai-dlc/tensorflow-training:1.12PAI-gpu-py36-cu101-ubuntu18.04 command=sleep 1000 workspace_id=**** worker_spec=ecs.g6.xlarge -
Disable pay-as-you-go stock check
name=test_skip_checking3 workers=1 worker_image=registry.cn-hangzhou.aliyuncs.com/pai-dlc/tensorflow-training:1.12PAI-gpu-py36-cu101-ubuntu18.04 command=sleep 1000 workspace_id=**** worker_spec=ecs.g6.xlarge disable_ecs_stock_check=true
-
Related documents
After a job is submitted successfully, you can manage the job using the client tool. For more information, see Command used to stop training jobs and Commands used to query logs or jobs.
You can also manage submitted jobs in the console. For more information, see Manage training jobs.