O Graph Database (GDB) Reader e o GDB Writer permitem a sincronização bidirecional de dados entre o DataWorks e fontes de dados do GDB.
Limites
|
Leitura de dados em lote |
Gravação de dados em lote |
|
|
Adicionar uma fonte de dados
Antes de desenvolver uma tarefa de sincronização no DataWorks, adicione a fonte de dados necessária seguindo as instruções em Gerenciamento de fontes de dados. Consulte as descrições dos parâmetros no console do DataWorks para compreender o significado de cada parâmetro ao adicionar uma fonte de dados.
Desenvolver uma tarefa de sincronização de dados
Para obter informações sobre o ponto de entrada e o procedimento de configuração de uma tarefa de sincronização, consulte os guias de configuração a seguir.
Configurar uma tarefa de sincronização em lote para sincronizar dados de tabela única
Para mais detalhes sobre o procedimento de configuração, consulte Configurar uma tarefa de sincronização em lote usando a interface visual sem código e Configurar uma tarefa de sincronização em lote usando o editor de código.
Para informações sobre todos os parâmetros configurados e o código executado ao usar o editor de código para configurar uma tarefa de sincronização em lote, consulte Apêndice: Código e parâmetros.
Apêndice: Código e parâmetros
Configurar uma tarefa de sincronização em lote usando o editor de código
Para configurar uma tarefa de sincronização em lote por meio do editor de código, defina os parâmetros relevantes no script conforme os requisitos unificados de formato de script. Para mais informações, consulte Configuração em modo de script. As informações a seguir descrevem os parâmetros que devem ser configurados para as fontes de dados ao utilizar o editor de código.
Exemplo de script do Reader
Ao configurar um trabalho de sincronização de dados para gravar no Graph Database (GDB), configure vértices e arestas separadamente:
-
Configure uma tarefa de sincronização para ler dados de vértices de uma instância do GDB
{ "order":{ "hops":[ { "from":"Reader", "to":"Writer" } ] }, "setting":{ "errorLimit":{ "record":"100" // Maximum number of dirty data records to tolerate. }, "jvmOption":"", "speed":{ "concurrent":3, "throttle":true, // If true, throttling is enabled. If false, throttling is disabled and the mbps parameter is ignored. "mbps":"12"// The throttling limit. 1 mbps is equal to 1 MB/s. } }, "steps":[ { "category":"reader", "name":"Reader", "parameter":{ "host": "gdb-xxxxxx.aliyuncs.com", // The endpoint of the GDB instance. "port": 8182, // The port of the GDB instance. "username": "gdb", // The username for accessing the GDB instance. "password": "gdb", // The password for the username. "labelType": "VERTEX", // The type of the label. VERTEX specifies a vertex. "labels": ["label1", "label2"], // A list of labels. If left empty, all vertices are exported. "column": [ { "name": "id", // The field name. "type": "string", // The field type. "columnType": "primaryKey" // Field category. Specifies the vertex primary key, which must be of the STRING type in GDB. }, { "name": "label", // The field name. "type": "string", // The field type. "columnType": "primaryLabel" // Field category. Specifies the vertex label, which must be of the STRING type in GDB. }, { "name": "age", // The property name. "type": "int", // The property type. "columnType": "vertexProperty" // Field category. Specifies a basic property of the vertex in GDB. } ] }, "stepType":"gdb" }, { "category":"writer", "name":"Writer", "parameter":{ "print": true }, "stepType":"stream" } ] } -
Configure uma tarefa de sincronização para ler dados de arestas de uma instância do GDB
{ "order":{ "hops":[ { "from":"Reader", "to":"Writer" } ] }, "setting":{ "errorLimit":{ "record":"100" // Maximum number of dirty data records to tolerate. }, "jvmOption":"", "speed":{ "concurrent":3, "throttle":true,// If true, throttling is enabled. If false, throttling is disabled and the mbps parameter is ignored. "mbps":"12"// The throttling limit. 1 mbps is equal to 1 MB/s. } }, "steps":[ { "category":"reader", "name":"Reader", "parameter":{ "host": "gdb-xxxxxx.aliyuncs.com", // The endpoint of the GDB instance. "port": 8182, // The port of the GDB instance. "username": "gdb", // The username for accessing the GDB instance. "password": "gdb", // The password for the username. "labelType": "EDGE", // The type of the label. EDGE specifies an edge. "labels": ["label1", "label2"], // A list of labels. If left empty, all edges are exported. "column": [ { "name": "id", // The field name. "type": "string", // The field type. "columnType": "primaryKey" // Field category. Specifies the edge primary key, which must be of the STRING type in GDB. }, { "name": "label", // The field name. "type": "string", // The field type. "columnType": "primaryLabel" // Field category. Specifies the edge label, which must be of the STRING type in GDB. }, { "name": "srcId", // The field name. "type": "string", // The field type. "columnType": "srcPrimaryKey" // Field category. Specifies the start vertex primary key, which must be of the STRING type in GDB. }, { "name": "srcLabel", // The field name. "type": "string", // The field type. "columnType": "srcPrimaryLabel" // Field category. Specifies the start vertex label, which must be of the STRING type in GDB. }, { "name": "dstId", // The field name. "type": "string", // The field type. "columnType": "dstPrimaryKey" // Field category. Specifies the end vertex primary key, which must be of the STRING type in GDB. }, { "name": "dstLabel", // The field name. "type": "string", // The field type. "columnType": "dstPrimaryLabel" // Field category. Specifies the end vertex label, which must be of the STRING type in GDB. }, { "name": "weight", // The property name. "type": "double", // The property type. "columnType": "edgeProperty" // Field category. Specifies an edge property. } ] }, "stepType":"gdb" }, { "category":"writer", "name":"Writer", "parameter":{ "print": true }, "stepType":"stream" } ] }
Parâmetros do script do Reader
|
Parâmetro |
Descrição |
Obrigatório |
Valor padrão |
|
host |
O endpoint de conexão da instância do GDB. No console do GDB, clique em Management na instância desejada para visualizar o Internal Endpoint (que corresponde ao host). |
Sim |
Sem valor padrão |
|
port |
A porta usada para conectar-se à instância do GDB. |
Sim |
8182 |
|
username |
O nome de usuário usado para conectar-se à instância do GDB. |
Sim |
Sem valor padrão |
|
password |
A senha usada para conectar-se à instância do GDB. |
Sim |
Sem valor padrão |
|
labels |
O rótulo (nome) do vértice ou da aresta. O GDB Reader pode ler dados de vários vértices ou arestas simultaneamente, especificados como um array, por exemplo ["label1", "label2"]. |
Sim |
Sem valor padrão |
|
labelType |
O tipo do rótulo. Valores válidos:
|
Sim |
Sem valor padrão |
|
column |
Os vértices ou arestas a serem sincronizados. |
Sim |
Sem valor padrão |
|
column -> name |
O nome da propriedade do vértice ou da aresta a ser sincronizada. Obrigatório se propriedades de vértice ou aresta estiverem incluídas. |
Sim |
Sem valor padrão |
|
column -> type |
O tipo de dado da propriedade do vértice ou da aresta a ser sincronizada.
|
Sim |
Sem valor padrão |
|
column -> columnType |
A categoria da propriedade do vértice ou da aresta a ser sincronizada.
|
Sim |
Sem valor padrão |
Exemplo de script do Writer
-
Configure uma tarefa de sincronização para gravar dados de vértices em um banco de dados GDB
{ "order":{ "hops":[ { "from":"Reader", "to":"Writer" } ] }, "setting":{ "errorLimit":{ "record":"100" // Maximum number of dirty data records to tolerate. }, "speed":{ "throttle":true,// If true, throttling is enabled. If false, throttling is disabled and the mbps parameter is ignored. "concurrent":3, // The job concurrency. "mbps":"12"// The throttling limit. 1 mbps is equal to 1 MB/s. } }, "steps":[ { "category":"reader", "name":"Reader", "parameter":{ "column":[ "*" ], "datasource":"_ODPS", "emptyAsNull":true, "guid":"", "isCompress":false, "partition":[], "table":"" }, "stepType":"odps" }, { "category":"writer", "name":"Writer", "parameter": { "datasource": "testGDB", // The name of the data source. "label": "person", // The label, which is the name of the vertex. "srcLabel": "", // Not applicable to vertices. "dstLabel": "", // Not applicable to vertices. "labelType": "VERTEX", // The type of the label. "VERTEX" specifies a vertex. "writeMode": "INSERT", // The policy for handling duplicate primary keys. "idTransRule": "labelPrefix", // The conversion rule for the vertex primary key. "srcIdTransRule": "none", // Not applicable to vertices. "dstIdTransRule": "none", // Not applicable to vertices. "column": [ { "name": "id", // The field name. "value": "#{0}", // Takes the value from the source column at index 0. Concatenation is supported. "type": "string", // The field type. "columnType": "primaryKey" // The field category. `primaryKey` specifies the primary key. }, // The primary key of the vertex. The field name must be `id`, the type must be STRING, and this record is required. { "name": "person_age", "value": "#{1}", // Takes the value from the source column at index 1. Concatenation is supported. "type": "int", "columnType": "vertexProperty" // The field category. `vertexProperty` specifies a vertex property. }, // A property of the vertex. Supported types: INT, LONG, FLOAT, DOUBLE, BOOLEAN, and STRING. { "name": "person_credit", "value": "#{2}", // Takes the value from the source column at index 2. Concatenation is supported. "type": "string", "columnType": "vertexProperty" } // A property of the vertex. ] } "stepType":"gdb" } ], "type":"job", "version":"2.0" } -
Configure uma tarefa de sincronização para gravar dados de arestas em um banco de dados GDB
{ "order":{ "hops":[ { "from":"Reader", "to":"Writer" } ] }, "setting":{ "errorLimit":{ "record":"100" // Maximum number of dirty data records to tolerate. }, "jvmOption":"", "speed":{ "throttle":true,// If true, throttling is enabled. If false, throttling is disabled and the mbps parameter is ignored. "concurrent":3, // The job concurrency. "mbps":"12"// The throttling limit. 1 mbps is equal to 1 MB/s. } }, "steps":[ { "category":"reader", "name":"Reader", "parameter":{ "column":[ "*" ], "datasource":"_ODPS", "emptyAsNull":true, "guid":"", "isCompress":false, "partition":[], "table":"" }, "stepType":"odps" }, { "category":"writer", "name":"Writer", "parameter": { "datasource": "testGDB", // The name of the data source. "label": "use", // The label, which is the name of the edge. "labelType": "EDGE", // The type of the label. `EDGE` specifies an edge. "srcLabel": "person", // The label of the start vertex. "dstLabel": "software", // The label of the end vertex. "writeMode": "INSERT", // The policy for handling duplicate primary keys. "idTransRule": "labelPrefix", // The conversion rule for the edge primary key. "srcIdTransRule": "labelPrefix", // The conversion rule for the primary key of the start vertex. "dstIdTransRule": "labelPrefix", // The conversion rule for the primary key of the end vertex. "column": [ { "name": "id", // The field name. "value": "#{0}", // Takes the value from the source column at index 0. Concatenation is supported. "type": "string", // The field type. "columnType": "primaryKey" // The field category. `primaryKey` specifies the primary key. }, // The primary key of the edge. The field name must be `id` and the type must be STRING. This record is optional. { "name": "id", "value": "#{1}", // Concatenation is supported. Ensure the mapping rule is consistent with the one used when importing vertices. "type": "string", "columnType": "srcPrimaryKey" // The field category. `srcPrimaryKey` specifies the primary key of the start vertex. }, // The primary key of the start vertex. The field name must be `id`, the type must be STRING, and this record is required. { "name": "id", "value": "#{2}", // Concatenation is supported. Ensure the mapping rule is consistent with the one used when importing vertices. "type": "string", "columnType": "dstPrimaryKey" // The field category. `dstPrimaryKey` specifies the primary key of the end vertex. }, // The primary key of the end vertex. The field name must be `id`, the type must be STRING, and this record is required. { "name": "person_use_software_time", "value": "#{3}", // Concatenation is supported. "type": "long", "columnType": "edgeProperty" // The field category. `edgeProperty` specifies an edge property. }, // A property of the edge. Supported types: INT, LONG, FLOAT, DOUBLE, BOOLEAN, and STRING. { "name": "person_regist_software_name", "value": "#{4}", // Concatenation is supported. "type": "string", "columnType": "edgeProperty" }, // A property of the edge. { "name": "id", "value": "#{5}", // Concatenation is supported. "type": "long", "columnType": "edgeProperty" } // A property of the edge with the field name `id`. This is a regular property, not a primary key, and is optional. ] } "stepType":"gdb" } ], "type":"job", "version":"2.0" }
Parâmetros do script do Writer
|
Parâmetro |
Descrição |
Obrigatório |
Valor padrão |
|
datasource |
O nome da fonte de dados. O valor deve corresponder exatamente ao nome da fonte de dados adicionada no editor de código. |
Sim |
Sem valor padrão |
|
label |
O rótulo, que corresponde ao nome do vértice ou da aresta. O GDB Writer pode obter rótulos de colunas na tabela de origem. Por exemplo, se você definir este parâmetro como #{0}, o GDB Writer usará o valor da primeira coluna como rótulo. O índice da coluna começa em 0. |
Sim |
Sem valor padrão |
|
labelType |
O tipo do rótulo. Valores válidos:
|
Sim |
Sem valor padrão |
|
srcLabel |
|
Não |
Sem valor padrão |
|
dstLabel |
|
Não |
Sem valor padrão |
|
writeMode |
Define como o GDB Writer trata registros de dados com chaves primárias duplicadas. Valores válidos:
|
Sim |
INSERT |
|
idTransRule |
A regra de conversão para a chave primária. Valores válidos:
|
Sim |
none |
|
srcIdTransRule |
A regra de conversão para a chave primária do vértice inicial quando labelType está definido como EDGE. Valores válidos:
|
Obrigatório quando o parâmetro labelType está definido como EDGE |
none |
|
dstIdTransRule |
A regra de conversão para a chave primária do vértice final quando labelType está definido como EDGE. Valores válidos:
|
Obrigatório quando o parâmetro labelType está definido como EDGE |
none |
|
column |
Os vértices ou arestas a serem sincronizados.
Exemplo de propriedades
|
Sim |
Sem valor padrão |