OSS-HDFS
OSS-HDFS Service (JindoFS Service) is a cloud-native data lake storage product. An OSS-HDFS data source provides a bidirectional channel to read from and write to OSS-HDFS. This topic describes the data synchronization capabilities that DataWorks provides for OSS-HDFS.
Supported capabilities
Capability | Supported |
Offline read | Yes |
Offline write | Yes |
Real-time write | Yes |
Limitations
Offline read
The network connection from a resource group to OSS-HDFS can be complex. To run data synchronization tasks, use a Serverless resource group (recommended) or an exclusive resource group for Data Integration. Ensure that your resource group can access OSS-HDFS over the network.
OSS-HDFS Reader supports the following:
Files in text, CSV, ORC, and Parquet formats. The file content must be a logical two-dimensional table.
Reading multiple data types and column constants.
Recursive reads and the wildcard characters
*and?.Concurrent reads from multiple files. The actual number of concurrent threads is the smaller value between the number of files to read and the
concurrentsetting.
OSS-HDFS Reader does not support multi-threaded concurrent reads from a single file due to the internal chunking algorithm for single files.
Offline write
OSS-HDFS Writer supports only text, ORC, and Parquet formats. The file content must be a logical two-dimensional table.
For text files, ensure that the field delimiter used for writing matches the delimiter used when creating the Hive table. This ensures that the data written to OSS-HDFS maps correctly to Hive table fields.
Real-time write
Supports real-time writes.
Supports real-time writes for Hudi format version 0.14.x.
Supported field types
Offline read
OSS-HDFS Reader converts data types from ParquetFile, ORCFile, TextFile, and CsvFile to the internal types that Data Integration supports.
Type category | OSS-HDFS data types |
Integer | TINYINT, SMALLINT, INT, BIGINT |
Floating-point | FLOAT, DOUBLE, DECIMAL |
String | STRING, CHAR, VARCHAR |
Date and time | DATE, TIMESTAMP |
Boolean | BOOLEAN |
The following examples illustrate internal type representations:
LONG: Integer data in an OSS-HDFS file, such as 123456789.
DOUBLE: Floating-point data in an OSS-HDFS file, such as 3.1415.
BOOLEAN: Boolean data in an OSS-HDFS file, such as true or false. Values are not case-sensitive.
DATE: Date and time data in an OSS-HDFS file, such as 2014-12-31 00:00:00.
Offline write
OSS-HDFS Writer writes files in TextFile, ORCFile, and ParquetFile formats to a specified path in the OSS-HDFS file system.
Type category | OSS-HDFS data types |
Integer | TINYINT, SMALLINT, INT, BIGINT |
Floating-point | FLOAT, DOUBLE |
String | CHAR, VARCHAR, STRING |
Boolean | BOOLEAN |
Date and time | DATE, TIMESTAMP |
Add a data source
Before developing a synchronization task in DataWorks, add the required data source by following the instructions in Data source management. Parameter descriptions are available in the DataWorks console when you add a data source.
Develop a data synchronization task
Configure an offline synchronization task for a single table
For step-by-step instructions, see Configure a task in the codeless UI and Configure a task in the code editor.
For a complete parameter list and script example, see Appendix: Script demos and parameter descriptions.
Configure a real-time synchronization task for a single table
See Configure real-time incremental synchronization for a single table.
Configure a full and incremental real-time synchronization task for an entire database
See Configure a real-time synchronization task for an entire database.
Appendix: Script demos and parameter descriptions
Reader script demo
All parameters follow the unified script format required by the code editor. For format details, see Configure a task in the code editor.
{
"type": "job",
"version": "2.0",
"steps": [
{
"stepType": "oss_hdfs",
"parameter": {
"path": "",
"datasource": "",
"column": [
{
"index": 0,
"type": "string"
},
{
"index": 1,
"type": "long"
},
{
"index": 2,
"type": "double"
},
{
"index": 3,
"type": "boolean"
},
{
"format": "yyyy-MM-dd HH:mm:ss",
"index": 4,
"type": "date"
}
],
"fieldDelimiter": ",",
"encoding": "UTF-8",
"fileFormat": ""
},
"name": "Reader",
"category": "reader"
},
{
"stepType": "stream",
"parameter": {},
"name": "Writer",
"category": "writer"
}
],
"setting": {
"errorLimit": {
"record": ""
},
"speed": {
"concurrent": 3,
"throttle": true,
"mbps": "12"
}
},
"order": {
"hops": [
{
"from": "Reader",
"to": "Writer"
}
]
}
}Reader script parameters
Parameter | Description | Required | Default value |
| The path of the file or directory to read. Three input styles are supported: OPTION 1: Single file — OSS-HDFS Reader uses a single thread to read the file. OPTION 2: Multiple files — OSS-HDFS Reader reads files concurrently. The actual thread count is the smaller value between the number of files and the | Yes | None |
| The file type. Valid values: | Yes | None |
| The list of fields to read. Set to | Yes | None |
| The field delimiter for reading TextFile data. Not required for ORC or Parquet files. | No |
|
| The file encoding. | No |
|
| The string to treat as a null value. For example, setting | No | None |
| The compression format. Valid values: | No | None |
Writer script demo
{
"type": "job",
"version": "2.0",
"steps": [
{
"stepType": "stream",
"parameter": {},
"name": "Reader",
"category": "reader"
},
{
"stepType": "oss_hdfs",
"parameter": {
"path": "",
"fileName": "",
"compress": "",
"datasource": "",
"column": [
{
"name": "col1",
"type": "string"
},
{
"name": "col2",
"type": "int"
},
{
"name": "col3",
"type": "double"
},
{
"name": "col4",
"type": "boolean"
},
{
"name": "col5",
"type": "date"
}
],
"writeMode": "",
"fieldDelimiter": ",",
"encoding": "",
"fileFormat": "text"
},
"name": "Writer",
"category": "writer"
}
],
"setting": {
"errorLimit": {
"record": ""
},
"speed": {
"concurrent": 3,
"throttle": false
}
},
"order": {
"hops": [
{
"from": "Reader",
"to": "Writer"
}
]
}
}Writer script parameters
Parameter | Description | Required | Default value |
| The file type. Valid values: | Yes | None |
| The path in the OSS-HDFS file system where data is stored. OSS-HDFS Writer writes multiple files to this directory based on the concurrency configuration. When associating with a Hive table, specify the Hive table's storage path on OSS-HDFS. | Yes | None |
| The base name for output files. A random suffix is appended to this name for each concurrent thread to create the actual file names. | Yes | None |
| The fields to write. Writing to a subset of columns is not supported. When associating with a Hive table, specify all field names and types. Use | Yes (not required if | None |
| How OSS-HDFS Writer handles existing files before writing. OSS-HDFS Writer uses a write-then-rename strategy: data is first written to a temporary directory named using the | Yes | None |
| The field delimiter for output files. Only single-character delimiters are supported; multiple characters cause a runtime error. Not required when | Yes (not required if | None |
| The compression format for text files. Valid values: | No | None |
| The file encoding. | No |
|
| The schema definition for Parquet output files. Only takes effect when | No | None |