Features

Updated at:

Data synchronization simplifies the ingestion and synchronization of batch and real-time data from heterogeneous data sources. The system offers comprehensive features for configuring data ingestion and monitoring task execution. This process ensures stable and manageable data ingestion to fulfill data aggregation requirements across various platforms, data sources, and applications. It also supports synchronizing and aggregating spatial data from mainstream Geographic Information System (GIS) platforms, such as ArcGIS and SuperMap, and open source PostGIS spatial databases (invitational preview).

The custom sync task configuration feature lets you quickly create scheduled and real-time sync tasks. To create a sync task, you can select data from a registered source and choose a destination. The system provides several features to simplify task configuration, such as automatic table creation at the destination, automatic field mapping between the source and destination, batch creation of sync tasks, and referencing data modules for configuration.

After you configure data synchronization, the module provides comprehensive operations and maintenance (O&M) and monitoring for task execution. It collects statistics on core metrics, such as the total number of tasks, online tasks, and running tasks. It also monitors the execution status of failed or abnormal tasks and the status of current instances. The module tracks and displays information, such as task start time, end time, duration, speed, status, and execution logs. You can also rerun scheduled tasks.

Data source management

  • You can register and manage source and destination data sources in a unified manner. You can also configure various data source types, such as relational databases, files, and message queues, and validate the connectivity of the configured data sources.

  • You can synchronize metadata and view data objects for configured data sources to understand your data assets.

Data template management

  • You can create data templates for semi-structured and unstructured data. You can also customize data fields and field types, and use features such as edit and delete.

  • You can reference existing data templates in offline and real-time data sync tasks. The tasks then run based on the data structure that is defined in the template.

Offline data synchronization

  • You can create offline tasks for single tables or in batches.

  • You can select data from registered sources and destinations. The feature supports various common offline sync paths, such as synchronizing spatial data from PostGIS, Ganos, SuperMap SDX, and ArcGIS SDE to PostGIS and Ganos (invitational preview).

  • You can use features such as automatic table creation for full synchronization, automatic mapping for fields with the same name, and task scheduling configuration.

  • You can monitor the runtime properties and operational logs of offline task instances.

  • You can rerun instances of non-incremental offline sync tasks.

Real-time data synchronization

  • You can create single-table real-time tasks with DataHub as the destination.

  • The feature supports five real-time data synchronization modes: Oracle Change Data Capture (CDC), MySQL binlog, SQL Server CDC, RabbitMQ, and Kafka.

  • The feature provides comprehensive configuration for real-time data ingestion to ensure the stability and manageability of real-time tasks.

Configurable synchronization

  • You can configure offline and real-time sync tasks using custom scripts. You can also reference source Reader specifications and destination Writer specifications to quickly configure sync task scripts.

  • You can configure scheduling, ancestor node dependencies, and runtime resources for script-based sync tasks.

  • After you test and publish script-based tasks, you can view runtime property metrics and operational logs in Service Monitoring.