Data Coder Agent
Data Coder Agent is a senior data expert agent that DataWorks Data Agent provides for data development scenarios. It understands the development tasks that you submit, proactively uses your enterprise-specific knowledge base to perform requirement analysis and data discovery, writes requirement and technical design documents, and completes multi-engine code writing, node and workflow creation, and run and scheduling configuration. Through automatic debugging and error correction, it keeps advancing the task until the data output meets expectations, and deploys only as you instruct, achieving an autonomous end-to-end closed loop from requirement understanding to data delivery.
Feature overview
Data Coder Agent is a senior data expert agent that DataWorks Data Agent provides for data development scenarios. Its capabilities are jointly provided by the following two official expert suites, which integrate platform operations and engine-level code development into a unified development workflow.
DataWorks Data Development suite: provides data development workflow orchestration and platform operation capabilities, covering requirement analysis, data discovery, requirement and technical design document writing, data upload, node and workflow creation, run and scheduling configuration, execution troubleshooting, and pre-deployment checks and deployment.
Data Coder Agent SQL Conventions suite: provides the SQL syntax conventions, best practices, and engine Coder sub-agents of each compute engine, covering MaxCompute, Flink, Hologres, EMR StarRocks, EMR Spark, and EMR Hive. Based on the target engines of a task, it automatically invokes the corresponding sub-agents and conventions to complete code writing and error correction, and to reduce issues such as incompatible functions, misused partitions, and data type differences.
Describe your business background and development goals in natural language. Data Agent then combines your enterprise-specific knowledge base and the actual data environment to coordinate the two suites as needed and advance the following workflow:
Requirement understanding → Data discovery → Requirement and technical design document writing → Node and workflow creation → Code writing and run configuration → Automatic debugging and error correction → Data output → Deployment as instructed
During development, Data Agent automatically connects the tasks of each stage, passes context and intermediate results, and keeps debugging and correcting errors based on execution feedback. This achieves an autonomous end-to-end closed loop from requirement understanding to data delivery, without the need for you to manually switch between agents or pass intermediate results.
Core capabilities
Data Coder Agent combines the platform operation capabilities of the DataWorks Data Development suite and the engine-level code development capabilities of the Data Coder Agent SQL Conventions suite to coordinate requirement analysis, data discovery, solution design, development debugging, and deployment delivery in a unified manner. You can start a development task in natural language without manually switching between agents or passing intermediate results.
Core capability | Description |
Requirement understanding and analysis | Understands the business background and development goals that you submit, combines your enterprise-specific knowledge base to sort out business calibers, expected outputs, and acceptance requirements, clarifies gaps and ambiguities in the requirement, and identifies the data scope and development objects that need further discovery. |
Data discovery and feasibility analysis | Around the development requirement, invokes discovery sub-agents to confirm data sources, target tables, columns, partitions, and related development objects, and to learn about task dependencies, compute resources, and scheduling configuration. It verifies the data conditions on which the requirement depends through actual data and metadata, identifies information gaps and implementation constraints, and provides the basis for solution design. When necessary, it further clarifies or adjusts the requirement. |
Requirement document and technical solution writing | Based on the requirement analysis and data discovery results, writes requirement and technical design documents that specify business calibers, data sources, processing logic, node and workflow design, run and scheduling configuration, and acceptance criteria. It keeps the documents consistent with the actual data environment and provides the basis for subsequent development and verification. |
Multi-engine code development | Based on the target compute engine of the task, invokes the corresponding Coder sub-agent and SQL convention skill as needed to complete code writing and correction. It covers MaxCompute, Flink, Hologres, EMR StarRocks, EMR Spark, and EMR Hive, and handles function, partition, and type system differences according to the syntax conventions and best practices of each engine to reduce cross-engine dialect mismatches. |
Node and workflow development | Invokes DataWorks platform capabilities to turn the development solution into executable nodes and workflows, completes task dependency orchestration, run parameters, and scheduling configuration, and performs data upload and resource, function, and component management as needed, connecting code development with platform execution. |
Run diagnosis and automatic error correction | Executes the development task, locates issues based on run logs, code, metadata, and dependencies, and coordinates platform troubleshooting capabilities and engine Coder sub-agents to fix them. Through the iterative process of "run—diagnose—fix—verify", it keeps advancing the task until the data output meets expectations. |
Output verification and deployment delivery | Checks the task execution results and data output, verifies them against the requirements and acceptance criteria, and performs pre-deployment checks based on the development content and run results. It completes deployment as you instruct, connecting development, verification, and delivery to form an end-to-end closed loop from requirement to data output. |
Supported compute engines and SQL development enhancements
Data development support scope
Data Coder Agent performs development tasks based on DataWorks Data Studio. Data Studio provides various compute engine nodes as well as general nodes such as Shell and Python. For the specific node types and applicable scope, see Node Development Overview. Node support may vary across editions, regions, and task forms. Refer to the actual interface.
SQL code development enhancements
On the basis of the preceding platform capabilities, the Data Coder Agent SQL Conventions expert suite adds SQL convention skills and Coder sub-agents for the following compute engines to enhance code writing and automatic error correction:
MaxCompute (ODPS)
Flink
Hologres
EMR StarRocks
EMR Spark
EMR Hive
When performing SQL development tasks for these engines, Data Agent invokes the corresponding convention skills and Coder sub-agents as needed, and handles differences in engine syntax, functions, data types, and partitions based on the existing convention materials. This reduces dialect mismatches and improves the accuracy and stability of code generation and fixing.
The preceding list indicates the enhancement scope of the SQL Conventions suite. It is not the complete engine support list of Data Coder Agent, and engines that are not listed are not necessarily unsupported for development. The convention coverage varies by engine, and the generated code still needs to be run and verified in your actual environment.
Prerequisites
Product activation: DataWorks is activated. All editions are supported, and you do not need to upgrade to a specific edition.
Instance state: A new Data Agent instance is created and started, and the instance is in the Running state.
Suite dependency: The two official expert suites DataWorks Data Development and Data Coder Agent SQL Conventions are enabled, and the skills required by the task are turned on.
Configuration scope: After you enable or disable an expert suite or a skill within a suite, the change takes effect only on new tasks and does not affect tasks in progress. To apply the changed configuration, create a new task.
Billing
Data Coder Agent applies to all editions of DataWorks. The usage fee is charged according to the billing rules of Data Agent:
Credit Plan subscription activated and a seat bound to the account: The Credits gifted with the seat are deducted first. After the quota is used up, you are billed based on actual usage.
Credit Plan subscription activated but no seat bound to the account, or no subscription activated: You are billed directly based on actual usage.
For more information, see Data Agent fees.
Get started
1. Enter Data Agent
Log on to the DataWorks console. In the left-side navigation pane, click Data Agent to enter Data Agent.
2. Confirm that the expert suites are enabled
In the left-side navigation pane, choose Extensions > Expert Suites. On the DataWorks tab, confirm that the following two official expert suites and the skills required by the task are enabled:
DataWorks Data Development
Data Coder Agent SQL Conventions
Both suites are enabled by default. If you adjust the enabled state of a suite or skill, create a new task for the configuration to take effect.
3. Describe the development requirement
Click New Task and describe the business background, existing metadata, development goals, and expected output in natural language. Example:
Use MaxCompute to develop a DWS summary table for live streaming room product transaction data based on the ctlive_ods.ods_ctlive_trd_order_pay table in the my_project project, and organize the tasks into a workflow that is scheduled daily. First sort out the requirements and discover the available data, clarify the statistical calibers and processing solution, and then start development. After completion, verify the data output. Do not deploy for now.
You can also enter @ in the input box to reference context such as expert suites, tables, nodes, or code files, or enter / to explicitly invoke a skill. For more information, see Extensions.
4. Clarify the requirement and confirm the solution
Data Agent analyzes the requirement based on your description, discovers the related data and resources, and then writes the requirement and technical design documents based on the discovery results. For gaps and ambiguities in business calibers, data sources, or implementation conditions, supplement the information as prompted, and confirm the requirement and solution.
5. Perform development and debugging
After the solution is confirmed, Data Agent coordinates code writing, node and workflow creation, and run and scheduling configuration, and keeps advancing the task through running, diagnosis, and fixing. During execution, confirm the related operations according to the current approval mode and the on-screen prompts. At key stages such as destructive operations, verify the operation scope and impact before you confirm.
6. Verify the output and deploy as needed
View the task execution results and output data to confirm whether they meet the business requirements and acceptance criteria. To deploy, continue to issue a deployment instruction to Data Agent, which performs pre-deployment checks and deploys after your confirmation.
How it works
Data Coder Agent completes development tasks through the collaboration of the Data Agent main agent, discovery sub-agents, and engine Coder sub-agents. The main agent organizes the development workflow, passes context, and performs platform operations. The specialized sub-agents handle specific tasks such as data discovery and code writing and fixing. Together, they connect the complete workflow from requirement understanding to data delivery.
Expert suite collaboration
The capabilities of Data Coder Agent are jointly provided by the following two official expert suites.
Expert suite | Main responsibilities |
DataWorks Data Development | Provides platform capabilities such as node and workflow development, task execution, troubleshooting, data upload, and deployment. |
Data Coder Agent SQL Conventions | Provides SQL syntax conventions, best practices, and engine Coder sub-agents to enhance code writing and fixing for MaxCompute, Flink, Hologres, EMR StarRocks, EMR Spark, and EMR Hive. |
The main agent invokes the corresponding capabilities as the development task requires and uniformly connects the inputs and outputs of each stage, so you do not need to manually pass intermediate results between agents. The enhancement scope of the SQL Conventions suite is not equivalent to the complete engine support scope of the data development platform.
Data discovery and solution design
After understanding the business requirement, the main agent invokes discovery sub-agents to verify the data, development objects, and resource conditions on which the development depends.
Discovery aspect | Main content |
Tables and data | Identifies target tables, confirms columns, partitions, and data sources, and clarifies ambiguities such as tables with the same name. |
Development objects and dependencies | Finds related nodes and workflows, and learns about task dependencies, upstream and downstream relationships, and scheduling configuration. |
Compute and scheduling resources | Learns about project data sources, compute resources, resource groups, and related running conditions. |
Discovery is performed in read-only mode, and the results are returned to the main agent as structured information. Based on the requirement analysis and the actual discovery results, the main agent further clarifies information gaps and then writes the requirement and technical design documents, so that the development solution is built on a verified data environment.
Layered SQL conventions
The SQL Conventions suite is organized in three layers—general conventions, syntax categories, and engine features—to provide the basis for code writing and fixing.
General convention layer: Defines cross-engine SQL writing conventions and code quality requirements.
Syntax category layer: Provides writing guidance by category, such as query (DQL), data definition (DDL), data manipulation (DML), and Script mode.
Engine feature layer: Supplements engine-specific rules such as native syntax, built-in functions, type systems, and partitions and indexes.
The layered conventions balance general development requirements and target engine differences, and reduce issues such as incompatible functions, improper type usage, and misused partitions. The specific convention coverage varies by engine.
Engine identification and code task delegation
When performing tasks such as SQL writing, table creation, query development, or code fixing, the main agent first identifies the target compute engine, loads the corresponding convention skill, and preferentially delegates the code task to the Coder sub-agent of that engine.
In an isolated context, the Coder sub-agent completes code writing or fixing based on the confirmed requirements, table structures, and related conventions, and returns a structured result to the main agent. When the target engine or required data structure information is unclear, it first performs discovery or clarifies with you, to avoid substituting the syntax and behavior of another engine for the target engine rules.
The Coder sub-agent is responsible only for code and does not directly perform platform write operations. Operations such as code saving, node creation, dry runs, and deployment are uniformly performed by the main agent according to the current approval mode.
Execution feedback and automatic error correction
The main agent organizes the running of nodes and workflows and coordinates subsequent handling based on the execution results. When a run fails, it locates the cause based on logs, code, metadata, and dependencies, then invokes the corresponding capabilities to fix the code or adjust the configuration, and runs the task again for verification.
This process forms a feedback closed loop of "run—diagnose—fix—rerun—output verification". The main agent continuously connects the execution results of each round, providing the basis for subsequent debugging, output verification, and pre-deployment checks.
Operation confirmation and security control
Data Coder Agent follows the interactive confirmation mechanism of the new Data Agent. Platform operations are performed according to the current approval mode. At key stages such as plan confirmation, destructive operations, production deployment, and high-cost operations, you must confirm as prompted before the process continues.
The automated workflow connects tasks and execution results, but does not replace your decisions on key operations. After development and verification are complete, the main agent performs pre-deployment checks as you instruct and deploys only after your confirmation.