Compaction Service
Compaction Service offloads compaction tasks from your primary compute group to a dedicated service in EMR Serverless StarRocks 3.5 (Beta), enabling workload isolation, auto scaling, and performance optimization.
Overview
Core benefits
|
Capability |
Description |
|
Workload isolation |
Running compaction tasks on a dedicated service prevents resource contention with business workloads such as queries and data ingestion, ensuring application stability. |
|
Auto scaling |
Scales automatically based on the compaction workload. Configure Min CU and Max CU to balance compaction timeliness and cost. |
|
Out-of-the-box |
Automatically created for clusters running version 3.5. Enable it with a single click in the console — no additional resource purchase or configuration required. |
Performance optimizations
The Compaction Service also provides these performance optimizations:
-
Peer cache reads: During a compaction task, the service directly pulls data from cached nodes in the primary compute group (peer cache). This avoids accessing object storage (remote I/O) and significantly improves read performance for compaction.
-
Cache push: After compaction is complete, the Compaction Service asynchronously pushes the merged data files to the nodes in the primary compute group. This prevents query performance from degrading due to cache misses that would otherwise require accessing object storage.
Prerequisites
-
Your EMR Serverless StarRocks cluster is version 3.5 or later.
-
Your cluster uses the storage-compute separation architecture.
Enable the Compaction Service
Enable the Compaction Service in the EMR Serverless StarRocks console:
-
Go to the E-MapReduce Serverless StarRocks instance list page.
-
Log on to the E-MapReduce console.
-
In the navigation pane on the left, choose .
-
In the top menu bar, select the required region.
-
-
Click the ID of the target instance.
-
Click the Compaction Service tab. On the Basic Information page, click Start the service.
-
In the panel that appears, configure Minimum CU and Maximum CU.
-
Click Start the service.
Stop the Compaction Service
In the console, click Disable the service. In the confirmation panel, select the Risk Confirmation checkbox, and then click Confirm service shutdown. After the service is stopped, compaction tasks revert to running on the primary compute group. In-progress tasks will complete normally.
Auto scaling
CU configuration
Configure the auto scaling range for the Compaction Service:
|
Parameter |
Description |
Recommendation |
|
Min CU |
The minimum number of compute units. The service scales in to this value during idle periods. |
Set this to the minimum value required to meet baseline compaction needs. |
|
Max CU |
The maximum number of compute units. The service scales out to this value during peak loads. |
Set this based on your peak write throughput and compaction score. |
Scaling policy
The Compaction Service scales automatically based on the following metrics:
-
compaction score: Reflects the accumulation of data versions. A higher score indicates greater compaction pressure.
-
Task load: The ratio of current compaction tasks to available resources.
The system automatically scales out when the compaction score consistently rises or when tasks are queued. After the load decreases, the system gradually scales in to the configured Min CU.
Best practices
Recommended scenarios for the Compaction Service:
-
High write-throughput workloads: Continuous, high-frequency writes cause the compaction score to rise, degrading query performance.
-
Query-sensitive workloads: The service prevents compaction from competing for query resources, which is ideal for latency-sensitive applications.
-
Cost-optimization scenarios: Auto scaling uses on-demand resources for compaction, reducing standing costs.
Usage notes
-
The Compaction Service is only available for clusters that use the storage-compute separation (shared-data) architecture.
-
After you enable the Compaction Service, compaction tasks for all tables run on the service.
-
The compute unit (CU) resources for the Compaction Service are billed separately based on usage. Configure Min CU and Max CU appropriately.
-
If you stop the Compaction Service, compaction tasks automatically revert to running on their respective primary compute groups. In-progress tasks will complete normally.
-
To minimize potential system impact, enable the Compaction Service for the first time during off-peak hours.