ALIYUN::CMS2::AggTaskGroup

更新时间:
复制 MD 格式

ALIYUN::CMS2::AggTaskGroup类型用于创建聚合任务组。

语法

{
  "Type": "ALIYUN::CMS2::AggTaskGroup",
  "Properties": {
    "AggTaskGroupName": String,
    "AggTaskGroupConfig": String,
    "InstanceId": String,
    "TargetPrometheusId": String,
    "AggTaskGroupConfigType": String,
    "CronExpr": String,
    "Description": String,
    "Delay": Integer,
    "FromTime": Integer,
    "MaxRetries": Integer,
    "MaxRunTimeInSeconds": Integer,
    "OverrideIfExists": Boolean,
    "PrecheckString": Map,
    "Status": String,
    "ScheduleTimeExpr": String,
    "ScheduleMode": String,
    "ToTime": Integer,
    "Tags": List
  }
}

属性

属性名称

类型

必须

允许更新

描述

约束

AggTaskGroupConfig

String

是

是

聚合任务组的配置。

目前仅支持 RecordingRuleYaml 格式,需要符合开源 Prometheus 的 RecordingRule 的格式要求。示例:

groups:
  - name: node.rules
    interval: 60s
    rules:
      - record: node_namespace_pod:kube_pod_info:
        expr: "max(label_replace(kube_pod_info{job=\"kubernetes-pods-kube-state-metrics\"}, \"pod\", \"$1\", \"pod\", \"(.*)\")) by (node, namespace, pod, cluster)"

AggTaskGroupName

String

是

是

聚合任务组的名称。

无

InstanceId

String

是

否

聚合任务组读取数据的源 Prometheus 实例ID。

无

TargetPrometheusId

String

是

是

聚合任务组的目标 Prometheus 实例ID。

无

AggTaskGroupConfigType

String

否

是

聚合任务组的配置类型。

默认值:

  • RecordingRuleYaml

CronExpr

String

否

是

当 ScheduleMode 为 "Cron" 时使用的 cron 表达式。

例如,"0/1 * * * *" 表示从第 0 分钟开始每 1 分钟调度一次。

Delay

Integer

否

是

调度的固定延迟。

单位为秒。默认值:30。

Description

String

否

是

聚合任务组的描述。

无

FromTime

Integer

否

是

调度开始时间对应的秒级时间戳。

无

MaxRetries

Integer

否

是

执行聚合任务的最大重试次数。

默认值:20。

MaxRunTimeInSeconds

Integer

否

是

执行聚合任务的最大重试时间。

单位为秒。默认值:600。

OverrideIfExists

Boolean

否

否

创建时是否覆盖并更新已存在的同名聚合任务组。

仅创建时有效,不可更新。

PrecheckString

Map

否

是

预检配置。

默认不配置。输入的字符串需要能被正确 JSON 解析。

示例:

{
  "policy": "skip",
  "prometheusId": "xxx",
  "query": "scalar(sum(count_over_time(up{job=\"_arms/kubelet/cadvisor\"}[15s])) / 21)",
  "threshold": 0.5,
  "timeout": 15,
  "type": "promql"
}

ScheduleMode

String

否

是

调度模式。

取值:

  • Cron

  • FixedRate

ScheduleTimeExpr

String

否

是

调度时间表达式。

推荐使用 "@s" 或 "@m",表示调度时间窗口的取整粒度。默认值:"@m"。

Status

String

否

是

聚合任务组的状态。

取值:

  • Running

  • Stopped

Tags

List

否

是

聚合任务组的标签。

更多信息,请参考Tags属性。

ToTime

Integer

否

是

调度结束时间对应的秒级时间戳。

0 表示调度永不停止。

Tags语法

"Tags": [
  {
    "Value": String,
    "Key": String
  }
]

Tags属性

属性名称

类型

必须

允许更新

描述

约束

Key

String

是

否

标签的键。

无

Value

String

否

否

标签的值。

无

返回值

Fn::GetAtt

  • AggTaskGroupName:聚合任务组的名称。

  • Status:聚合任务组的当前状态。

  • AggTaskGroupConfigHash:聚合任务组的配置哈希值。

  • AggTaskGroupId:聚合任务组的ID。

  • SourcePrometheusId:聚合任务组的源 Prometheus 实例ID。

示例

场景 1 :为已有 Prometheus 实例创建聚合任务组

ROSTemplateFormatVersion: '2015-09-01'
Description:
  zh-cn: 为已有的Prometheus实例创建一个简单的聚合任务组,使用默认调度配置和基础Recording Rule。
  en: Create a simple aggregation task group for an existing Prometheus instance with default schedule configuration and basic recording rules.
Parameters:
  SourcePrometheusId:
    Type: String
    Label:
      zh-cn: 源Prometheus实例ID
      en: Source Prometheus Instance ID
    Description:
      zh-cn: 聚合任务组读取数据的源Prometheus实例ID。
      en: The ID of the source Prometheus instance that the agg task group reads data from.
  TargetPrometheusId:
    Type: String
    Label:
      zh-cn: 目标Prometheus实例ID
      en: Target Prometheus Instance ID
    Description:
      zh-cn: 聚合结果写入的目标Prometheus实例ID。
      en: The ID of the target Prometheus instance that the agg task group writes data to.
  AggTaskGroupName:
    Type: String
    Label:
      zh-cn: 聚合任务组名称
      en: Agg Task Group Name
    Description:
      zh-cn: 聚合任务组的名称。
      en: The name of the aggregation task group.
    Default: simple-agg-task-group
Resources:
  AggTaskGroup:
    Type: ALIYUN::CMS2::AggTaskGroup
    Properties:
      AggTaskGroupName:
        Ref: AggTaskGroupName
      InstanceId:
        Ref: SourcePrometheusId
      TargetPrometheusId:
        Ref: TargetPrometheusId
      AggTaskGroupConfig: |-
        groups:
          - name: basic_rules
            rules:
              - record: job:node_cpu_usage:avg_rate5m
                expr: avg by (job) (rate(node_cpu_seconds_total{mode!="idle"}[5m]))
      Description: Simple aggregation task group
Outputs:
  AggTaskGroupId:
    Description:
      zh-cn: 聚合任务组ID。
      en: Agg task group ID.
    Label:
      zh-cn: 聚合任务组ID
      en: Agg Task Group ID
    Value:
      Fn::GetAtt:
        - AggTaskGroup
        - AggTaskGroupId
  AggTaskGroupName:
    Description:
      zh-cn: 聚合任务组名称。
      en: Agg task group name.
    Label:
      zh-cn: 聚合任务组名称
      en: Agg Task Group Name
    Value:
      Fn::GetAtt:
        - AggTaskGroup
        - AggTaskGroupName
  Status:
    Description:
      zh-cn: 聚合任务组状态。
      en: Agg task group status.
    Label:
      zh-cn: 聚合任务组状态
      en: Status
    Value:
      Fn::GetAtt:
        - AggTaskGroup
        - Status
Metadata:
  ALIYUN::ROS::Interface:
    ParameterGroups:
      - Parameters:
          - SourcePrometheusId
          - TargetPrometheusId
          - AggTaskGroupName
{
  "ROSTemplateFormatVersion": "2015-09-01",
  "Description": {
    "zh-cn": "为已有的Prometheus实例创建一个简单的聚合任务组,使用默认调度配置和基础Recording Rule。",
    "en": "Create a simple aggregation task group for an existing Prometheus instance with default schedule configuration and basic recording rules."
  },
  "Parameters": {
    "SourcePrometheusId": {
      "Type": "String",
      "Label": {
        "zh-cn": "源Prometheus实例ID",
        "en": "Source Prometheus Instance ID"
      },
      "Description": {
        "zh-cn": "聚合任务组读取数据的源Prometheus实例ID。",
        "en": "The ID of the source Prometheus instance that the agg task group reads data from."
      }
    },
    "TargetPrometheusId": {
      "Type": "String",
      "Label": {
        "zh-cn": "目标Prometheus实例ID",
        "en": "Target Prometheus Instance ID"
      },
      "Description": {
        "zh-cn": "聚合结果写入的目标Prometheus实例ID。",
        "en": "The ID of the target Prometheus instance that the agg task group writes data to."
      }
    },
    "AggTaskGroupName": {
      "Type": "String",
      "Label": {
        "zh-cn": "聚合任务组名称",
        "en": "Agg Task Group Name"
      },
      "Description": {
        "zh-cn": "聚合任务组的名称。",
        "en": "The name of the aggregation task group."
      },
      "Default": "simple-agg-task-group"
    }
  },
  "Resources": {
    "AggTaskGroup": {
      "Type": "ALIYUN::CMS2::AggTaskGroup",
      "Properties": {
        "AggTaskGroupName": {
          "Ref": "AggTaskGroupName"
        },
        "InstanceId": {
          "Ref": "SourcePrometheusId"
        },
        "TargetPrometheusId": {
          "Ref": "TargetPrometheusId"
        },
        "AggTaskGroupConfig": "groups:\n  - name: basic_rules\n    rules:\n      - record: job:node_cpu_usage:avg_rate5m\n        expr: avg by (job) (rate(node_cpu_seconds_total{mode!=\"idle\"}[5m]))",
        "Description": "Simple aggregation task group"
      }
    }
  },
  "Outputs": {
    "AggTaskGroupId": {
      "Description": {
        "zh-cn": "聚合任务组ID。",
        "en": "Agg task group ID."
      },
      "Label": {
        "zh-cn": "聚合任务组ID",
        "en": "Agg Task Group ID"
      },
      "Value": {
        "Fn::GetAtt": [
          "AggTaskGroup",
          "AggTaskGroupId"
        ]
      }
    },
    "AggTaskGroupName": {
      "Description": {
        "zh-cn": "聚合任务组名称。",
        "en": "Agg task group name."
      },
      "Label": {
        "zh-cn": "聚合任务组名称",
        "en": "Agg Task Group Name"
      },
      "Value": {
        "Fn::GetAtt": [
          "AggTaskGroup",
          "AggTaskGroupName"
        ]
      }
    },
    "Status": {
      "Description": {
        "zh-cn": "聚合任务组状态。",
        "en": "Agg task group status."
      },
      "Label": {
        "zh-cn": "聚合任务组状态",
        "en": "Status"
      },
      "Value": {
        "Fn::GetAtt": [
          "AggTaskGroup",
          "Status"
        ]
      }
    }
  },
  "Metadata": {
    "ALIYUN::ROS::Interface": {
      "ParameterGroups": [
        {
          "Parameters": [
            "SourcePrometheusId",
            "TargetPrometheusId",
            "AggTaskGroupName"
          ]
        }
      ]
    }
  }
}

场景 2 :创建 Prometheus 实例并配置 Cron 调度聚合任务组

ROSTemplateFormatVersion: '2015-09-01'
Description:
  zh-cn: 创建源和目标Prometheus实例,并配置聚合任务组,使用Cron调度模式、自定义重试策略和多条Recording Rule聚合规则。
  en: Create source and target Prometheus instances, configure aggregation task group with Cron schedule mode, custom retry policy, and multiple recording rules.
Parameters:
  PrometheusInstanceNamePrefix:
    Type: String
    Label:
      zh-cn: Prometheus实例名称前缀
      en: Prometheus Instance Name Prefix
    Description:
      zh-cn: 源和目标Prometheus实例的名称前缀。
      en: Name prefix for source and target Prometheus instances.
    Default: monitoring
    MinLength: 1
    MaxLength: 100
  StorageDuration:
    Type: Number
    Label:
      zh-cn: 存储时长(天)
      en: Storage Duration (Days)
    Description:
      zh-cn: Prometheus实例的数据存储时长。
      en: Data storage duration of the Prometheus instance.
    AllowedValues:
      - 15
      - 30
      - 60
      - 90
      - 180
    Default: 90
  AggTaskGroupName:
    Type: String
    Label:
      zh-cn: 聚合任务组名称
      en: Agg Task Group Name
    Description:
      zh-cn: 聚合任务组的名称。
      en: The name of the aggregation task group.
    Default: cron-agg-task-group
  CronExpr:
    Type: String
    Label:
      zh-cn: Cron表达式
      en: Cron Expression
    Description:
      zh-cn: 定时调度表达式,例如 "0/5 * * * *" 表示每5分钟执行一次。
      en: Cron schedule expression, e.g. "0/5 * * * *" means every 5 minutes.
    Default: 0/5 * * * *
  MaxRetries:
    Type: Number
    Label:
      zh-cn: 最大重试次数
      en: Max Retries
    Description:
      zh-cn: 聚合任务执行的最大重试次数。
      en: The maximum number of retries for executing the agg task.
    Default: 20
  MaxRunTimeInSeconds:
    Type: Number
    Label:
      zh-cn: 最大运行时间(秒)
      en: Max Run Time (Seconds)
    Description:
      zh-cn: 聚合任务执行的最大运行时间。
      en: The maximum run time for executing the agg task in seconds.
    Default: 600
Resources:
  SourcePrometheus:
    Type: ALIYUN::CMS2::PrometheusInstance
    Properties:
      PrometheusInstanceName:
        Fn::Sub: ${PrometheusInstanceNamePrefix}-source
      StorageDuration:
        Ref: StorageDuration
      ArchiveDuration: 0
      PaymentType: POSTPAY
      EnableAuthFreeRead: true
      EnableAuthFreeWrite: false
      EnableAuthToken: true
  TargetPrometheus:
    Type: ALIYUN::CMS2::PrometheusInstance
    Properties:
      PrometheusInstanceName:
        Fn::Sub: ${PrometheusInstanceNamePrefix}-target
      StorageDuration:
        Ref: StorageDuration
      ArchiveDuration: 0
      PaymentType: POSTPAY
      EnableAuthFreeRead: true
      EnableAuthFreeWrite: false
      EnableAuthToken: true
  AggTaskGroup:
    Type: ALIYUN::CMS2::AggTaskGroup
    DependsOn:
      - SourcePrometheus
      - TargetPrometheus
    Properties:
      AggTaskGroupName:
        Ref: AggTaskGroupName
      InstanceId:
        Fn::GetAtt:
          - SourcePrometheus
          - PrometheusInstanceId
      TargetPrometheusId:
        Fn::GetAtt:
          - TargetPrometheus
          - PrometheusInstanceId
      ScheduleMode: Cron
      CronExpr:
        Ref: CronExpr
      MaxRetries:
        Ref: MaxRetries
      MaxRunTimeInSeconds:
        Ref: MaxRunTimeInSeconds
      Delay: 60
      ScheduleTimeExpr: '@m'
      OverrideIfExists: true
      Description: Aggregation task group with Cron schedule and custom retry policy
      AggTaskGroupConfig: |-
        groups:
          - name: cpu_aggregation
            rules:
              - record: job:node_cpu_usage:avg_rate5m
                expr: avg by (job) (rate(node_cpu_seconds_total{mode!="idle"}[5m]))
              - record: job:node_cpu_usage:max_rate5m
                expr: max by (job) (rate(node_cpu_seconds_total{mode!="idle"}[5m]))
          - name: memory_aggregation
            rules:
              - record: job:node_memory_usage:avg
                expr: avg by (job) (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes)
              - record: job:node_memory_usage:percent
                expr: avg by (job) ((node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / node_memory_MemTotal_bytes * 100)
      Tags:
        - Key: Environment
          Value: Production
        - Key: Team
          Value: SRE
Outputs:
  AggTaskGroupId:
    Description:
      zh-cn: 聚合任务组ID。
      en: Agg task group ID.
    Label:
      zh-cn: 聚合任务组ID
      en: Agg Task Group ID
    Value:
      Fn::GetAtt:
        - AggTaskGroup
        - AggTaskGroupId
  SourcePrometheusId:
    Description:
      zh-cn: 源Prometheus实例ID。
      en: Source Prometheus instance ID.
    Label:
      zh-cn: 源Prometheus实例ID
      en: Source Prometheus ID
    Value:
      Fn::GetAtt:
        - SourcePrometheus
        - PrometheusInstanceId
  TargetPrometheusId:
    Description:
      zh-cn: 目标Prometheus实例ID。
      en: Target Prometheus instance ID.
    Label:
      zh-cn: 目标Prometheus实例ID
      en: Target Prometheus ID
    Value:
      Fn::GetAtt:
        - TargetPrometheus
        - PrometheusInstanceId
  AggTaskGroupStatus:
    Description:
      zh-cn: 聚合任务组状态。
      en: Agg task group status.
    Label:
      zh-cn: 聚合任务组状态
      en: Status
    Value:
      Fn::GetAtt:
        - AggTaskGroup
        - Status
Metadata:
  ALIYUN::ROS::Interface:
    ParameterGroups:
      - Parameters:
          - PrometheusInstanceNamePrefix
          - StorageDuration
      - Parameters:
          - AggTaskGroupName
          - CronExpr
          - MaxRetries
          - MaxRunTimeInSeconds
{
  "ROSTemplateFormatVersion": "2015-09-01",
  "Description": {
    "zh-cn": "创建源和目标Prometheus实例,并配置聚合任务组,使用Cron调度模式、自定义重试策略和多条Recording Rule聚合规则。",
    "en": "Create source and target Prometheus instances, configure aggregation task group with Cron schedule mode, custom retry policy, and multiple recording rules."
  },
  "Parameters": {
    "PrometheusInstanceNamePrefix": {
      "Type": "String",
      "Label": {
        "zh-cn": "Prometheus实例名称前缀",
        "en": "Prometheus Instance Name Prefix"
      },
      "Description": {
        "zh-cn": "源和目标Prometheus实例的名称前缀。",
        "en": "Name prefix for source and target Prometheus instances."
      },
      "Default": "monitoring",
      "MinLength": 1,
      "MaxLength": 100
    },
    "StorageDuration": {
      "Type": "Number",
      "Label": {
        "zh-cn": "存储时长(天)",
        "en": "Storage Duration (Days)"
      },
      "Description": {
        "zh-cn": "Prometheus实例的数据存储时长。",
        "en": "Data storage duration of the Prometheus instance."
      },
      "AllowedValues": [
        15,
        30,
        60,
        90,
        180
      ],
      "Default": 90
    },
    "AggTaskGroupName": {
      "Type": "String",
      "Label": {
        "zh-cn": "聚合任务组名称",
        "en": "Agg Task Group Name"
      },
      "Description": {
        "zh-cn": "聚合任务组的名称。",
        "en": "The name of the aggregation task group."
      },
      "Default": "cron-agg-task-group"
    },
    "CronExpr": {
      "Type": "String",
      "Label": {
        "zh-cn": "Cron表达式",
        "en": "Cron Expression"
      },
      "Description": {
        "zh-cn": "定时调度表达式,例如 \"0/5 * * * *\" 表示每5分钟执行一次。",
        "en": "Cron schedule expression, e.g. \"0/5 * * * *\" means every 5 minutes."
      },
      "Default": "0/5 * * * *"
    },
    "MaxRetries": {
      "Type": "Number",
      "Label": {
        "zh-cn": "最大重试次数",
        "en": "Max Retries"
      },
      "Description": {
        "zh-cn": "聚合任务执行的最大重试次数。",
        "en": "The maximum number of retries for executing the agg task."
      },
      "Default": 20
    },
    "MaxRunTimeInSeconds": {
      "Type": "Number",
      "Label": {
        "zh-cn": "最大运行时间(秒)",
        "en": "Max Run Time (Seconds)"
      },
      "Description": {
        "zh-cn": "聚合任务执行的最大运行时间。",
        "en": "The maximum run time for executing the agg task in seconds."
      },
      "Default": 600
    }
  },
  "Resources": {
    "SourcePrometheus": {
      "Type": "ALIYUN::CMS2::PrometheusInstance",
      "Properties": {
        "PrometheusInstanceName": {
          "Fn::Sub": "${PrometheusInstanceNamePrefix}-source"
        },
        "StorageDuration": {
          "Ref": "StorageDuration"
        },
        "ArchiveDuration": 0,
        "PaymentType": "POSTPAY",
        "EnableAuthFreeRead": true,
        "EnableAuthFreeWrite": false,
        "EnableAuthToken": true
      }
    },
    "TargetPrometheus": {
      "Type": "ALIYUN::CMS2::PrometheusInstance",
      "Properties": {
        "PrometheusInstanceName": {
          "Fn::Sub": "${PrometheusInstanceNamePrefix}-target"
        },
        "StorageDuration": {
          "Ref": "StorageDuration"
        },
        "ArchiveDuration": 0,
        "PaymentType": "POSTPAY",
        "EnableAuthFreeRead": true,
        "EnableAuthFreeWrite": false,
        "EnableAuthToken": true
      }
    },
    "AggTaskGroup": {
      "Type": "ALIYUN::CMS2::AggTaskGroup",
      "DependsOn": [
        "SourcePrometheus",
        "TargetPrometheus"
      ],
      "Properties": {
        "AggTaskGroupName": {
          "Ref": "AggTaskGroupName"
        },
        "InstanceId": {
          "Fn::GetAtt": [
            "SourcePrometheus",
            "PrometheusInstanceId"
          ]
        },
        "TargetPrometheusId": {
          "Fn::GetAtt": [
            "TargetPrometheus",
            "PrometheusInstanceId"
          ]
        },
        "ScheduleMode": "Cron",
        "CronExpr": {
          "Ref": "CronExpr"
        },
        "MaxRetries": {
          "Ref": "MaxRetries"
        },
        "MaxRunTimeInSeconds": {
          "Ref": "MaxRunTimeInSeconds"
        },
        "Delay": 60,
        "ScheduleTimeExpr": "@m",
        "OverrideIfExists": true,
        "Description": "Aggregation task group with Cron schedule and custom retry policy",
        "AggTaskGroupConfig": "groups:\n  - name: cpu_aggregation\n    rules:\n      - record: job:node_cpu_usage:avg_rate5m\n        expr: avg by (job) (rate(node_cpu_seconds_total{mode!=\"idle\"}[5m]))\n      - record: job:node_cpu_usage:max_rate5m\n        expr: max by (job) (rate(node_cpu_seconds_total{mode!=\"idle\"}[5m]))\n  - name: memory_aggregation\n    rules:\n      - record: job:node_memory_usage:avg\n        expr: avg by (job) (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes)\n      - record: job:node_memory_usage:percent\n        expr: avg by (job) ((node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / node_memory_MemTotal_bytes * 100)",
        "Tags": [
          {
            "Key": "Environment",
            "Value": "Production"
          },
          {
            "Key": "Team",
            "Value": "SRE"
          }
        ]
      }
    }
  },
  "Outputs": {
    "AggTaskGroupId": {
      "Description": {
        "zh-cn": "聚合任务组ID。",
        "en": "Agg task group ID."
      },
      "Label": {
        "zh-cn": "聚合任务组ID",
        "en": "Agg Task Group ID"
      },
      "Value": {
        "Fn::GetAtt": [
          "AggTaskGroup",
          "AggTaskGroupId"
        ]
      }
    },
    "SourcePrometheusId": {
      "Description": {
        "zh-cn": "源Prometheus实例ID。",
        "en": "Source Prometheus instance ID."
      },
      "Label": {
        "zh-cn": "源Prometheus实例ID",
        "en": "Source Prometheus ID"
      },
      "Value": {
        "Fn::GetAtt": [
          "SourcePrometheus",
          "PrometheusInstanceId"
        ]
      }
    },
    "TargetPrometheusId": {
      "Description": {
        "zh-cn": "目标Prometheus实例ID。",
        "en": "Target Prometheus instance ID."
      },
      "Label": {
        "zh-cn": "目标Prometheus实例ID",
        "en": "Target Prometheus ID"
      },
      "Value": {
        "Fn::GetAtt": [
          "TargetPrometheus",
          "PrometheusInstanceId"
        ]
      }
    },
    "AggTaskGroupStatus": {
      "Description": {
        "zh-cn": "聚合任务组状态。",
        "en": "Agg task group status."
      },
      "Label": {
        "zh-cn": "聚合任务组状态",
        "en": "Status"
      },
      "Value": {
        "Fn::GetAtt": [
          "AggTaskGroup",
          "Status"
        ]
      }
    }
  },
  "Metadata": {
    "ALIYUN::ROS::Interface": {
      "ParameterGroups": [
        {
          "Parameters": [
            "PrometheusInstanceNamePrefix",
            "StorageDuration"
          ]
        },
        {
          "Parameters": [
            "AggTaskGroupName",
            "CronExpr",
            "MaxRetries",
            "MaxRunTimeInSeconds"
          ]
        }
      ]
    }
  }
}

场景 3 :创建完整监控数据聚合方案

ROSTemplateFormatVersion: '2015-09-01'
Description:
  zh-cn: 创建完整的监控数据聚合方案,包含SLS日志项目、CMS2工作空间、多组Prometheus实例、多个聚合任务组(不同粒度的Recording Rule),实现多维度监控数据聚合与集中管理。
  en: Create a complete monitoring data aggregation solution including SLS log project, CMS2 workspace, multiple Prometheus instance groups, and multiple aggregation task groups with different granularity recording rules for multi-dimensional monitoring data aggregation and centralized management.
Parameters:
  CommonName:
    Type: String
    Label:
      zh-cn: 资源名称前缀
      en: Resource Name Prefix
    Description:
      zh-cn: 所有资源名称的统一前缀。
      en: Unified name prefix for all resources.
    Default: monitor-center
    MinLength: 1
    MaxLength: 50
  StorageDuration:
    Type: Number
    Label:
      zh-cn: 存储时长(天)
      en: Storage Duration (Days)
    Description:
      zh-cn: Prometheus实例的数据存储时长。
      en: Data storage duration of the Prometheus instance.
    AllowedValues:
      - 15
      - 30
      - 60
      - 90
      - 180
    Default: 90
  SourcePrometheusCount:
    Type: Number
    Label:
      zh-cn: 源Prometheus实例数量
      en: Source Prometheus Instance Count
    Description:
      zh-cn: 创建的源Prometheus实例数量。
      en: Number of source Prometheus instances to create.
    AllowedValues:
      - 1
      - 2
    Default: 2
Resources:
  SlsProject:
    Type: ALIYUN::SLS::Project
    Properties:
      Name:
        Fn::Sub: ${CommonName}-sls
      Description: SLS project for monitoring workspace
  Workspace:
    Type: ALIYUN::CMS2::Workspace
    Properties:
      WorkspaceName:
        Fn::Sub: ${CommonName}-ws
      DisplayName: Monitoring Center Workspace
      Description: Centralized monitoring workspace for metric aggregation
      SlsProject:
        Ref: SlsProject
  SourcePrometheusA:
    Type: ALIYUN::CMS2::PrometheusInstance
    Properties:
      PrometheusInstanceName:
        Fn::Sub: ${CommonName}-src-a
      StorageDuration:
        Ref: StorageDuration
      ArchiveDuration: 0
      PaymentType: POSTPAY
      EnableAuthFreeRead: true
      EnableAuthFreeWrite: false
      EnableAuthToken: true
      Workspace:
        Ref: Workspace
  SourcePrometheusB:
    Type: ALIYUN::CMS2::PrometheusInstance
    Properties:
      PrometheusInstanceName:
        Fn::Sub: ${CommonName}-src-b
      StorageDuration:
        Ref: StorageDuration
      ArchiveDuration: 0
      PaymentType: POSTPAY
      EnableAuthFreeRead: true
      EnableAuthFreeWrite: false
      EnableAuthToken: true
      Workspace:
        Ref: Workspace
  TargetPrometheus:
    Type: ALIYUN::CMS2::PrometheusInstance
    Properties:
      PrometheusInstanceName:
        Fn::Sub: ${CommonName}-target
      StorageDuration:
        Ref: StorageDuration
      ArchiveDuration: 365
      PaymentType: POSTPAY
      EnableAuthFreeRead: true
      EnableAuthFreeWrite: false
      EnableAuthToken: true
      Workspace:
        Ref: Workspace
  SystemMetricsAgg:
    Type: ALIYUN::CMS2::AggTaskGroup
    DependsOn:
      - SourcePrometheusA
      - TargetPrometheus
    Properties:
      AggTaskGroupName: system-metrics-agg
      InstanceId:
        Fn::GetAtt:
          - SourcePrometheusA
          - PrometheusInstanceId
      TargetPrometheusId:
        Fn::GetAtt:
          - TargetPrometheus
          - PrometheusInstanceId
      ScheduleMode: Cron
      CronExpr: 0/1 * * * *
      ScheduleTimeExpr: '@m'
      Delay: 30
      MaxRetries: 20
      MaxRunTimeInSeconds: 600
      OverrideIfExists: true
      Status: Running
      Description: Aggregate system-level metrics (CPU, memory, disk, network)
      AggTaskGroupConfig: |-
        groups:
          - name: cpu_rules
            interval: 1m
            rules:
              - record: job:node_cpu_usage:avg_rate5m
                expr: avg by (job) (rate(node_cpu_seconds_total{mode!="idle"}[5m]))
              - record: job:node_cpu_usage:max_rate5m
                expr: max by (job) (rate(node_cpu_seconds_total{mode!="idle"}[5m]))
              - record: instance:node_cpu_usage:rate5m
                expr: rate(node_cpu_seconds_total{mode!="idle"}[5m])
          - name: memory_rules
            interval: 1m
            rules:
              - record: job:node_memory_usage:avg
                expr: avg by (job) (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes)
              - record: job:node_memory_usage:percent
                expr: avg by (job) ((node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / node_memory_MemTotal_bytes * 100)
              - record: job:node_memory_available:avg
                expr: avg by (job) (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100)
          - name: disk_rules
            interval: 1m
            rules:
              - record: job:node_disk_usage:percent
                expr: avg by (job) ((node_filesystem_size_bytes - node_filesystem_avail_bytes) / node_filesystem_size_bytes * 100)
              - record: job:node_disk_io:rate5m
                expr: avg by (job) (rate(node_disk_io_time_seconds_total[5m]))
          - name: network_rules
            interval: 1m
            rules:
              - record: job:node_network_receive:rate5m
                expr: avg by (job) (rate(node_network_receive_bytes_total[5m]))
              - record: job:node_network_transmit:rate5m
                expr: avg by (job) (rate(node_network_transmit_bytes_total[5m]))
      Tags:
        - Key: Category
          Value: SystemMetrics
        - Key: Environment
          Value: Production
        - Key: Team
          Value: SRE
  AppMetricsAgg:
    Type: ALIYUN::CMS2::AggTaskGroup
    DependsOn:
      - SourcePrometheusB
      - TargetPrometheus
    Properties:
      AggTaskGroupName: app-metrics-agg
      InstanceId:
        Fn::GetAtt:
          - SourcePrometheusB
          - PrometheusInstanceId
      TargetPrometheusId:
        Fn::GetAtt:
          - TargetPrometheus
          - PrometheusInstanceId
      ScheduleMode: FixedRate
      Delay: 60
      MaxRetries: 30
      MaxRunTimeInSeconds: 900
      OverrideIfExists: true
      Status: Running
      Description: Aggregate application-level metrics (HTTP requests, latency, errors)
      AggTaskGroupConfig: |-
        groups:
          - name: http_rules
            interval: 1m
            rules:
              - record: job:http_requests:rate5m
                expr: sum by (job) (rate(http_requests_total[5m]))
              - record: job:http_requests:error_rate5m
                expr: sum by (job) (rate(http_requests_total{status=~"5.."}[5m])) / sum by (job) (rate(http_requests_total[5m])) * 100
              - record: job:http_request_duration:p99
                expr: histogram_quantile(0.99, sum by (job, le) (rate(http_request_duration_seconds_bucket[5m])))
              - record: job:http_request_duration:p95
                expr: histogram_quantile(0.95, sum by (job, le) (rate(http_request_duration_seconds_bucket[5m])))
              - record: job:http_request_duration:avg
                expr: avg by (job) (rate(http_request_duration_seconds_sum[5m]) / rate(http_request_duration_seconds_count[5m]))
          - name: availability_rules
            interval: 1m
            rules:
              - record: job:up:avg_rate5m
                expr: avg by (job) (avg_over_time(up[5m]))
              - record: job:up:min_rate5m
                expr: min by (job) (avg_over_time(up[5m]))
      Tags:
        - Key: Category
          Value: AppMetrics
        - Key: Environment
          Value: Production
        - Key: Team
          Value: DevOps
Outputs:
  WorkspaceName:
    Description:
      zh-cn: 工作空间名称。
      en: Workspace name.
    Label:
      zh-cn: 工作空间名称
      en: Workspace Name
    Value:
      Fn::GetAtt:
        - Workspace
        - WorkspaceName
  SystemMetricsAggId:
    Description:
      zh-cn: 系统指标聚合任务组ID。
      en: System metrics aggregation task group ID.
    Label:
      zh-cn: 系统指标聚合任务组ID
      en: System Metrics Agg ID
    Value:
      Fn::GetAtt:
        - SystemMetricsAgg
        - AggTaskGroupId
  AppMetricsAggId:
    Description:
      zh-cn: 应用指标聚合任务组ID。
      en: App metrics aggregation task group ID.
    Label:
      zh-cn: 应用指标聚合任务组ID
      en: App Metrics Agg ID
    Value:
      Fn::GetAtt:
        - AppMetricsAgg
        - AggTaskGroupId
  TargetPrometheusId:
    Description:
      zh-cn: 目标Prometheus实例ID。
      en: Target Prometheus instance ID.
    Label:
      zh-cn: 目标Prometheus实例ID
      en: Target Prometheus ID
    Value:
      Fn::GetAtt:
        - TargetPrometheus
        - PrometheusInstanceId
  TargetPrometheusRemoteWriteUrl:
    Description:
      zh-cn: 目标Prometheus实例Remote Write内网地址。
      en: Target Prometheus instance Remote Write internal URL.
    Label:
      zh-cn: 目标Prometheus Remote Write地址
      en: Target Prometheus Remote Write URL
    Value:
      Fn::GetAtt:
        - TargetPrometheus
        - RemoteWriteIntraUrl
  TargetPrometheusHttpApiUrl:
    Description:
      zh-cn: 目标Prometheus实例HTTP API内网地址。
      en: Target Prometheus instance HTTP API internal URL.
    Label:
      zh-cn: 目标Prometheus HTTP API地址
      en: Target Prometheus HTTP API URL
    Value:
      Fn::GetAtt:
        - TargetPrometheus
        - HttpApiIntraUrl
Metadata:
  ALIYUN::ROS::Interface:
    ParameterGroups:
      - Parameters:
          - CommonName
          - StorageDuration
          - SourcePrometheusCount
{
  "ROSTemplateFormatVersion": "2015-09-01",
  "Description": {
    "zh-cn": "创建完整的监控数据聚合方案,包含SLS日志项目、CMS2工作空间、多组Prometheus实例、多个聚合任务组(不同粒度的Recording Rule),实现多维度监控数据聚合与集中管理。",
    "en": "Create a complete monitoring data aggregation solution including SLS log project, CMS2 workspace, multiple Prometheus instance groups, and multiple aggregation task groups with different granularity recording rules for multi-dimensional monitoring data aggregation and centralized management."
  },
  "Parameters": {
    "CommonName": {
      "Type": "String",
      "Label": {
        "zh-cn": "资源名称前缀",
        "en": "Resource Name Prefix"
      },
      "Description": {
        "zh-cn": "所有资源名称的统一前缀。",
        "en": "Unified name prefix for all resources."
      },
      "Default": "monitor-center",
      "MinLength": 1,
      "MaxLength": 50
    },
    "StorageDuration": {
      "Type": "Number",
      "Label": {
        "zh-cn": "存储时长(天)",
        "en": "Storage Duration (Days)"
      },
      "Description": {
        "zh-cn": "Prometheus实例的数据存储时长。",
        "en": "Data storage duration of the Prometheus instance."
      },
      "AllowedValues": [
        15,
        30,
        60,
        90,
        180
      ],
      "Default": 90
    },
    "SourcePrometheusCount": {
      "Type": "Number",
      "Label": {
        "zh-cn": "源Prometheus实例数量",
        "en": "Source Prometheus Instance Count"
      },
      "Description": {
        "zh-cn": "创建的源Prometheus实例数量。",
        "en": "Number of source Prometheus instances to create."
      },
      "AllowedValues": [
        1,
        2
      ],
      "Default": 2
    }
  },
  "Resources": {
    "SlsProject": {
      "Type": "ALIYUN::SLS::Project",
      "Properties": {
        "Name": {
          "Fn::Sub": "${CommonName}-sls"
        },
        "Description": "SLS project for monitoring workspace"
      }
    },
    "Workspace": {
      "Type": "ALIYUN::CMS2::Workspace",
      "Properties": {
        "WorkspaceName": {
          "Fn::Sub": "${CommonName}-ws"
        },
        "DisplayName": "Monitoring Center Workspace",
        "Description": "Centralized monitoring workspace for metric aggregation",
        "SlsProject": {
          "Ref": "SlsProject"
        }
      }
    },
    "SourcePrometheusA": {
      "Type": "ALIYUN::CMS2::PrometheusInstance",
      "Properties": {
        "PrometheusInstanceName": {
          "Fn::Sub": "${CommonName}-src-a"
        },
        "StorageDuration": {
          "Ref": "StorageDuration"
        },
        "ArchiveDuration": 0,
        "PaymentType": "POSTPAY",
        "EnableAuthFreeRead": true,
        "EnableAuthFreeWrite": false,
        "EnableAuthToken": true,
        "Workspace": {
          "Ref": "Workspace"
        }
      }
    },
    "SourcePrometheusB": {
      "Type": "ALIYUN::CMS2::PrometheusInstance",
      "Properties": {
        "PrometheusInstanceName": {
          "Fn::Sub": "${CommonName}-src-b"
        },
        "StorageDuration": {
          "Ref": "StorageDuration"
        },
        "ArchiveDuration": 0,
        "PaymentType": "POSTPAY",
        "EnableAuthFreeRead": true,
        "EnableAuthFreeWrite": false,
        "EnableAuthToken": true,
        "Workspace": {
          "Ref": "Workspace"
        }
      }
    },
    "TargetPrometheus": {
      "Type": "ALIYUN::CMS2::PrometheusInstance",
      "Properties": {
        "PrometheusInstanceName": {
          "Fn::Sub": "${CommonName}-target"
        },
        "StorageDuration": {
          "Ref": "StorageDuration"
        },
        "ArchiveDuration": 365,
        "PaymentType": "POSTPAY",
        "EnableAuthFreeRead": true,
        "EnableAuthFreeWrite": false,
        "EnableAuthToken": true,
        "Workspace": {
          "Ref": "Workspace"
        }
      }
    },
    "SystemMetricsAgg": {
      "Type": "ALIYUN::CMS2::AggTaskGroup",
      "DependsOn": [
        "SourcePrometheusA",
        "TargetPrometheus"
      ],
      "Properties": {
        "AggTaskGroupName": "system-metrics-agg",
        "InstanceId": {
          "Fn::GetAtt": [
            "SourcePrometheusA",
            "PrometheusInstanceId"
          ]
        },
        "TargetPrometheusId": {
          "Fn::GetAtt": [
            "TargetPrometheus",
            "PrometheusInstanceId"
          ]
        },
        "ScheduleMode": "Cron",
        "CronExpr": "0/1 * * * *",
        "ScheduleTimeExpr": "@m",
        "Delay": 30,
        "MaxRetries": 20,
        "MaxRunTimeInSeconds": 600,
        "OverrideIfExists": true,
        "Status": "Running",
        "Description": "Aggregate system-level metrics (CPU, memory, disk, network)",
        "AggTaskGroupConfig": "groups:\n  - name: cpu_rules\n    interval: 1m\n    rules:\n      - record: job:node_cpu_usage:avg_rate5m\n        expr: avg by (job) (rate(node_cpu_seconds_total{mode!=\"idle\"}[5m]))\n      - record: job:node_cpu_usage:max_rate5m\n        expr: max by (job) (rate(node_cpu_seconds_total{mode!=\"idle\"}[5m]))\n      - record: instance:node_cpu_usage:rate5m\n        expr: rate(node_cpu_seconds_total{mode!=\"idle\"}[5m])\n  - name: memory_rules\n    interval: 1m\n    rules:\n      - record: job:node_memory_usage:avg\n        expr: avg by (job) (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes)\n      - record: job:node_memory_usage:percent\n        expr: avg by (job) ((node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / node_memory_MemTotal_bytes * 100)\n      - record: job:node_memory_available:avg\n        expr: avg by (job) (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100)\n  - name: disk_rules\n    interval: 1m\n    rules:\n      - record: job:node_disk_usage:percent\n        expr: avg by (job) ((node_filesystem_size_bytes - node_filesystem_avail_bytes) / node_filesystem_size_bytes * 100)\n      - record: job:node_disk_io:rate5m\n        expr: avg by (job) (rate(node_disk_io_time_seconds_total[5m]))\n  - name: network_rules\n    interval: 1m\n    rules:\n      - record: job:node_network_receive:rate5m\n        expr: avg by (job) (rate(node_network_receive_bytes_total[5m]))\n      - record: job:node_network_transmit:rate5m\n        expr: avg by (job) (rate(node_network_transmit_bytes_total[5m]))",
        "Tags": [
          {
            "Key": "Category",
            "Value": "SystemMetrics"
          },
          {
            "Key": "Environment",
            "Value": "Production"
          },
          {
            "Key": "Team",
            "Value": "SRE"
          }
        ]
      }
    },
    "AppMetricsAgg": {
      "Type": "ALIYUN::CMS2::AggTaskGroup",
      "DependsOn": [
        "SourcePrometheusB",
        "TargetPrometheus"
      ],
      "Properties": {
        "AggTaskGroupName": "app-metrics-agg",
        "InstanceId": {
          "Fn::GetAtt": [
            "SourcePrometheusB",
            "PrometheusInstanceId"
          ]
        },
        "TargetPrometheusId": {
          "Fn::GetAtt": [
            "TargetPrometheus",
            "PrometheusInstanceId"
          ]
        },
        "ScheduleMode": "FixedRate",
        "Delay": 60,
        "MaxRetries": 30,
        "MaxRunTimeInSeconds": 900,
        "OverrideIfExists": true,
        "Status": "Running",
        "Description": "Aggregate application-level metrics (HTTP requests, latency, errors)",
        "AggTaskGroupConfig": "groups:\n  - name: http_rules\n    interval: 1m\n    rules:\n      - record: job:http_requests:rate5m\n        expr: sum by (job) (rate(http_requests_total[5m]))\n      - record: job:http_requests:error_rate5m\n        expr: sum by (job) (rate(http_requests_total{status=~\"5..\"}[5m])) / sum by (job) (rate(http_requests_total[5m])) * 100\n      - record: job:http_request_duration:p99\n        expr: histogram_quantile(0.99, sum by (job, le) (rate(http_request_duration_seconds_bucket[5m])))\n      - record: job:http_request_duration:p95\n        expr: histogram_quantile(0.95, sum by (job, le) (rate(http_request_duration_seconds_bucket[5m])))\n      - record: job:http_request_duration:avg\n        expr: avg by (job) (rate(http_request_duration_seconds_sum[5m]) / rate(http_request_duration_seconds_count[5m]))\n  - name: availability_rules\n    interval: 1m\n    rules:\n      - record: job:up:avg_rate5m\n        expr: avg by (job) (avg_over_time(up[5m]))\n      - record: job:up:min_rate5m\n        expr: min by (job) (avg_over_time(up[5m]))",
        "Tags": [
          {
            "Key": "Category",
            "Value": "AppMetrics"
          },
          {
            "Key": "Environment",
            "Value": "Production"
          },
          {
            "Key": "Team",
            "Value": "DevOps"
          }
        ]
      }
    }
  },
  "Outputs": {
    "WorkspaceName": {
      "Description": {
        "zh-cn": "工作空间名称。",
        "en": "Workspace name."
      },
      "Label": {
        "zh-cn": "工作空间名称",
        "en": "Workspace Name"
      },
      "Value": {
        "Fn::GetAtt": [
          "Workspace",
          "WorkspaceName"
        ]
      }
    },
    "SystemMetricsAggId": {
      "Description": {
        "zh-cn": "系统指标聚合任务组ID。",
        "en": "System metrics aggregation task group ID."
      },
      "Label": {
        "zh-cn": "系统指标聚合任务组ID",
        "en": "System Metrics Agg ID"
      },
      "Value": {
        "Fn::GetAtt": [
          "SystemMetricsAgg",
          "AggTaskGroupId"
        ]
      }
    },
    "AppMetricsAggId": {
      "Description": {
        "zh-cn": "应用指标聚合任务组ID。",
        "en": "App metrics aggregation task group ID."
      },
      "Label": {
        "zh-cn": "应用指标聚合任务组ID",
        "en": "App Metrics Agg ID"
      },
      "Value": {
        "Fn::GetAtt": [
          "AppMetricsAgg",
          "AggTaskGroupId"
        ]
      }
    },
    "TargetPrometheusId": {
      "Description": {
        "zh-cn": "目标Prometheus实例ID。",
        "en": "Target Prometheus instance ID."
      },
      "Label": {
        "zh-cn": "目标Prometheus实例ID",
        "en": "Target Prometheus ID"
      },
      "Value": {
        "Fn::GetAtt": [
          "TargetPrometheus",
          "PrometheusInstanceId"
        ]
      }
    },
    "TargetPrometheusRemoteWriteUrl": {
      "Description": {
        "zh-cn": "目标Prometheus实例Remote Write内网地址。",
        "en": "Target Prometheus instance Remote Write internal URL."
      },
      "Label": {
        "zh-cn": "目标Prometheus Remote Write地址",
        "en": "Target Prometheus Remote Write URL"
      },
      "Value": {
        "Fn::GetAtt": [
          "TargetPrometheus",
          "RemoteWriteIntraUrl"
        ]
      }
    },
    "TargetPrometheusHttpApiUrl": {
      "Description": {
        "zh-cn": "目标Prometheus实例HTTP API内网地址。",
        "en": "Target Prometheus instance HTTP API internal URL."
      },
      "Label": {
        "zh-cn": "目标Prometheus HTTP API地址",
        "en": "Target Prometheus HTTP API URL"
      },
      "Value": {
        "Fn::GetAtt": [
          "TargetPrometheus",
          "HttpApiIntraUrl"
        ]
      }
    }
  },
  "Metadata": {
    "ALIYUN::ROS::Interface": {
      "ParameterGroups": [
        {
          "Parameters": [
            "CommonName",
            "StorageDuration",
            "SourcePrometheusCount"
          ]
        }
      ]
    }
  }
}