Daily O&M management
O&M ensures the continuous, stable, and secure operation of a cloud platform. The process involves meeting cloud infrastructure demands, monitoring and creating alerts for cloud applications, promptly identifying and resolving issues, and performing continuous optimization. These tasks are the primary daily work of a cloud management team.
A Cloud Center of Excellence (CCoE) adapts a company's existing O&M management process to the three unique characteristics of cloud O&M:
Different O&M objects: Cloud infrastructure O&M no longer involves physical devices. The O&M interface shifts from various device-specific interfaces to the cloud's unified management platform.
Different O&M methods: Compared to traditional infrastructure, cloud platforms provide a rich set of OpenAPI, which enables O&M teams to use more automation for O&M tasks.
Different optimization methods: Because cloud infrastructure provides resources and capabilities as services, it offers more methods for O&M and continuous optimization, such as purchasing cloud-native O&M tools, buying managed services, and following best practices.
Requirements management and resource provisioning
The elasticity of a cloud platform makes resource provisioning more convenient and enables agile responses to various needs. However, companies still need to standardize the demand for and provisioning of cloud resources. Companies can manage cloud resource demand and provisioning in two ways:
Platform-operated model. In this model, the cloud management team grants resource provisioning permissions to the requesters, allowing them to provision resources directly. Management is performed after provisioning.
O&M-managed model. In this model, the cloud management team's work is similar to traditional O&M. They review formal requests from requesters and then provision the resources.
The first model can easily cause the architecture, which is built on the landing zone and the sound cloud principles pursued by the CCoE, to become uncontrolled. This can prevent the full achievement of cloud strategic goals. However, it fully leverages the elasticity of cloud resources, which allows requesters to meet their resource needs quickly. The second model sacrifices the elasticity and agility of cloud resources for direct users in exchange for standardized management.
A basic strategy that balances both models for different situations is generally recommended:
Use the first model for non-sensitive environments, such as development and testing.
Use the second model for production environments or sensitive applications.
To avoid technical debt, review cloud resources for compliance. This can be done after a request is made in the second model or before an application goes online in the first model.
Create templates for commonly used cloud resources and provision resources based on these templates.
If you have the technical capability, you can use infrastructure automation for systematic management.
Regardless of the model used, the O&M team must perform management actions that align with the landing zone. Even with the first model, post-provisioning supervision and management are required.
Compared to traditional resources, the O&M team also needs to consider how to fully leverage the following characteristics of cloud resources:
The elasticity of cloud resources simplifies the adjustment of environment configurations. Therefore, adjusting configurations should be part of the O&M service.
Most cloud resources are out-of-the-box. Do not hoard resources. Regularly remove idle resources.
For small-scale cloud resource needs that deviate from the landing zone and other rules, you can allow for innovation if costs are controllable and the external impact is limited. This approach can serve as a basis for the continuous introduction of new technologies.
Monitoring and observability
Monitoring and observability, along with related alerting, are key to most daily O&M tasks. Key cloud platform capabilities are divided into two parts:
Infrastructure monitoring: This involves using the monitoring capabilities provided by the cloud platform to monitor IT infrastructure and network quality. This includes business monitoring based on events, custom metrics, and logs. These capabilities provide companies with more efficient and comprehensive monitoring services to promptly detect faults, increase system service uptime, and reduce IT O&M monitoring costs. On Alibaba Cloud, this corresponds to the Cloud Monitor product.
Application monitoring: For cloud-native applications, this involves implementing full-stack performance monitoring and end-to-end trace diagnostics to improve monitoring efficiency and reduce the O&M workload. It covers various observable environments and scenarios, such as browsers, miniapps, mobile apps, distributed applications, and containers, for both application monitoring and Real User Monitoring (RUM). On Alibaba Cloud, this corresponds to the Application Real-Time Monitoring Service (ARMS) product.
The cloud management team should prioritize implementing infrastructure monitoring that covers all cloud resources. When a company uses cloud computing on a large scale, application monitoring also becomes very important. It helps the company establish unified application monitoring for cloud-native applications.
Inspection
An inspection is a periodic review of the cloud architecture and resource status. It is based on monitoring and observability data over a period, combined with various events that occurred during that time. The CCoE uses the inspection process to focus on sound cloud architecture. It reviews technology adoption, landing zone baselines, technology adoption and expansion, and architectural debt. The O&M management team uses the inspection process to focus on the state of cloud resources. It reviews cloud resources, security, permissions, and costs. The goal of the review is to provide input for threat and problem identification, protection, and administration.
The O&M management team's inspection process is typically comprehensive. A company's IT department needs to establish standards that cover all aspects, including infrastructure, technical frameworks, and applications. The inspection of cloud resources should also cover the following areas:
Security
The inspection should check whether newly provisioned or modified resources, and their associated north-south and internal traffic protection, deviate from the security protection baseline and schedule fixes if necessary.
Verifies whether the permission configurations for cloud resources and the identity permissions for accounts meet security protection baseline requirements.
The inspection should determine if security products and measures that were not adopted for various reasons are related to security events that occurred during the inspection period. It should also assess whether new security products and technologies are needed for protection.
It is important to pay close attention to and resolve alerts from Security Center, which provides a closed loop of automated security operations, including threat detection, response, and source tracing.
Resources
Exporting the cloud resource checklist and comparing it with the one from the last inspection helps verify that actual resource changes are consistent with change requests. This is especially important when requesters have self-service permissions for resource provisioning and changes.
The inspection should assess the business continuity situation during the inspection period. It should also evaluate the stability of cloud resources and review the high availability (HA) architecture.
The inspection should identify the root causes of anomalous activity from monitoring, such as high SLB latency or 5xx errors. Such events may not immediately affect business but pose a risk of failure.
The inspection should also check for non-anomalous events to find opportunities for resource optimization, such as high or low ECS load. The findings help with continuous optimization.
Monitoring and automation
If a problem found during an inspection is determined to be a false positive or false negative caused by missing monitoring or an incorrect configuration, you should adjust the metrics and thresholds.
If common patterns are found in the problems discovered during inspections, you can consider developing automated scripts and applications for continuous tracking and optimization.
Optimization
Based on monitoring and observability actions, combined with inspections, the CCoE focuses on architectural optimization. This work involves changes to the landing zone baseline and the continuous introduction of technology to achieve cloud strategy goals.
The cloud management team focuses on optimizing configuration baselines, security, stability, and cost through risk detection, protection, administration, and issue handling to ultimately enhance infrastructure automation.
Problems found in post-event inspections should drive process optimization for cloud resource planning, provisioning, and O&M.
Fees and costs
Because cloud resources provide better digital support for cost management and optimization, the inspection and resource reconciliation processes can be integrated. This helps better evaluate the investment in the cloud strategy.
Using methods such as tags, cloud resources can provide specific cost information for each application on the cloud. This information helps the CCoE and business teams evaluate the ROI of the business.
The items above do not cover all aspects of an inspection. Whether it is Infrastructure as a Service (IaaS) such as compute, network, and storage, or Platform as a Service (PaaS) such as databases, middleware, and big data, inspections must include the application, architecture, and resource aspects. However, inspecting the capacity and configuration of cloud resources is a primary responsibility of the O&M management team. Most inspection work can be done by collecting and analyzing data using the API tools provided by the cloud platform. For many companies, the first step in moving big data to the cloud is moving their O&M big data to the cloud.
Backup
Backup is the ultimate guarantee of stability. In the recovery process after a failure, it ensures minimal data loss in the production environment.
The main objects for backup are the structured and unstructured data that a company needs to store persistently. This includes not only business data but also IT data related to operations and maintenance, such as applications and configuration information. Backup is a routine O&M task. Cloud technology reduces the difficulty and cost of backups, which helps O&M management provide more robust backup capabilities for the company.
You should evaluate any cloud resource that stores persistent company data for backup needs. This involves products such as compute, storage, middleware, databases, and big data. Specific backup solutions and implementation steps are described in product documents and best practices. Backup methods vary by product. The O&M management team needs to establish separate policies and baselines for backups and establish and rehearse procedures for backup and recovery.
Storage and database backups are already very familiar to traditional IT O&M. Some cloud technologies help O&M management teams improve backups for hybrid cloud scenarios:
Database Backup (DBS) provides robust protection for databases in various environments, including data centers, other cloud vendors, public clouds, and hybrid clouds.
Hybrid Backup Recovery (HBR) provides backup, disaster recovery, and policy-based archive management for Alibaba Cloud ECS instances, ECS databases, file systems, NAS, OSS, and Tablestore. It also supports files, databases, virtual machines, and large-scale NAS in on-premises data centers.
The core product for data backup on the cloud platform is a cloud storage service, such as Alibaba Cloud Object Storage Service (OSS). Storage services usually offer multiple storage classes with different features and costs. When creating a backup policy, you can use low-cost storage classes to meet non-critical needs and reduce backup costs.
External O&M resources
Cloud service providers offer expert services related to cloud O&M. To ensure that cloud systems run stably and efficiently and to handle business peaks, they provide O&M management services such as architecture checks for various cloud products. These services include the following:

As a partner of the cloud provider, a Managed Service Provider (MSP) can also deliver products or services for a company. They can design, architect, build, migrate, and manage a customer's workloads and applications on the cloud.
The services provided by an MSP can help a company:
Fill skill or personnel gaps in the company's cloud management team through on-site and personalized services. Some companies outsource their entire cloud management work to an MSP.
Provide a faster response. They can be on-site as soon as a problem occurs to quickly locate the issue and resolve most common problems.
Provide professional O&M management and continuous optimization suggestions based on a deep understanding of the company's cloud strategy and application status. This helps align cloud technology more closely with the company's needs.
For cross-platform scenarios, they provide hybrid or multicloud management platforms and services. This helps companies establish one-stop infrastructure O&M and management.
O&M management tools
A company's daily O&M requires a platform for managing and operating its business continuity. This platform should have features such as monitoring integration, alert denoising, event notification and routing, and ITIL-based incident management. It helps a company achieve real-time digital management, faster fault response, shorter fault duration, and improved business continuity.

An O&M event center platform supports these needs by providing the capabilities shown in the figure above:
Multi-system monitoring integration: Supports integration with more than 10 common monitoring systems. Integration can be completed quickly with simple configuration.
Flexible alert denoising: Supports horizontal suppression and vertical convergence to fully control alert storms, ensuring no critical alerts are missed.
Greatly reduced transactional operations: A complete event assignment and notification mechanism avoids repetitive transactional operations and improves O&M efficiency.
Based on Alibaba's incident management best practices: Helps companies on the cloud build an incident management system and continuously improve business continuity.
The platform's practical use cases fall into two main categories:
1. One-stop O&M event management
In this scenario, the O&M event center platform can meet the need for unified event-based management of alerts from various monitoring scenarios. It supports integration with various monitoring systems and allows servers to push anomalous activity reports. It provides end-to-end, one-stop management of alerts, events, and incidents, which improves a company's O&M efficiency.

In the one-stop O&M event management scenario, the O&M event center platform can help a company address the following typical problems:
Multi-source monitoring integration: Integration with multiple common monitoring systems can be completed with simple configuration.
Unified alert processing: Centrally denoise alerts with suppression and convergence to avoid alert storms.
Closed-loop event management: Manage the full lifecycle of events generated from alerts to ensure no events are missed.
2. Systematic closed-loop incident management
The O&M event center platform meets a company's need for process-based, online management of major incidents, continuously improving business continuity.

In the systematic closed-loop incident management scenario, the O&M event center platform can help a company address the following typical problems:
Incident response: Sending global emergency notifications for incidents through multiple channels such as phone calls, text messages, emails, and instant messaging to speed up information flow.
Incident tracking: Managing and collaborating online on incident progress, impact scope, public sentiment feedback, and timelines to improve incident handling efficiency.
Post-incident review: Based on best practices, this feature establishes structured requirements for in-depth incident reviews. It creates online checkpoints and implements the process as a product feature.
Incident improvement: Defining clear improvement and acceptance measures, owners, and completion times for each incident. This ensures that every in-depth post-incident review leads to an improvement in business continuity.