Advanced O&M

更新时间:
复制 MD 格式

The TRaaS Technology Risk Control Platform is a platform developed from the long-term practical methodologies and internal tools of Ant Group's Site Reliability Engineering (SRE) team. It is designed to solve O&M challenges that arise during cloud migration and distributed transformation, such as observability, incident response, disaster recovery, chaos engineering, fund security, and stress testing.

image

High Availability Management Platform

High Availability Service (HAS) is a management and control platform that focuses on high availability and disaster recovery. It provides end-to-end capabilities for disaster recovery plans, including switchover and recovery, planning, and simulation drills for your entire stack, from customer services to middleware, Platform-as-a-Service (PaaS), and Infrastructure-as-a-Service (IaaS). HAS also provides monitoring of the overall data center and disaster recovery status, a disaster recovery dashboard, environment inspections, and risk response capabilities.

HAS provides a disaster recovery service view, plan orchestration, and switchover and recovery capabilities. It supports one-click, data center-level disaster recovery switchover and recovery in a multi-data center deployment architecture.

Service architecture

BB2DA5~1

Benefits

  • Ant's Technical Risk Management: Methodology and Platform Tools in Practice

    The HAS platform leverages Ant Group's extensive experience in technology risk control. It helps you build a technology risk control system that fits your specific needs and improves your overall risk control capabilities.

  • Improved efficiency in technology risk control

    HAS automates and standardizes daily O&M. This reduces operational complexity, clarifies O&M results, and enables closed-loop management of risk events.

    • Proactively detect service operation risks through routine inspections and address them before they impact your services.

    • Automated fault diagnosis and standardized emergency plans help you quickly locate and recover from faults, which reduces service downtime.

    • Fault drills proactively test the high availability of your applications.

    • Supports dual data center disaster recovery switchovers for Ant Group products to meet regulatory compliance requirements.

  • Rapid updates to the technology risk control content library

    The Alibaba Cloud and Ant Group technology risk teams jointly maintain a content library for routine inspections, fault diagnosis, and emergency plans. This library is based on internal and external risk control experience and provides you with the latest content for technology risk control.

  • Financial-grade disaster recovery

    Achieves a disaster recovery rating of up to Level 5.

    Provides comprehensive disaster recovery capabilities, including a disaster recovery dashboard with monitoring and alerts, simulation drills, and inspections.

    Validated at scale by Alipay and MYbank.

  • End-to-end disaster recovery

    Supports end-to-end disaster recovery capabilities from customer applications to core services. This allows for complete disaster recovery, monitoring, and O&M without requiring multi-platform integration. The service provides end-to-end, multi-layer disaster recovery for user applications, middleware, PaaS, and IaaS.

  • Multi-scenario disaster recovery

    Supports all disaster recovery scenarios in the finance industry:

    • Active-active in the same city

    • Active-standby across regions

    • Three data centers across two regions

    • LDC unitization

Scenarios

Daily risk control

In daily O&M scenarios, use the interconnected multi-functional modules to automate routine O&M scripts. This achieves regular and controllable routine inspections. The content for routine inspections, fault diagnosis, and emergency plans is continuously updated and optimized. This enriches and improves the application technology risk control system and simplifies daily application O&M.

Fault drills

To continuously improve product high availability, use the fault drill module in HAS to design and plan drill and recovery plans. During the drills, you can discover and resolve issues in your disaster recovery plans. This reduces the probability of faults during product use, improves fault recovery efficiency, and effectively enhances product high availability.

Data-center-level disaster recovery

  • Active-active in the same city: Two data centers are built in the same city, less than 50 km apart. They are interconnected by 10 Gbps dedicated optical fiber leased lines. At the application layer, both data centers can provide services simultaneously. If one data center fails, services in the other data center are not affected.

  • Active-standby across regions: To meet disaster recovery needs, two data centers are built in different cities over 1000 km apart. One is the primary data center and the other is the secondary data center. The primary data center handles service traffic, while the secondary data center carries no service traffic and serves only as a backup. If the primary data center fails, you can switch traffic to the secondary data center to quickly recover services. After the primary data center recovers, you can fail back the traffic.

  • Three data centers across two regions: This architecture, also known as an active-active same-city plus active-standby cross-region solution, involves two data centers in the same city deployed in an active-active configuration. An additional data center in a different region serves only as a backup. It does not carry any service traffic and is mainly used for cold data backup. This design provides the highest level of high availability for data backups.

  • LDC unitization (active geo-redundancy): The LDC unitization architecture enables active geo-redundancy and high-concurrency scenarios. LDC stands for Logical Data Center, a concept proposed in contrast to the traditional Internet Data Center (IDC). The core idea of an LDC is that the entire data center is logically coordinated and unified, regardless of its physical distribution. This architecture is mainly suitable for supporting the online transaction systems of large Internet companies, such as Taobao, Alipay, and Ctrip.

End-to-end Stress Testing

End-to-end Stress Testing (Loadcenter) is a one-stop stress testing service for enterprises that covers performance stress testing, report generation, and risk control. Based on Ant Group's years of experience with online end-to-end stress testing, Loadcenter provides a high-fidelity, low-cost online stress testing experience with effective risk identification.

Service architectureImage 47

Benefits

  • Complex scenario modeling capabilities

    • Supports multiple traffic models, allowing you to quickly import and configure traffic.

    • Supports both templated and custom-developed scripts to meet the stress testing needs of business scenarios with varying complexity.

  • Powerful report analysis capabilities

    • Centrally manage and archive stress test records.

    • Combined with the real-time monitoring service, stress test results include standard application monitoring data and custom business monitoring data. This helps you quickly identify bottlenecked applications and related performance metrics.

    • Use comparative report analysis to track the evolution of application performance baselines.

  • Stable pressure output

    Horizontally scale load generators and dynamically adjust the load in seconds. This meets stress testing demands of up to tens of millions of Transactions Per Second (TPS).

  • Support for stress testing internal network interfaces

    The load generator resource pool supports public and private tenant modes:

    • In public mode, you do not need to provide your own load generators and can run stress tests at any time.

    • In private mode, you can use your own load generators to save bandwidth costs and reduce network latency. This mode also supports internal interface-level testing without exposure to the public network, which provides better security.

  • Support for multiple protocols

    • Supports standard protocols such as HTTP, HTTPS, and SOFARPC.

    • Provides a custom script development mode based on the Java language. You can extend proprietary protocols on your own.

  • Reliable risk control for production stress testing

    • Integrates with multiple FinTech products, combining application monitoring, business monitoring, and O&M capabilities. If a risk is detected, the stress test can be stopped automatically.

    • When combined with the FinTech SOFA middleware product, you can use the shadow end-to-end stress testing solution. This solution isolates stress traffic from normal traffic, which lets you run stress tests in a production environment.

Scenarios

End-to-end Stress Testing is suitable for any application scenario that requires stress testing or simulated traffic.

New system launch testing

Before a new system goes live, perform stress and load testing on the system based on the expected business model. Test whether the system has performance issues and whether the expected capacity can handle the service pressure after launch.

Existing system baseline regression

Periodically perform performance regression tests on the online system in a constant scenario. Observe whether the system's performance has changed. Promptly detect performance degradation caused by iterations or technology upgrades.

System capacity assessment

Before launching promotional activities, assess the system capacity through stress testing. Continuously increase the pressure based on the business scenario to evaluate the system's capacity level. This allows for early optimization and scaling. If you have traffic limiting measures, you can also validate them through stress testing.

System fault drills

Use continuous stress traffic to verify whether services are affected when the system is abnormal. You can use stress traffic in combination with fault injection drills and data center disaster recovery drills to observe the extent of service impact and the recovery capability.

Fund Security Monitoring

The Fund Security Monitoring platform is a real-time verification platform that uses a bypass method to analyze the flow of funds in business processes and provide real-time alerts. It ensures fund security at a technical level and prevents fund loss within business systems.

Service architectureImage 76

Benefits

  • The platform is non-intrusive to production systems because it collects verification data through a bypass.

  • Rules are configurable and require no coding. You can add or modify rules at any time to meet various verification needs.

  • Supports multiple verification schedules, such as real-time, near real-time, T+1, and T+H. This meets your different timeliness requirements for fund loss risk monitoring.

  • Provides comprehensive management features, including a verification dashboard and coverage measurement capabilities.

  • Supports notification channels such as text message, email, and DingTalk. This provides instant monitoring and emergency response for core services.

  • Provides measurement features for fund loss risk monitoring coverage. It also includes expert consulting services that provide cloud users with years of accumulated fund loss prevention experience.

Scenarios

Service protection

Helps you periodically or regularly review core services that involve fund links. Configure verification rules to cross-check various types of data or perform logical checks on data content. This ensures that core services run correctly.

Change risk inspection

Before releasing a change, add verification rules for the changed business tables and their associated tables. You can also add check rules for the data in the changed business tables. This ensures there are no blind spots in fund loss risk monitoring after the change goes live.

Historical data audit

Perform batch checks on the historical data of existing services to see if discrepancies already exist. Promptly analyze the causes of discrepancies, fix vulnerabilities, and recover losses.

Data quality monitoring

Missing or incomplete data can also indirectly cause fund loss. You can configure verification rules to check data integrity, monitor data quality, and promptly detect faults.