The TRaaS Technology Risk Control Platform is a platform developed from the long-term practical methodologies and internal tools of Ant Group's Site Reliability Engineering (SRE) team. It is designed to solve O&M challenges that arise during cloud migration and distributed transformation, such as observability, incident response, disaster recovery, chaos engineering, fund security, and stress testing.

High Availability Management Platform
High Availability Service (HAS) is a management and control platform that focuses on high availability and disaster recovery. It provides end-to-end capabilities for disaster recovery plans, including switchover and recovery, planning, and simulation drills for your entire stack, from customer services to middleware, Platform-as-a-Service (PaaS), and Infrastructure-as-a-Service (IaaS). HAS also provides monitoring of the overall data center and disaster recovery status, a disaster recovery dashboard, environment inspections, and risk response capabilities.
HAS provides a disaster recovery service view, plan orchestration, and switchover and recovery capabilities. It supports one-click, data center-level disaster recovery switchover and recovery in a multi-data center deployment architecture.
Service architecture
Benefits
Ant's Technical Risk Management: Methodology and Platform Tools in Practice
The HAS platform leverages Ant Group's extensive experience in technology risk control. It helps you build a technology risk control system that fits your specific needs and improves your overall risk control capabilities.
Improved efficiency in technology risk control
HAS automates and standardizes daily O&M. This reduces operational complexity, clarifies O&M results, and enables closed-loop management of risk events.
Proactively detect service operation risks through routine inspections and address them before they impact your services.
Automated fault diagnosis and standardized emergency plans help you quickly locate and recover from faults, which reduces service downtime.
Fault drills proactively test the high availability of your applications.
Supports dual data center disaster recovery switchovers for Ant Group products to meet regulatory compliance requirements.
Rapid updates to the technology risk control content library
The Alibaba Cloud and Ant Group technology risk teams jointly maintain a content library for routine inspections, fault diagnosis, and emergency plans. This library is based on internal and external risk control experience and provides you with the latest content for technology risk control.
Financial-grade disaster recovery
Achieves a disaster recovery rating of up to Level 5.
Provides comprehensive disaster recovery capabilities, including a disaster recovery dashboard with monitoring and alerts, simulation drills, and inspections.
Validated at scale by Alipay and MYbank.
End-to-end disaster recovery
Supports end-to-end disaster recovery capabilities from customer applications to core services. This allows for complete disaster recovery, monitoring, and O&M without requiring multi-platform integration. The service provides end-to-end, multi-layer disaster recovery for user applications, middleware, PaaS, and IaaS.
Multi-scenario disaster recovery
Supports all disaster recovery scenarios in the finance industry:
Active-active in the same city
Active-standby across regions
Three data centers across two regions
LDC unitization
Scenarios
Daily risk control |
In daily O&M scenarios, use the interconnected multi-functional modules to automate routine O&M scripts. This achieves regular and controllable routine inspections. The content for routine inspections, fault diagnosis, and emergency plans is continuously updated and optimized. This enriches and improves the application technology risk control system and simplifies daily application O&M. |
Fault drills |
To continuously improve product high availability, use the fault drill module in HAS to design and plan drill and recovery plans. During the drills, you can discover and resolve issues in your disaster recovery plans. This reduces the probability of faults during product use, improves fault recovery efficiency, and effectively enhances product high availability. |
Data-center-level disaster recovery |
|
End-to-end Stress Testing
End-to-end Stress Testing (Loadcenter) is a one-stop stress testing service for enterprises that covers performance stress testing, report generation, and risk control. Based on Ant Group's years of experience with online end-to-end stress testing, Loadcenter provides a high-fidelity, low-cost online stress testing experience with effective risk identification.
Service architecture
Benefits
Complex scenario modeling capabilities
Supports multiple traffic models, allowing you to quickly import and configure traffic.
Supports both templated and custom-developed scripts to meet the stress testing needs of business scenarios with varying complexity.
Powerful report analysis capabilities
Centrally manage and archive stress test records.
Combined with the real-time monitoring service, stress test results include standard application monitoring data and custom business monitoring data. This helps you quickly identify bottlenecked applications and related performance metrics.
Use comparative report analysis to track the evolution of application performance baselines.
Stable pressure output
Horizontally scale load generators and dynamically adjust the load in seconds. This meets stress testing demands of up to tens of millions of Transactions Per Second (TPS).
Support for stress testing internal network interfaces
The load generator resource pool supports public and private tenant modes:
In public mode, you do not need to provide your own load generators and can run stress tests at any time.
In private mode, you can use your own load generators to save bandwidth costs and reduce network latency. This mode also supports internal interface-level testing without exposure to the public network, which provides better security.
Support for multiple protocols
Supports standard protocols such as HTTP, HTTPS, and SOFARPC.
Provides a custom script development mode based on the Java language. You can extend proprietary protocols on your own.
Reliable risk control for production stress testing
Integrates with multiple FinTech products, combining application monitoring, business monitoring, and O&M capabilities. If a risk is detected, the stress test can be stopped automatically.
When combined with the FinTech SOFA middleware product, you can use the shadow end-to-end stress testing solution. This solution isolates stress traffic from normal traffic, which lets you run stress tests in a production environment.
Scenarios
End-to-end Stress Testing is suitable for any application scenario that requires stress testing or simulated traffic.
New system launch testing |
Before a new system goes live, perform stress and load testing on the system based on the expected business model. Test whether the system has performance issues and whether the expected capacity can handle the service pressure after launch. |
Existing system baseline regression |
Periodically perform performance regression tests on the online system in a constant scenario. Observe whether the system's performance has changed. Promptly detect performance degradation caused by iterations or technology upgrades. |
System capacity assessment |
Before launching promotional activities, assess the system capacity through stress testing. Continuously increase the pressure based on the business scenario to evaluate the system's capacity level. This allows for early optimization and scaling. If you have traffic limiting measures, you can also validate them through stress testing. |
System fault drills |
Use continuous stress traffic to verify whether services are affected when the system is abnormal. You can use stress traffic in combination with fault injection drills and data center disaster recovery drills to observe the extent of service impact and the recovery capability. |
Fund Security Monitoring
The Fund Security Monitoring platform is a real-time verification platform that uses a bypass method to analyze the flow of funds in business processes and provide real-time alerts. It ensures fund security at a technical level and prevents fund loss within business systems.
Service architecture
Benefits
The platform is non-intrusive to production systems because it collects verification data through a bypass.
Rules are configurable and require no coding. You can add or modify rules at any time to meet various verification needs.
Supports multiple verification schedules, such as real-time, near real-time, T+1, and T+H. This meets your different timeliness requirements for fund loss risk monitoring.
Provides comprehensive management features, including a verification dashboard and coverage measurement capabilities.
Supports notification channels such as text message, email, and DingTalk. This provides instant monitoring and emergency response for core services.
Provides measurement features for fund loss risk monitoring coverage. It also includes expert consulting services that provide cloud users with years of accumulated fund loss prevention experience.
Scenarios
Service protection |
Helps you periodically or regularly review core services that involve fund links. Configure verification rules to cross-check various types of data or perform logical checks on data content. This ensures that core services run correctly. |
Change risk inspection |
Before releasing a change, add verification rules for the changed business tables and their associated tables. You can also add check rules for the data in the changed business tables. This ensures there are no blind spots in fund loss risk monitoring after the change goes live. |
Historical data audit |
Perform batch checks on the historical data of existing services to see if discrepancies already exist. Promptly analyze the causes of discrepancies, fix vulnerabilities, and recover losses. |
Data quality monitoring |
Missing or incomplete data can also indirectly cause fund loss. You can configure verification rules to check data integrity, monitor data quality, and promptly detect faults. |
