Modernizing traditional applications: Beyond heterogeneous governance
During the migration from a traditional monolithic architecture to a microservices model, the growing number of microservices introduces new challenges in communication, monitoring, and security. A service mesh acts as a bridge between applications and infrastructure. It moves away from traditional software development kit (SDK) connection types and handles communication between services, and between services and the infrastructure, in a way that is transparent to applications. This achieves maximum decoupling between application development and infrastructure, allowing application developers to focus only on business logic.
Heterogeneous compatibility and standardized control
SOFA Mesh provides a platform-agnostic, language-agnostic, lightweight, and non-intrusive cloud-native solution. It lets you smoothly migrate traditional applications to the cloud, supports interconnection between heterogeneous systems, and facilitates the smooth upgrade of traditional applications to a distributed architecture. It provides dual-mode service management and governance capabilities that support both traditional architectures based on an Enterprise Service Bus (ESB) and distributed architectures such as Spring Cloud and Dubbo. Traditional applications can interconnect without modification. This integration lets you build a highly scalable, high-performance, low-cost, lightweight, and non-intrusive distributed system.
Message Mesh: Smooth cloud migration for traditional applications
Cloud-native mesh technology allows traditional applications to benefit from the technical advantages of a distributed architecture at zero or low cost. These benefits include decentralization, observability, and heterogeneous system service governance. Message Mesh takes over IBM MQ traffic to help you migrate smoothly to the cloud and reap the benefits of a distributed architecture. At the same time, you can gradually remove dependencies on the ESB. This helps you achieve the dual goals of a distributed transformation and an innovative information technology (IT) architecture, while building an enterprise-level communication network.
Unified control plane with support for multiple registries
SOFA Mesh not only supports service discovery in Kubernetes but also integrates with industry-standard service registries such as ZooKeeper, SOFA Registry, Nacos, Eureka, and Consul. This integration maximizes compatibility with existing enterprise infrastructure and minimizes implementation costs.
Agile application delivery: Dual-mode application O&M and control
The SOFAStack application Platform as a Service (PaaS) platform (CAFE) is designed to reuse existing IT assets and O&M processes. It provides fine-grained application O&M and control capabilities for dual-mode IT architectures. This capability ensures a smooth evolution from traditional O&M to containerized O&M.
Dual-mode hybrid release: From traditional VMs to cloud-native containerization
The SOFA application PaaS platform (CAFE) provides two release modes: virtual machine (VM) applications and containerized applications. You can seamlessly switch application metadata, monitoring integrations, and middleware configurations. By retaining the package-based release model of classic application O&M delivery systems, it reduces image modification and building costs for developers and O&M engineers. It also lets you select and customize technology stacks to easily associate base images and achieve near-real-time image building.

Instance-level canary release: Fine-grained control over rolling updates
The application PaaS platform supports fine-grained, group-based releases with custom topologies. This capability is an important safeguard for changes to online applications. It ensures that changes at the container instance level can be controlled and released in phases. Each stage can be paused for full online validation before continuing the release. If an exception occurs during validation, you can perform an immediate rollback to reduce the risk of changes.
End-to-end canary governance: Unit-based traffic shifting and isolation
The unit-based traffic shifting capability provided by the application PaaS platform ensures that traffic is directed and progressively validated across the entire microservice call chain.
Canary release: Flexible traffic shifting for canary deployments
For canary and blue-green releases, you can distribute traffic between new and old versions by weight using Layer 7 load balancing and traffic percentages.
You can fully validate the new version with a small amount of production traffic. This approach allows for immediate failback to prevent issues from escalating, ensures a smooth product launch, and improves release efficiency.
Blue-green deployment: Isolate old and new call units
For A/B testing, you can run versions in parallel and distribute traffic based on business semantics using Layer 7 load balancing and HTTP labels.
You can isolate calls between units by enclosing the logical context within logical release units and middleware, such as microservice invocation control. This prevents compatibility issues with east-west traffic between new and old versions during a release.
Intelligent production O&M: Technical risk framework for business continuity
TRaaS (Tech Riskdefend as a Service) is a technical risk prevention and control platform. It is based on the long-term practical methodologies and internal tools developed by Ant Group's Site Reliability Engineering (SRE) teams. TRaaS addresses the O&M challenges that users face during cloud migration and distributed transformation, such as observability, emergency response, disaster recovery, chaos engineering, financial security, and stress testing.
Unified risk emergency response for people, events, and processes
The risk emergency response capability of the TRaaS technical risk prevention and control platform focuses on handling user risks and faults. It effectively connects people, events, and platform capabilities through emergency response. This connection enables emergency initiation, personnel notification, diagnostic pushes, and plan execution. The platform also visualizes the emergency process to improve the efficiency of emergency responses.

Three-in-one business observability
The TRaaS technical risk prevention and control platform provides multiple framework protocols to collect diverse data such as monitoring metrics, traces, and logs. It also supports multi-dimensional aggregation based on business scenarios. It establishes a business continuity assurance system centered on business monitoring. Through monitoring drill-downs, trace analysis, log correlation, and fault decision tree diagnostics, these features jointly build a fault localization and analysis system. The platform provides comprehensive, real-time monitoring from various perspectives, including business, application, underlying resources, and cloud-native, and offers one-stop O&M capabilities.

Efficient financial security and risk assurance
Any financial loss is a disaster for a company's brand. To prevent this, a near-real-time, non-intrusive engineering system is required that provides real-time detection, real-time alerting, real-time circuit breaking, and real-time remediation. The TRaaS financial security and risk assurance platform establishes a bypass monitoring system for financial security to safeguard critical financial operations. This system ensures that financial errors are detected in real time. Based on the experience of Ant Group and MYbank, it provides consulting services for building and implementing reconciliation rules.
High-fidelity performance and capacity risk identification
TRaaS end-to-end stress testing is based on an end-to-end business model. It includes the entire system environment, including frontend systems, backend applications, middleware adaptation layers, and databases, in the scope of the test. Using HTTP requests, it simulates real user behavior to generate realistic, large-scale access traffic in the production environment. Pressure is applied based on the end-to-end stress testing model until the target peak is reached, which helps identify system bottlenecks and validate system capabilities. Compared to single-application or single-link stress testing, end-to-end stress testing provides a broader view of business scenarios and more comprehensively and accurately reflects the current business support capabilities.

End-to-end disaster recovery switchover
The TRaaS technical risk platform provides end-to-end disaster recovery switchover capabilities. It supports disaster recovery plans for Infrastructure as a Service (IaaS), PaaS, and Software as a Service (SaaS) through unified orchestration and regular drills. You can use a unified platform to perform switchover actions for different disaster recovery scenarios.

Chaos engineering-based red vs. blue team capabilities
TRaaS uses chaos engineering to establish a drill mechanism. It provides fault injection and drill orchestration capabilities. It supports proactively injecting faults into business systems in different environments and at different stages. This approach lets you proactively observe the robustness of a single application while also validating the entire system's capabilities for fault detection, emergency response, and self-healing.
Hybrid cloud active geo-redundancy: Unit-based architecture for disaster recovery and capacity
SOFAStack leverages cloud-native technology to deliver Ant Group's financial-grade unit-based architecture, developed over many years. This architecture includes capabilities for agile development, distributed architecture, release and O&M, monitoring and analysis, and disaster recovery and emergency response. It supports multiple key scenarios, such as same-city active-active, two-region, three-data-center, active geo-redundancy, and hybrid cloud on heterogeneous infrastructure. It also provides a sustainable and evolutionary path for architecture upgrades.
Active geo-redundancy with unit-based disaster recovery
The active geo-redundancy unit-based disaster recovery architecture is designed to solve the following challenges based on application and data splitting in a distributed architecture:
Reduces access latency in geo-distributed scenarios to achieve active geo-redundancy.
Overcomes database connection limits in a single data center to break through physical constraints.
Simplifies capacity estimation and scaling by enabling estimation and scaling by unit.
Addresses challenges with phased releases by supporting canary releases.

Hybrid cloud: Federated release and control for active-active applications on heterogeneous infrastructure
In a narrow sense, hybrid cloud refers to a mix of public cloud and private deployments. Through technologies and products for platform capability abstraction and unification, automated deployment, and configuration management, it shifts the focus of developers and O&M engineers away from the underlying infrastructure. This allows applications and data to be deployed and managed in a hybrid data center environment.
Hybrid cloud active geo-redundancy architecture
SOFAStack hybrid cloud provides a comprehensive solution that covers DevOps workflows, active-active disaster recovery, and the O&M, deployment, and management of applications and workloads through the unification and standardization of PaaS. It uses Kubernetes to abstract away the differences in the underlying IaaS. This approach improves resource utilization across clouds, enables new business opportunities, and better serves business needs.

Business value
Extends data center boundaries with global delivery, multi-data center redundancy, active-active deployment, and cost reduction through the public cloud.
Provides multicloud management with a single solution to manage diverse IT and cloud assets.
Offers unified application release and control, and integrates with CI/CD pipelines.
Delivers high availability and disaster recovery capabilities across clusters, data centers, and regions.
Meets multi-cluster governance needs for scenarios that require financial security domain isolation.
Enables elastic scaling. When combined with the active geo-redundancy architecture, it allows businesses to scale out horizontally at the data center level as needed.
Unified container network RAMA: Flexible and high-performance deployments for complex scenarios
In Apsara Stack PaaS solution scenarios, the main focus of implementation is often adapting to the underlying heterogeneous infrastructure, such as networking and storage. Although Kubernetes defines the generic Container Network Interface (CNI), it leaves the implementation complexity to specific network plugins. To address the needs for performance, adaptability, and flexible IP management at the network layer in Apsara Stack PaaS solutions, Alibaba developed RAMA, a collaborative, cross-departmental container network project.
RAMA was launched in October 2019. The first phase reached General Availability (GA) in February 2020. After several iterations, it now supports deployment in complex scenarios with multiple vSwitches, multiple VLANs, and multiple CIDR blocks. It also supports advanced policy-based allocation such as IP address assignment and static IP addresses, ARP protection, MAC address spoofing, high-performance networking, and full compatibility with the native kube-proxy. Future plans for RAMA include adding support for dual-stack networks, further optimizing the data plane, exploring support for mixed overlay deployments, and introducing eBPF to enhance network observability.

Core advantages
High network performance: The underlay network has no encapsulation or decapsulation overhead. It offers significant performance advantages over solutions such as flannel-ipip and kube-ovn. This helps users deploy production applications in containers and integrate with their existing network facilities in a manner very similar to traditional network O&M.
Strong adaptability: With MAC NAT and ARP Protection features, it bypasses the MAC address filtering restrictions of underlying cloud platforms and isolates malicious ARP requests. It is verified to run successfully in heterogeneous environments with physical servers and OpenStack or VMware virtual machines.
Flexible network model: A custom multi-level model of Network/Subnet/IPInstance corresponds to the underlying vSwitches. It also supports pre-reserved IPs and blacklisted IPs, allowing for more flexible network planning and usage.
Advanced IP management policies: Supports IP assignment and static IP policies for stateful pods (workloads). This helps users more easily migrate applications that require a fixed IP address to the cloud.
Compatible with Kubernetes Service semantics: Using kernel features such as policy-based routing and transparent bridging, it ensures the underlay network is compatible with the iptables-based ClusterIP in Kubernetes. This provides users with a more comprehensive and user-friendly Kubernetes experience.
Supports IPv4/IPv6 dual-stack
Comparison with community and industry solutions
Feature comparison
From a feature perspective, RAMA's underlay network design gives it unique advantages in static IPs, IP assignment, support for multiple CIDR blocks and vSwitches, and integration with existing networks. It matches or surpasses many benchmark container network solutions in the industry. In addition, RAMA also includes many improvements in adaptability, security, and compatibility with Kubernetes. These improvements have resulted in additional features, such as MAC NAT, ARP Protection, and full compatibility with kube-proxy.
Performance comparison
From a performance perspective, RAMA is based on fundamental kernel networking capabilities such as linuxbridge, vlan, iptables, ip rule, and ebtables. The data link is simple and efficient. It has a performance loss of only about 10% compared to the host network, surpassing most benchmark container network solutions in the industry.