Operations management

Updated at:

SOFAStack CAFE (Cloud Application Fabric Engine) is a Platform as a Service (PaaS) that provides full lifecycle management. This includes application management, release and deployment, O&M orchestration, monitoring and analysis, and disaster recovery. CAFE meets the O&M requirements for both classic and cloud-native architectures in the finance industry. It helps you transition smoothly from traditional architectures and mitigates technical risks in financial services.

image

Scenarios

Unified application runtime platform

Use the platform to solve challenges in releases, monitoring, and auditing for large-scale O&M. The platform also integrates various cloud-native features, such as containers, Serverless, and Mesh, to improve O&M efficiency.

Financial-grade high-availability architecture platform support

Provides PaaS support for intra-city active-active, unit-based, and active geo-redundancy architectures.

Upgrade from a classic to a cloud-native architecture

Supports the transition of financial infrastructure to cloud-native containerization. This reduces the technical risks associated with adopting new architectures and O&M models.

Unitized application service

Unitized Application Service (LDC Hybrid Cloud), or LHC, runs on a cloud-native infrastructure. In multi-data center and multi-region Kubernetes cluster scenarios, it provides capabilities such as application management, release and O&M, traffic scheduling, and configuration synchronization. LHC is designed to help you evolve from a single Kubernetes cluster to an active-active federated cluster. It offers disaster recovery capabilities for intra-city active-active, two-region three-data-center, and other multi-data center active-active disaster recovery scenarios. LHC can be combined with SOFAStack middleware products and the OceanBase distributed database to create a unitized active geo-redundancy architecture solution.

Service architecture

Image 74

Benefits

  • Financial-grade releases

    The release process is secure and reliable. It supports retries, canary releases, rollbacks, and traceability.

    Supports hybrid releases of virtual machines and containers. This provides a transition plan from virtual machines to containers.

  • Runtime monitoring

    You can use custom business dashboards to monitor business dynamics at any time.

    Monitors basic application metrics in real time, such as page views (PV), service invocations (Service), and calls to external services (SAL).

    Collects comprehensive basic resource metrics, such as CPU, memory, and I/O traffic.

  • Microservice framework

    Deeply integrates with Ant SOFA Mesh for service registration, discovery, and cross-language communication.

  • Network modes

    Supports VPC and Overlay network modes.

    Supports Server Load Balancer-type services and Ingress.

  • High availability and disaster recovery

    Supports intra-city active-active and two-region three-data-center disaster recovery plans.

    Supports upgrades to the unitized high-availability disaster recovery plan developed by Ant Group.

Scenarios

LHC operates in a cloud-native model. It uses a single PaaS platform to provide unified application and resource management and a unified view for releases and O&M. This enables multi-cluster management, cross-cluster application O&M and releases, resource management, and traffic management.

Intra-city active-active

Establish multiple Kubernetes clusters in two or more zones within the same region.

Two-region three-data-center

  • Build on an intra-city active-active setup by adding a data center in a different region for data and application backups. Based on network latency and bandwidth, choose a hot, warm, or cold standby solution.

  • If latency to the remote site is within an acceptable range for your services, choose a hot standby solution. Under normal conditions, users access the intra-city active-active sites. During a disaster, fail over to the geo-disaster recovery data center to maintain service availability.

Active geo-redundancy (Multi-region active-active)

The data layer is partitioned using sharding. Different zones can be divided into more logical units, or Logic Data Centers, to process different data shards. This design aims to ensure that the data access path from the access layer to the application layer and then to the data layer does not cross zones. This architecture allows for active-active deployment across any number of regions.

Hybrid cloud with heterogeneous infrastructure

Kubernetes abstracts away the differences in the underlying Infrastructure as a Service (IaaS). This lets you fully use public cloud resources by deploying business applications on both Apsara Stack and public clouds with unified O&M. In this scenario, you can achieve the following goals:

  • Reduce investment in developer and test resources. Deploy production applications on Apsara Stack and deploy developer and test applications on the public cloud as needed.

  • Meet offline disaster recovery needs. Deploy an offline environment to handle unexpected events on the public cloud, as required by national regulations. A customer example is Tianhong Yu'e Bao.

  • Achieve elastic scaling. Combine with an active geo-redundancy architecture to enable unlimited horizontal scaling of your business at the data center level as needed.

Classic application service

Classic Application Service (CAS) provides an application-centric view for visually and automatically managing application versions, deployment packages, and resources. CAS offers automated and intelligent DevOps support for the entire application lifecycle. This improves efficiency, reduces costs, and minimizes human error, allowing developers to focus on business logic.

Service architecture93A798~1

Benefits

  • Application-centric DevOps

    Provides automated DevOps support for the full application lifecycle. It shifts the management perspective from being IT resource-centric to being application-centric and business-centric. This allows users to focus on business value while improving R&D efficiency and reducing the chance of human error.

  • Customizable automated O&M

    Provides customizable automated O&M with custom technology stack solutions. This enhances the flexibility of the cloud platform and its compatibility with your existing systems. It lets you use familiar technology frameworks on the platform that are not natively provided by SOFAStack.

  • Powerful release and deployment capabilities

    Offers flexible deployment strategies such as group releases, beta releases, canary releases, single data center releases, and blue-green deployments. It supports visual and automated releases and deployments with retry and rollback capabilities to meet various requirements.

  • Flexible O&M pipeline capabilities

    Provides a channel to enter and execute custom O&M commands and scripts. This lets you perform custom O&M instructions.

Scenarios

Support for traditional O&M

Most core business applications in traditional enterprises have not been containerized. They are still deployed on virtual machines or physical servers using traditional code packages. CAS supports a smooth transition from traditional O&M methods to containerized O&M.

Decoupling IaaS and PaaS

In classic O&M scenarios, CAS supports IaaS from Alibaba Cloud and Huawei Cloud. Integration with other IaaS providers is ongoing. This allows users to avoid strong dependencies on the underlying infrastructure, achieving true decoupling of IaaS and PaaS.

CICD integration

Provides comprehensive application lifecycle APIs for integration with upstream continuous integration (CI) platforms to form a closed-loop CICD process.

Container application service

Container Application Service (Application Kubernetes Service), or AKS, fully integrates with Kubernetes. It provides complete platform capabilities for cluster management, authentication and authorization, container networking, and persistent volume (PV) storage. While ensuring standardized and consistent Kubernetes capabilities, AKS also delivers production-ready release and deployment capabilities for the full application lifecycle, based on practical experience.

Service architecture03997B~1

Benefits

  • Financial-grade releases

    The release process is secure and reliable. It supports retries, canary releases, rollbacks, and traceability.

    Supports hybrid releases of virtual machines and containers. This provides a transition plan from virtual machines to containers.

  • Runtime monitoring

    You can use custom business dashboards to monitor business dynamics at any time.

    Monitors basic application metrics in real time, such as page views (PV), Service, and SAL.

    Collects comprehensive basic resource metrics, such as CPU, memory, and I/O traffic.

  • Microservice framework

    Deeply integrates with Ant SOFA Mesh for service registration, discovery, and cross-language communication.

    Supports native deployment of Istio to provide microservice capabilities with Service Mesh.

  • Network modes

    Supports VPC and classic network modes.

  • High availability and disaster recovery

    Supports classic intra-city active-active and two-region three-data-center disaster recovery plans.

    Supports the unitized high-availability disaster recovery plan from Alibaba Cloud.

Scenarios

Traditional R&D and O&M systems using the SOFA technology stack

Applications in these systems are developed with SOFABoot or use SOFA Mesh directly. These systems have complex relationships and dependencies. They are deeply integrated with SOFAStack products and require seamless integration with the release and deployment capabilities of the existing PaaS:

  • Simultaneous release and O&M for multiple applications, with advanced capabilities such as application grouping and application dependency adjustments.

  • Requires blue-green deployment and unitized release capabilities.

Note

In this scenario, create an application service in AKS. Use the SOFABoot runtime image as the base image to build an application image. Deploy the application using in-place upgrades. This allows the application service to run on AKS, interact with services provided by virtual machines, and seamlessly connect with existing SOFAStack products.

Lightweight R&D and O&M systems using the SOFA technology stack

Applications in these systems are developed with SOFABoot or use SOFA Mesh and are tightly integrated with SOFAStack products. These applications have the following features:

  • Do not require simultaneous release of multiple applications. Each application can be released independently.

  • Require zero-downtime releases.

  • Have no legacy burdens and can adopt cloud-native O&M methods.

Note

In this scenario, create an application service in AKS. Use the SOFABoot runtime image as the base image to build an application image. Deploy and manage the application using in-place upgrades. This allows the application service to run on AKS and seamlessly connect with existing SOFAStack products.

Systems using a cloud-native technology stack

Applications in these systems typically use traditional Spring or SpringBoot technology stacks. They use Eureka or ZooKeeper for service registration and discovery and are paired with monitoring and tracing tools from the CNCF ecosystem to form a self-contained system. These applications have the following features:

  • They have a high tolerance for business errors or are non-critical systems.

  • They are released independently.

  • They have no legacy burdens and can adopt cloud-native O&M methods.

  • Do not require integration with SOFAStack products, such as monitoring, elastic scaling, disaster recovery, or middleware.

Note

In this scenario, create an application service in AKS. Use the SOFABoot runtime image as the base image to build an application image. Deploy and manage the application using in-place upgrades. This allows the application service to run on AKS and seamlessly connect with existing SOFAStack products.

Business Monitoring

Real-time Monitoring Service (RMS) is a financial-grade monitoring product with visual monitoring capabilities.

It performs multi-dimensional aggregation on massive amounts of data, including logs, metrics, and traces. It provides visual monitoring from multiple perspectives, such as business monitoring, application monitoring, cloud-native monitoring, basic resource monitoring, log query and analysis, and distributed tracing. RMS offers a rich set of visual dashboards and an alert subscription feature.

This service helps O&M, developer, and Site Reliability Engineer (SRE) teams quickly discover, locate, analyze, and resolve problems. It helps ensure the availability of online systems.

Tested in large-scale and complex business scenarios at Ant Group, RMS provides comprehensive observability and insight analysis.

Service architectureImage 75

Benefits

  • Comprehensive real-time monitoring

    Provides monitoring from various perspectives, including business, application, basic resources, and cloud-native. It can monitor key metrics at a second-level granularity and regular metrics at a minute-level granularity. The monitoring is highly reliable and timely, with low latency.

  • Flexible alert rules

    You can set alert rules based on dimensions such as business features, time periods, and importance levels. This helps prevent false positives and false negatives.

  • Convenient custom configurations

    Offers a rich set of custom configuration features, allowing for convenient and efficient configuration of products and alerts.

  • Open technology stack configuration

    Enables monitoring upon deployment for Kubernetes and SOFA technology stack applications. You can access and monitor non-standard business applications with simple technology stack configurations.

  • Visual dashboards

    A rich set of visual dashboards helps you create personalized monitoring dashboards.

  • Distributed tracing

    Provides application topology and trace query features. You can observe complex invocation relationships, performance metrics, error information, and associated logs between applications and services. This supports DevOps work such as root cause analysis, service administration, application development and testing, performance management, performance tuning, architecture management, and fault attribution.

  • Log query and association

    Provides log query and log association features. You can not only query logs but also perform historical and contextual queries. You can also view error logs associated with Error metrics and business logs associated with traces. This makes problem analysis and root cause identification more convenient and efficient.

  • Low resource consumption

    Consumes minimal host resources, such as CPU and memory, while reliably transmitting large amounts of monitoring data.

  • High availability

    You can deploy monitoring for tens of thousands of devices within minutes. It features automatic fault recovery and a scalable cluster.

  • Stable and efficient time series and data storage

    Continuously aggregates data online to ensure data capacity is controllable. It provides intelligent tiered storage and storage policies.

Scenarios

Multi-dimensional O&M

Deeply integrates with application services from technology stacks such as Kubernetes and SOFA. It provides one-stop collection of infrastructure, middleware, application runtime, and business data. Through features like metric monitoring, log analysis, distributed tracing, and alert subscription, it offers multi-dimensional O&M analysis of application performance, running status, and resource usage. This helps to promptly discover and locate problems with applications, resources, and the platform.

  • One-stop analysis: View application error metric trends, application traces, application metrics, and system metrics in the application overview. This provides one-stop application analysis capabilities.

  • Comprehensive monitoring: Covers monitoring across multiple dimensions, including infrastructure, ApsaraDB, cloud middleware, and applications. This provides one-stop O&M capabilities.

  • Fault association analysis: Provides application-centric, multi-dimensional association analysis covering components, instances, hosts, and cloud resources to quickly find abnormal faults.

Fault analysis and rapid identification

In a distributed environment, service invocations are complex. This complexity makes it difficult to analyze and identify problems. A distributed tracing system can quickly locate faulty services and help resolve issues.

  • Complete application invocation topology: The system automatically discovers a service's historical calls and its interactions with all middleware. It then creates a topology graph of the entire system's call relationships.

  • Quickly locate unhealthy applications: Unhealthy applications are highlighted in the call topology graph. This makes it easy to find and analyze problematic applications.

  • Service performance details: Drill down into any application in the topology for a detailed analysis. Analyze application performance using metrics such as throughput, error rate, and response time.

Application performance optimization

In the call relationship topology, analyze the number of calls and time consumed for each application. Find applications with high and low loads to make reasonable use of resources.

  • Call link aggregation and summary: Aggregates and summarizes all call information. Analyze the call and response situations of each application.

  • Critical path: Quickly discover the path of critical applications in the entire system call topology.

  • Optimize unreasonable calls: Promptly discover and handle unreasonable calls, such as frequent database operations.