What is Real-time Monitoring Service
Real-time Monitoring Service (RMS) is a financial-grade monitoring product with visualization capabilities. It performs multi-dimensional aggregation on massive amounts of data, such as logs, metrics, and traces. The service provides visualization features from multiple perspectives, such as business monitoring, application monitoring, cloud-native monitoring, basic resource monitoring, log query and analysis, and distributed tracing. It includes rich visualization dashboards and an alert subscription feature. This service helps DevOps, developers, and Site Reliability Engineers (SREs) quickly identify, locate, analyze, and resolve issues to ensure the availability of online systems.
Real-time Monitoring Service supports the following features:
Comprehensive real-time monitoring: Monitors business, applications, basic resources, and cloud-native environments from various perspectives. It monitors key metrics at the second level and common metrics at the minute level. The service is highly reliable, timely, and has low latency.
Flexible alerting rules: You can set alerting rules based on dimensions such as business features, time periods, and importance levels. This helps ensure accurate alerts while minimizing false positives and false negatives.
Convenient custom configuration: You can configure products and alerts easily and efficiently using a wide range of custom options.
Open technology stack configuration: You can monitor Kubernetes and SOFAStack applications as soon as they are deployed. You can also connect and monitor non-standard business applications with simple configurations and use rich visualization tools to create custom monitoring dashboards.
Distributed tracing: Provides application topology and trace query features that allow you to observe complex call relationships between applications and services, performance metrics, error information, and associated logs. This enables DevOps tasks such as root cause analysis, service administration, application development and debugging, performance management, performance tuning, architecture management, and fault attribution.
Log query and correlation: Provides log query and correlation features. You can query logs, perform historical and contextual queries, and view error logs associated with error metrics and business logs associated with traces. This makes problem analysis and identification more efficient.
Low resource consumption: The service ensures very low consumption of host resources, such as CPU and memory, while reliably transferring large amounts of monitoring data.
High availability: The service supports minute-level monitoring deployments for tens of thousands of devices and features automatic fault recovery and cluster scalability.
Stable and efficient time series and data storage: The service continuously aggregates data online to keep the data volume manageable and includes intelligent tiered storage and placement policies.