Platform Operations
๐๏ธ Production Deployment
4 items
๐๏ธ Maintenance
This article lists common issues related to TapData maintenance.
๐๏ธ Task Scheduling Balance
TapData schedules tasks based on the number of running tasks on each live engine. When a task starts or is taken over after an exception, TapData assigns it to an engine with fewer running tasks. After cluster scaling, node recovery, or long-running workloads, task distribution can become uneven. Use Task Scheduling Balance to review the current distribution and manually move eligible tasks between engines.
๐๏ธ Integrating Prometheus Monitoring
TapData enables exposing runtime metrics in Prometheus format, allowing seamless integration into existing monitoring systems for unified observability, trend analysis, and alerting. This guide details how to enable monitoring, collect component metrics, and integrate with Prometheus, with optional Grafana dashboards for visualization.
๐๏ธ Troubleshooting
3 items
๐๏ธ Emergency Plans
This document provides a comprehensive emergency handling process and contingency strategies for TapData products, aiming to help you respond quickly and effectively in the event of an emergency or product issue, thereby mitigating the impact of failures and enhancing the overall stability and security of the product.