返回 Skill 列表
extension
分类: 开发与工程无需 API Key

altinity-expert-clickhouse-metrics

实时监控ClickHouse的指标、事件和异步指标。用于负载平均值、连接、队列监控和资源饱和度。

person作者: jakexiaohubgithub

Real-Time Metrics Monitoring

Real-time monitoring of ClickHouse metrics, events, and asynchronous metrics.


Diagnostics

Run all queries from checks.sql in this skill's directory and analyze the results.


Ad-Hoc Query Guidelines

Key Tables

  • system.metrics - Current gauge values
  • system.events - Cumulative counters since restart
  • system.asynchronous_metrics - System-level metrics
  • system.metric_log - Historical metrics
  • system.asynchronous_metric_log - Historical async metrics

Useful Patterns

-- Find metrics by pattern
select * from system.metrics where metric like '%pattern%'
select * from system.asynchronous_metrics where metric like '%pattern%'
select * from system.events where event like '%pattern%'

Cross-Module Triggers

| Finding | Load Module | Reason | |---------|-------------|--------| | High memory metrics | altinity-expert-clickhouse-memory | Memory analysis | | High replica delay | altinity-expert-clickhouse-replication | Replication issues | | High parts count | altinity-expert-clickhouse-merges | Merge backlog | | High load average | altinity-expert-clickhouse-reporting | Query analysis | | High connections | altinity-expert-clickhouse-reporting | Connection analysis |


Monitoring Recommendations

Key Metrics to Alert On

| Metric | Warning | Critical | |--------|---------|----------| | ReadonlyReplica | - | > 0 | | Query | > 75% max | > 90% max | | MemoryResident | > 80% RAM | > 90% RAM | | MaxPartCountForPartition | > parts_to_delay | > parts_to_throw | | ReplicasMaxAbsoluteDelay | > 5 min | > 1 hour | | LoadAverage1 | > CPU count | > 2x CPU count |

Prometheus/Grafana Export

ClickHouse exposes metrics at :9363/metrics in Prometheus format when enabled.