《Prometheus 监控系统部署》
Prometheus 是 CNCF 孵化的开源监控与告警系统,以拉取(pull)模型采集时序数据。
1 特点
- 多维数据模型:由指标名称和键/值对标签标识的时间序列数据。
- PromQL:一种灵活的查询语言,可利用标签维度做聚合与过滤。
- 不依赖分布式存储,单服务器节点自治。
- 时间序列通过 HTTP 上的 Pull 模型采集。
- 通过中间网关支持 Push 方式上报。
- 通过服务发现或静态配置发现目标。
- 多种图形和仪表板支持模式。
2 组件
Prometheus 生态包含多个可选组件:
- Prometheus 主服务器:采集并存储时序数据。项目:https://github.com/prometheus/prometheus
- 客户端库:应用通过这些库暴露符合规范的 metric。官方支持 Go/Java/Scala/Python/Ruby,第三方还包括 nginx 的 lua 版 https://github.com/knyar/nginx-lua-prometheus。
- 推送网关(Pushgateway):面向”运行一段时间后退出”的短生命周期任务,程序退出时把指标 push 到 pushgateway,再由 Prometheus 采集。
- 专用导出器(Exporter):本身无法暴露 http endpoint 的服务,通过 exporter 转换采集。
- 告警处理(Alertmanager):接收 Prometheus 发来的告警,向第三方转发邮件、短信等。
导出器完整列表:https://prometheus.io/docs/instrumenting/exporters/
3 部署(Debian)
(1)下载软件包:
wget https://github.com/prometheus/prometheus/releases/download/v2.20.1/prometheus-2.20.1.linux-amd64.tar.gz
tar zxf prometheus-2.20.1.linux-amd64.tar.gz
cd prometheus-2.20.1.linux-amd64
下载页面:https://prometheus.io/download/
(2)准备配置文件 prometheus.yml:
# my global config
global:
scrape_interval: 5s # 采集间隔,默认 1 分钟
evaluation_interval: 5s # 规则评估间隔,默认 1 分钟
# scrape_timeout 默认 10s
# Alertmanager 配置
alerting:
alertmanagers:
- static_configs:
- targets:
- 127.0.0.1:8501
# 加载规则文件并按 evaluation_interval 周期评估
rule_files:
# - "first_rules.yml"
# - "second_rules.yml"
# 抓取配置
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['127.0.0.1:8500']
- job_name: 'app-front'
metrics_path: /metrics
scheme: http
static_configs:
- targets: ['127.0.0.1:8200']
配置说明:https://prometheus.io/docs/prometheus/latest/configuration/configuration/
(3)告警规则示例:
groups:
- name: security
rules:
- alert: ratelimit
expr: nginx_front_security_count{kind="ratelimit", result="challenge"} > 2
for: 1m
labels:
status: warning
annotations:
summary: ": ratelimit challenge 超标!!!"
description: ":ratelimit challenge 超标!!!"
(4)启动:
./prometheus \
--config.file=./prometheus.yml \
--web.listen-address=0.0.0.0:8500 \
--web.enable-admin-api \
--log.format=logfmt \
--log.level=debug
(5)浏览器访问 http://<server-ip>:8500/graph。
4 常用选项
--config.file="prometheus.yml" Prometheus configuration file path.
--web.listen-address="0.0.0.0:9090" Address to listen on for UI, API, and telemetry.
--web.enable-admin-api Enable API endpoints for admin control actions.
--web.enable-lifecycle Enable shutdown and reload via HTTP request.
--storage.tsdb.path="data/" Base path for metrics storage.
--storage.tsdb.retention.time=15d 数据保留时间(默认 15 天)。
--log.level=info One of: [debug, info, warn, error]
--log.format=logfmt One of: [logfmt, json]
5 Prometheus 级联(Federation)
通过 /federate 接口实现中心服务器聚合多个子集群的指标。
中心 server 配置:
scrape_configs:
- job_name: federate
scrape_interval: 3s
scrape_timeout: 2s
params:
match[]:
- '{__name__=~".*"}'
honor_labels: true
metrics_path: /federate
static_configs:
- targets:
- <edge-server>:8500
注意:
metrics_path必须为/federate。params match[]必须有,若不过滤数据就匹配任意,缺了这个配置无法正常显示。
6 教程
阅读 —
·
全站 —