《Prometheus 监控系统部署》

《Prometheus 监控系统部署》

Prometheus 是 CNCF 孵化的开源监控与告警系统,以拉取(pull)模型采集时序数据。

1 特点

  • 多维数据模型:由指标名称和键/值对标签标识的时间序列数据。
  • PromQL:一种灵活的查询语言,可利用标签维度做聚合与过滤。
  • 不依赖分布式存储,单服务器节点自治。
  • 时间序列通过 HTTP 上的 Pull 模型采集。
  • 通过中间网关支持 Push 方式上报。
  • 通过服务发现或静态配置发现目标。
  • 多种图形和仪表板支持模式。

2 组件

Prometheus 生态包含多个可选组件:

  • Prometheus 主服务器:采集并存储时序数据。项目:https://github.com/prometheus/prometheus
  • 客户端库:应用通过这些库暴露符合规范的 metric。官方支持 Go/Java/Scala/Python/Ruby,第三方还包括 nginx 的 lua 版 https://github.com/knyar/nginx-lua-prometheus。
  • 推送网关(Pushgateway):面向”运行一段时间后退出”的短生命周期任务,程序退出时把指标 push 到 pushgateway,再由 Prometheus 采集。
  • 专用导出器(Exporter):本身无法暴露 http endpoint 的服务,通过 exporter 转换采集。
  • 告警处理(Alertmanager):接收 Prometheus 发来的告警,向第三方转发邮件、短信等。

导出器完整列表:https://prometheus.io/docs/instrumenting/exporters/

3 部署(Debian)

(1)下载软件包:

wget https://github.com/prometheus/prometheus/releases/download/v2.20.1/prometheus-2.20.1.linux-amd64.tar.gz tar zxf prometheus-2.20.1.linux-amd64.tar.gz cd prometheus-2.20.1.linux-amd64

下载页面:https://prometheus.io/download/

(2)准备配置文件 prometheus.yml:

# my global config global: scrape_interval: 5s # 采集间隔,默认 1 分钟 evaluation_interval: 5s # 规则评估间隔,默认 1 分钟 # scrape_timeout 默认 10s # Alertmanager 配置 alerting: alertmanagers: - static_configs: - targets: - 127.0.0.1:8501 # 加载规则文件并按 evaluation_interval 周期评估 rule_files: # - "first_rules.yml" # - "second_rules.yml" # 抓取配置 scrape_configs: - job_name: 'prometheus' static_configs: - targets: ['127.0.0.1:8500'] - job_name: 'app-front' metrics_path: /metrics scheme: http static_configs: - targets: ['127.0.0.1:8200']

配置说明:https://prometheus.io/docs/prometheus/latest/configuration/configuration/

(3)告警规则示例:

groups: - name: security rules: - alert: ratelimit expr: nginx_front_security_count{kind="ratelimit", result="challenge"} > 2 for: 1m labels: status: warning annotations: summary: ": ratelimit challenge 超标!!!" description: ":ratelimit challenge 超标!!!"

(4)启动:

./prometheus \ --config.file=./prometheus.yml \ --web.listen-address=0.0.0.0:8500 \ --web.enable-admin-api \ --log.format=logfmt \ --log.level=debug

(5)浏览器访问 http://<server-ip>:8500/graph。

4 常用选项

--config.file="prometheus.yml" Prometheus configuration file path. --web.listen-address="0.0.0.0:9090" Address to listen on for UI, API, and telemetry. --web.enable-admin-api Enable API endpoints for admin control actions. --web.enable-lifecycle Enable shutdown and reload via HTTP request. --storage.tsdb.path="data/" Base path for metrics storage. --storage.tsdb.retention.time=15d 数据保留时间(默认 15 天)。 --log.level=info One of: [debug, info, warn, error] --log.format=logfmt One of: [logfmt, json]

5 Prometheus 级联(Federation)

通过 /federate 接口实现中心服务器聚合多个子集群的指标。

中心 server 配置:

scrape_configs: - job_name: federate scrape_interval: 3s scrape_timeout: 2s params: match[]: - '{__name__=~".*"}' honor_labels: true metrics_path: /federate static_configs: - targets: - <edge-server>:8500

注意:

  1. metrics_path 必须为 /federate。
  2. params match[] 必须有,若不过滤数据就匹配任意,缺了这个配置无法正常显示。

6 教程

阅读 — · 全站 —
🎸 我的歌单 0 首