commit 762b4b49a0cf9cd0bd32569dadb0bfe55f9c3164 Author: wangchuanli Date: Mon Aug 24 15:00:43 2026 +0800 feat: 初始化 GPU Monitor 项目 搭建基于 Flask + SocketIO 的 GPU 集群实时监控系统,包含远程采集、定时调度、Web 前端可视化及项目文档。 diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..d35803a --- /dev/null +++ b/.gitignore @@ -0,0 +1,27 @@ +# Python +__pycache__/ +*.py[cod] +*$py.class +*.so +.Python +build/ +develop/ +dist/ +*.egg-info/ +.installed.cfg +*.egg + +# Virtual Environments +venv/ +.venv/ +env/ +.env/ + +# IDEs +.vscode/ +.idea/ + +# Config & Secrets +config.yaml +.env +*.log diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md new file mode 100644 index 0000000..990e25e --- /dev/null +++ b/CONTRIBUTING.md @@ -0,0 +1,55 @@ +# 贡献指南 (Contributing) + +感谢你考虑为 **GPU Monitor** 做出贡献!本文档说明如何参与本项目。 + +## 行为准则 + +请在所有交流中保持友善、尊重与包容。我们致力于为所有人提供友好的协作环境。 + +## 如何贡献 + +### 报告问题 (Bug Report) + +如果你发现了 Bug,请先搜索 [Issues](../../issues) 确认是否已被报告。若没有,请新建 Issue 并提供: + +- 清晰的问题描述与复现步骤 +- 操作系统、Python 版本、依赖版本 (`pip freeze`) +- 相关日志或截图 +- 期望行为与实际情况 + +### 功能建议 (Feature Request) + +欢迎提出新功能建议。请描述使用场景与期望效果,便于社区讨论。 + +### 提交代码 (Pull Request) + +1. Fork 本仓库并克隆到本地。 +2. 基于 `master` 分支创建特性分支: + ```bash + git checkout -b feature/your-feature-name + ``` +3. 安装开发依赖并确保代码可运行: + ```bash + pip install -r requirements.txt + python app.py + ``` +4. 请确保: + - 不提交任何敏感信息(如真实 IP、密码、`config.yaml`)。 + - 新增依赖需同步更新 `requirements.txt`。 + - 代码风格保持与现有代码一致(4 空格缩进、清晰的英文/中文注释)。 +5. 提交信息清晰描述改动(建议使用 Conventional Commits,如 `feat:`, `fix:`, `docs:`)。 +6. 推送分支并发起 Pull Request,描述改动内容与测试情况。 + +## 开发规范 + +- **配置安全**:所有示例配置使用 `config.example.yaml`,真实配置放 `config.yaml`(已被忽略)。 +- **密钥管理**:严禁硬编码任何密钥 / 密码,统一通过环境变量读取。 +- **代码质量**:保持函数职责单一,关键逻辑添加注释。 + +## 许可证 + +提交贡献即表示你同意你的贡献在 [MIT License](LICENSE) 下授权。 + +--- + +再次感谢你的贡献!🎉 diff --git a/LICENSE b/LICENSE new file mode 100644 index 0000000..64ff02a --- /dev/null +++ b/LICENSE @@ -0,0 +1,21 @@ +MIT License + +Copyright (c) 2026 GPU Monitor Contributors + +Permission is hereby granted, free of charge, to any person obtaining a copy +of this software and associated documentation files (the "Software"), to deal +in the Software without restriction, including without limitation the rights +to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +copies of the Software, and to permit persons to whom the Software is +furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all +copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +SOFTWARE. diff --git a/README.md b/README.md new file mode 100644 index 0000000..6bc7e4d --- /dev/null +++ b/README.md @@ -0,0 +1,166 @@ +# GPU Monitor + +[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE) + +**GPU Monitor** 是一个基于 Flask + SocketIO 的轻量级 NVIDIA GPU 集群实时监控系统。它通过 SSH 远程连接多台服务器,定时执行 `nvidia-smi` 采集 GPU 状态(利用率、温度、显存等),并通过 WebSocket 实时推送到浏览器前端,提供可视化卡片与历史趋势图表。 + +## 功能特性 + +- 🚀 **实时监控**:基于 WebSocket(SocketIO)的秒级数据推送,无需手动刷新 +- 🖥️ **多服务器支持**:在 `config.yaml` 中配置任意数量的远程服务器 +- 📊 **可视化卡片**:每台服务器的多张 GPU 卡片,展示利用率、温度、显存占用 +- 📈 **趋势图表**:点击 GPU 卡片可打开详情模态框,查看利用率历史曲线 +- 🔌 **连接状态指示**:前端实时显示 WebSocket 连接状态与服务器在线/离线 +- 🌡️ **温度阈值告警**:可配置温度告警阈值 +- 🔧 **零前端构建**:前端使用 CDN 引入 Tailwind / Socket.IO / Chart.js,无需打包 + +## 技术栈 + +| 层级 | 技术 | +|------|------| +| 后端 | Python 3.8+ / Flask / Flask-SocketIO | +| 调度 | APScheduler(后台定时任务) | +| 采集 | paramiko(SSH)调用 `nvidia-smi` | +| 前端 | HTML + Tailwind CSS + Chart.js + Socket.IO(均通过 CDN) | + +## 目录结构 + +``` +GPUMonitor/ +├── app.py # Flask 应用入口,启动 SocketIO 服务 +├── config.example.yaml # 配置示例文件(提交到仓库) +├── config.yaml # 本地配置(含敏感信息,已被 .gitignore 忽略) +├── core/ +│ ├── __init__.py +│ ├── collector.py # 通过 SSH 远程采集 GPU 数据 (GPUCollector) +│ └── scheduler.py # 定时调度与数据分发 (GPUScheduler) +├── static/ # 静态资源(js / css) +├── templates/ +│ └── index.html # 前端监控页面 +├── requirements.txt # Python 依赖 +├── README.md +├── LICENSE +└── .gitignore +``` + +## 环境要求 + +- Python 3.8 及以上 +- 目标服务器已安装 NVIDIA 驱动并可用 `nvidia-smi` 命令 +- 监控机可通过 SSH 访问目标服务器(支持密码或密钥登录) + +## 快速开始 + +### 1. 克隆仓库 + +```bash +git clone https://github.com/your-username/GPUMonitor.git +cd GPUMonitor +``` + +### 2. 创建虚拟环境并安装依赖 + +```bash +python -m venv venv +source venv/bin/activate # Linux / macOS +# venv\Scripts\activate # Windows (PowerShell/CMD) + +pip install -r requirements.txt +``` + +### 3. 配置服务器列表 + +复制示例配置文件并重命名为本地配置(**请勿将 `config.yaml` 提交到仓库**): + +```bash +cp config.example.yaml config.yaml +``` + +编辑 `config.yaml`,填入你的服务器信息: + +```yaml +servers: + - alias: 'AIServer' # 服务器别名(前端展示用) + ip: '192.168.1.10' # 服务器 IP + port: 22 # SSH 端口 + username: 'root' # SSH 用户名 + password: 'your_password' # SSH 密码(建议改用 SSH Key) + # 可继续添加更多服务器 ... + +settings: + interval: 5 # 采集间隔(秒) + timeout: 10 # SSH 连接超时(秒) + temperature_threshold: 80 # 温度告警阈值(°C) +``` + +> ⚠️ **安全提示**:配置文件中的密码为明文,请仅在受信任的本地环境使用,或改用 SSH 公钥认证。生产环境务必通过环境变量 `GPU_MONITOR_SECRET_KEY` 设置 Flask 会话密钥。 + +### 4. 运行 + +```bash +python app.py +``` + +启动后访问浏览器: + +``` +http://<监控机IP>:8099 +``` + +前端将自动通过 WebSocket 实时刷新 GPU 数据。 + +## 配置说明 + +| 配置项 | 说明 | 默认值 | +|--------|------|--------| +| `servers[].alias` | 服务器展示别名 | 必填 | +| `servers[].ip` | 服务器 IP 地址 | 必填 | +| `servers[].port` | SSH 端口 | 22 | +| `servers[].username` | SSH 用户名 | 必填 | +| `servers[].password` | SSH 密码 | 必填 | +| `settings.interval` | 采集间隔(秒) | 5 | +| `settings.timeout` | SSH 连接超时(秒) | 10 | +| `settings.temperature_threshold` | 温度告警阈值(°C) | 80 | + +## 工作原理 + +1. `GPUScheduler` 使用 APScheduler 在后台按 `interval` 周期触发采集。 +2. 每个采集周期遍历 `config.yaml` 中的服务器,由 `GPUCollector` 通过 paramiko 建立 SSH 连接。 +3. 远程执行 `nvidia-smi --query-gpu=...` 并以 CSV 格式返回,解析为结构化数据。 +4. 所有服务器结果汇总到 `last_data`,通过 SocketIO 的 `gpu_update` 事件推送给所有已连接的浏览器。 +5. 前端接收数据后动态渲染服务器卡片,并维护每张 GPU 的利用率趋势图表。 + +## 环境变量 + +| 变量名 | 说明 | +|--------|------| +| `GPU_MONITOR_SECRET_KEY` | Flask `SECRET_KEY`,用于会话签名;未设置时每次启动随机生成 | + +## 常见问题(FAQ) + +**Q: 前端显示「无法连接服务器」?** +A: 检查目标服务器 IP / 端口 / 用户名 / 密码是否正确,以及监控机到目标机的网络与 SSH 权限。 + +**Q: 卡片显示但无数据 / 数据为空?** +A: 确认目标服务器已安装 NVIDIA 驱动且 `nvidia-smi` 在登录 Shell(`bash -l`)的 PATH 中可用。 + +**Q: 如何修改监听端口?** +A: 编辑 `app.py` 末尾 `socketio.run(app, host='0.0.0.0', port=8099)` 中的 `port`。 + +**Q: 支持 HTTPS / 反向代理吗?** +A: 生产环境建议在 Nginx 等反向代理后部署,并为 SocketIO 配置 `cors_allowed_origins` 与 WebSocket 转发。 + +## 安全建议 + +- 不要将包含真实 IP / 密码的 `config.yaml` 提交到任何公开仓库。 +- 优先使用 SSH 公钥认证,避免明文密码。 +- 通过反向代理启用 HTTPS,不要将服务直接暴露在公网 `0.0.0.0`。 +- 设置强随机的 `GPU_MONITOR_SECRET_KEY` 环境变量。 + +## 贡献 + +欢迎提交 Issue 与 Pull Request!请参阅 [CONTRIBUTING.md](CONTRIBUTING.md)。 + +## 许可证 + +本项目基于 [MIT License](LICENSE) 开源。 diff --git a/app.py b/app.py new file mode 100644 index 0000000..ab117c7 --- /dev/null +++ b/app.py @@ -0,0 +1,36 @@ +from flask import Flask, render_template +from flask_socketio import SocketIO +from core.scheduler import GPUScheduler +import os + +app = Flask(__name__) +# SECRET_KEY 从环境变量读取,避免硬编码密钥;未设置时使用随机值(生产环境务必配置) +app.config['SECRET_KEY'] = os.environ.get('GPU_MONITOR_SECRET_KEY', os.urandom(24).hex()) +socketio = SocketIO(app, cors_allowed_origins="*") + +# 初始化调度器 +CONFIG_PATH = os.path.join(os.path.dirname(__file__), 'config.yaml') +gpu_scheduler = GPUScheduler(CONFIG_PATH, socketio=socketio) + +@app.route('/') +def index(): + """主监控页面""" + return render_template('index.html') + +@socketio.on('connect') +def handle_connect(): + print("Client connected") + # 客户端连接时,立即发送一次当前缓存的数据 + if gpu_scheduler.last_data: + socketio.emit('gpu_update', gpu_scheduler.last_data) + +if __name__ == '__main__': + # 启动定时采集任务 + if gpu_scheduler.start(): + print("GPU Scheduler started successfully.") + else: + print("Failed to start GPU Scheduler. Please check config.yaml.") + + # 运行 Flask 应用 + # 注意:使用 socketio.run 而不是 app.run 以支持 WebSocket + socketio.run(app, host='0.0.0.0', port=8099, debug=True) diff --git a/config.example.yaml b/config.example.yaml new file mode 100644 index 0000000..8528cdb --- /dev/null +++ b/config.example.yaml @@ -0,0 +1,24 @@ +# GPU Monitor Configuration (示例文件) +# ----------------------------------------------------------------------------- +# 复制本文件为 config.yaml 并填入你自己的服务器信息 +# 注意:config.yaml 已被 .gitignore 忽略,不会被提交到代码仓库 +# ----------------------------------------------------------------------------- + +# 目标服务器配置列表 +servers: + - alias: 'AIServer' + ip: '192.168.1.10' # 目标服务器 IP,请替换为你的真实地址 + port: 22 # SSH 端口 + username: 'root' # SSH 用户名 + password: 'your_ssh_password' # SSH 密码,建议使用 SSH Key 替代明文密码 + - alias: 'NginxAgent' + ip: '192.168.1.11' + port: 22 + username: 'root' + password: 'your_ssh_password' + +# 全局监控设置 +settings: + interval: 5 # 采集间隔 (秒) + timeout: 10 # SSH 连接超时 (秒) + temperature_threshold: 80 # 温度告警阈值 (Celsius) diff --git a/core/__init__.py b/core/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/core/collector.py b/core/collector.py new file mode 100644 index 0000000..e9508ae --- /dev/null +++ b/core/collector.py @@ -0,0 +1,88 @@ +import paramiko +import logging + +# 配置日志 +logging.basicConfig(level=logging.INFO) +logger = logging.getLogger("GPUCollector") + +class GPUCollector: + """ + 负责通过 SSH 远程连接到服务器并采集 NVIDIA GPU 状态的类 + """ + def __init__(self, server_config): + self.alias = server_config.get('alias', 'Unknown') + self.ip = server_config.get('ip') + self.port = server_config.get('port', 22) + self.username = server_config.get('username') + self.password = server_config.get('password') + self.timeout = 10 + + def fetch_gpu_data(self): + """ + 执行远程 nvidia-smi 命令并解析结果 + 返回: List[Dict] 包含每张显卡的详细信息,失败则返回 None + """ + ssh = paramiko.SSHClient() + ssh.set_missing_host_key_policy(paramiko.AutoAddPolicy()) + + try: + # 建立连接 + ssh.connect( + hostname=self.ip, + port=self.port, + username=self.username, + password=self.password, + timeout=self.timeout + ) + + # 使用 bash -l -c 强制加载登录环境 (Login Shell) + # 这样可以确保加载 /etc/profile 和 ~/.bash_profile,从而获取正确的 PATH + raw_cmd = ( + "nvidia-smi --query-gpu=index,name,temperature.gpu," + "utilization.gpu,utilization.memory,memory.total," + "memory.used,memory.free --format=csv,noheader,nounits" + ) + full_cmd = f'bash -l -c "{raw_cmd}"' + + stdin, stdout, stderr = ssh.exec_command(full_cmd) + output = stdout.read().decode('utf-8').strip() + error = stderr.read().decode('utf-8').strip() + + if error and not output: + logger.error(f"[{self.alias}] SSH Command Error: {error}") + return None + + if not output: + logger.warning(f"[{self.alias}] No output received from nvidia-smi") + return None + + # 解析 CSV 数据 + lines = output.split('\n') + gpu_list = [] + for line in lines: + if not line: continue + parts = [p.strip() for p in line.split(',')] + if len(parts) == 8: + gpu_list.append({ + "index": int(parts[0]), + "name": parts[1], + "temp": int(parts[2]), + "util": int(parts[3]), + "mem_util": int(parts[4]), + "mem_total": int(parts[5]), + "mem_used": int(parts[6]), + "mem_free": int(parts[7]) + }) + + return gpu_list + + except paramiko.AuthenticationException: + logger.error(f"[{self.alias}] SSH Authentication failed for {self.ip}") + except paramiko.SSHException as e: + logger.error(f"[{self.alias}] SSH Exception: {e}") + except Exception as e: + logger.error(f"[{self.alias}] Unexpected error: {e}") + finally: + ssh.close() + + return None diff --git a/core/scheduler.py b/core/scheduler.py new file mode 100644 index 0000000..af0d135 --- /dev/null +++ b/core/scheduler.py @@ -0,0 +1,82 @@ +import yaml +import logging +from apscheduler.schedulers.background import BackgroundScheduler +from .collector import GPUCollector + +# 配置日志 +logging.basicConfig(level=logging.INFO) +logger = logging.getLogger("GPUScheduler") + +class GPUScheduler: + """ + 负责定时触发远程采集任务并将结果分发的调度类 + """ + def __init__(self, config_path, socketio=None): + self.config_path = config_path + self.socketio = socketio # 可选:如果传入 socketio 实例,则直接推送数据到前端 + self.scheduler = BackgroundScheduler() + self.last_data = {} # 存储各服务器最后一次采集的数据 {server_alias: [gpu_info]} + + def load_config(self): + """加载本地配置文件""" + try: + with open(self.config_path, 'r', encoding='utf-8') as f: + return yaml.safe_load(f) + except Exception as e: + logger.error(f"Failed to load config file: {e}") + return None + + def collect_all_servers(self): + """遍历所有服务器并采集数据""" + config = self.load_config() + if not config: + return + + servers = config.get('servers', []) + all_results = {} + + for s_conf in servers: + alias = s_conf.get('alias', 'Unknown') + logger.info(f"Collecting data from {alias}...") + + collector = GPUCollector(s_conf) + data = collector.fetch_gpu_data() + + if data is not None: + all_results[alias] = data + else: + all_results[alias] = None # 标记为采集失败 + + # 更新内存状态 + self.last_data = all_results + + # 如果配置了 socketio,则实时推送给前端 + if self.socketio: + self.socketio.emit('gpu_update', all_results) + logger.info("GPU data broadcasted via SocketIO") + + def start(self): + """启动定时任务""" + config = self.load_config() + if not config: + logger.error("Cannot start scheduler: config file missing or invalid") + return False + + interval = config.get('settings', {}).get('interval', 5) + + # 添加定时任务 + self.scheduler.add_job( + self.collect_all_servers, + 'interval', + seconds=interval, + id='gpu_collection_job' + ) + + self.scheduler.start() + logger.info(f"Scheduler started. Interval: {interval}s") + return True + + def stop(self): + """停止调度器""" + self.scheduler.shutdown() + logger.info("Scheduler stopped") diff --git a/requirements.txt b/requirements.txt new file mode 100644 index 0000000..7f7335d --- /dev/null +++ b/requirements.txt @@ -0,0 +1,5 @@ +Flask +Flask-SocketIO +paramiko +PyYAML +APScheduler diff --git a/static/css/.gitkeep b/static/css/.gitkeep new file mode 100644 index 0000000..e69de29 diff --git a/static/js/.gitkeep b/static/js/.gitkeep new file mode 100644 index 0000000..e69de29 diff --git a/templates/index.html b/templates/index.html new file mode 100644 index 0000000..760b393 --- /dev/null +++ b/templates/index.html @@ -0,0 +1,201 @@ + + + + + + A100 GPU 实时监控系统 + + + + + + +
+
+
+

GPU Cluster Monitor

+

实时监控 NVIDIA A100 状态

+
+
+ + 连接中... +
+
+ + +
+ +
+
+ + + + + + +