feat: 初始化 GPU Monitor 项目
搭建基于 Flask + SocketIO 的 GPU 集群实时监控系统,包含远程采集、定时调度、Web 前端可视化及项目文档。
这个提交包含在:
+27
@@ -0,0 +1,27 @@
|
|||||||
|
# Python
|
||||||
|
__pycache__/
|
||||||
|
*.py[cod]
|
||||||
|
*$py.class
|
||||||
|
*.so
|
||||||
|
.Python
|
||||||
|
build/
|
||||||
|
develop/
|
||||||
|
dist/
|
||||||
|
*.egg-info/
|
||||||
|
.installed.cfg
|
||||||
|
*.egg
|
||||||
|
|
||||||
|
# Virtual Environments
|
||||||
|
venv/
|
||||||
|
.venv/
|
||||||
|
env/
|
||||||
|
.env/
|
||||||
|
|
||||||
|
# IDEs
|
||||||
|
.vscode/
|
||||||
|
.idea/
|
||||||
|
|
||||||
|
# Config & Secrets
|
||||||
|
config.yaml
|
||||||
|
.env
|
||||||
|
*.log
|
||||||
+55
@@ -0,0 +1,55 @@
|
|||||||
|
# 贡献指南 (Contributing)
|
||||||
|
|
||||||
|
感谢你考虑为 **GPU Monitor** 做出贡献!本文档说明如何参与本项目。
|
||||||
|
|
||||||
|
## 行为准则
|
||||||
|
|
||||||
|
请在所有交流中保持友善、尊重与包容。我们致力于为所有人提供友好的协作环境。
|
||||||
|
|
||||||
|
## 如何贡献
|
||||||
|
|
||||||
|
### 报告问题 (Bug Report)
|
||||||
|
|
||||||
|
如果你发现了 Bug,请先搜索 [Issues](../../issues) 确认是否已被报告。若没有,请新建 Issue 并提供:
|
||||||
|
|
||||||
|
- 清晰的问题描述与复现步骤
|
||||||
|
- 操作系统、Python 版本、依赖版本 (`pip freeze`)
|
||||||
|
- 相关日志或截图
|
||||||
|
- 期望行为与实际情况
|
||||||
|
|
||||||
|
### 功能建议 (Feature Request)
|
||||||
|
|
||||||
|
欢迎提出新功能建议。请描述使用场景与期望效果,便于社区讨论。
|
||||||
|
|
||||||
|
### 提交代码 (Pull Request)
|
||||||
|
|
||||||
|
1. Fork 本仓库并克隆到本地。
|
||||||
|
2. 基于 `master` 分支创建特性分支:
|
||||||
|
```bash
|
||||||
|
git checkout -b feature/your-feature-name
|
||||||
|
```
|
||||||
|
3. 安装开发依赖并确保代码可运行:
|
||||||
|
```bash
|
||||||
|
pip install -r requirements.txt
|
||||||
|
python app.py
|
||||||
|
```
|
||||||
|
4. 请确保:
|
||||||
|
- 不提交任何敏感信息(如真实 IP、密码、`config.yaml`)。
|
||||||
|
- 新增依赖需同步更新 `requirements.txt`。
|
||||||
|
- 代码风格保持与现有代码一致(4 空格缩进、清晰的英文/中文注释)。
|
||||||
|
5. 提交信息清晰描述改动(建议使用 Conventional Commits,如 `feat:`, `fix:`, `docs:`)。
|
||||||
|
6. 推送分支并发起 Pull Request,描述改动内容与测试情况。
|
||||||
|
|
||||||
|
## 开发规范
|
||||||
|
|
||||||
|
- **配置安全**:所有示例配置使用 `config.example.yaml`,真实配置放 `config.yaml`(已被忽略)。
|
||||||
|
- **密钥管理**:严禁硬编码任何密钥 / 密码,统一通过环境变量读取。
|
||||||
|
- **代码质量**:保持函数职责单一,关键逻辑添加注释。
|
||||||
|
|
||||||
|
## 许可证
|
||||||
|
|
||||||
|
提交贡献即表示你同意你的贡献在 [MIT License](LICENSE) 下授权。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
再次感谢你的贡献!🎉
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
MIT License
|
||||||
|
|
||||||
|
Copyright (c) 2026 GPU Monitor Contributors
|
||||||
|
|
||||||
|
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||||
|
of this software and associated documentation files (the "Software"), to deal
|
||||||
|
in the Software without restriction, including without limitation the rights
|
||||||
|
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||||
|
copies of the Software, and to permit persons to whom the Software is
|
||||||
|
furnished to do so, subject to the following conditions:
|
||||||
|
|
||||||
|
The above copyright notice and this permission notice shall be included in all
|
||||||
|
copies or substantial portions of the Software.
|
||||||
|
|
||||||
|
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||||
|
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||||
|
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||||
|
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||||
|
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||||
|
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||||
|
SOFTWARE.
|
||||||
@@ -0,0 +1,166 @@
|
|||||||
|
# GPU Monitor
|
||||||
|
|
||||||
|
[](LICENSE)
|
||||||
|
|
||||||
|
**GPU Monitor** 是一个基于 Flask + SocketIO 的轻量级 NVIDIA GPU 集群实时监控系统。它通过 SSH 远程连接多台服务器,定时执行 `nvidia-smi` 采集 GPU 状态(利用率、温度、显存等),并通过 WebSocket 实时推送到浏览器前端,提供可视化卡片与历史趋势图表。
|
||||||
|
|
||||||
|
## 功能特性
|
||||||
|
|
||||||
|
- 🚀 **实时监控**:基于 WebSocket(SocketIO)的秒级数据推送,无需手动刷新
|
||||||
|
- 🖥️ **多服务器支持**:在 `config.yaml` 中配置任意数量的远程服务器
|
||||||
|
- 📊 **可视化卡片**:每台服务器的多张 GPU 卡片,展示利用率、温度、显存占用
|
||||||
|
- 📈 **趋势图表**:点击 GPU 卡片可打开详情模态框,查看利用率历史曲线
|
||||||
|
- 🔌 **连接状态指示**:前端实时显示 WebSocket 连接状态与服务器在线/离线
|
||||||
|
- 🌡️ **温度阈值告警**:可配置温度告警阈值
|
||||||
|
- 🔧 **零前端构建**:前端使用 CDN 引入 Tailwind / Socket.IO / Chart.js,无需打包
|
||||||
|
|
||||||
|
## 技术栈
|
||||||
|
|
||||||
|
| 层级 | 技术 |
|
||||||
|
|------|------|
|
||||||
|
| 后端 | Python 3.8+ / Flask / Flask-SocketIO |
|
||||||
|
| 调度 | APScheduler(后台定时任务) |
|
||||||
|
| 采集 | paramiko(SSH)调用 `nvidia-smi` |
|
||||||
|
| 前端 | HTML + Tailwind CSS + Chart.js + Socket.IO(均通过 CDN) |
|
||||||
|
|
||||||
|
## 目录结构
|
||||||
|
|
||||||
|
```
|
||||||
|
GPUMonitor/
|
||||||
|
├── app.py # Flask 应用入口,启动 SocketIO 服务
|
||||||
|
├── config.example.yaml # 配置示例文件(提交到仓库)
|
||||||
|
├── config.yaml # 本地配置(含敏感信息,已被 .gitignore 忽略)
|
||||||
|
├── core/
|
||||||
|
│ ├── __init__.py
|
||||||
|
│ ├── collector.py # 通过 SSH 远程采集 GPU 数据 (GPUCollector)
|
||||||
|
│ └── scheduler.py # 定时调度与数据分发 (GPUScheduler)
|
||||||
|
├── static/ # 静态资源(js / css)
|
||||||
|
├── templates/
|
||||||
|
│ └── index.html # 前端监控页面
|
||||||
|
├── requirements.txt # Python 依赖
|
||||||
|
├── README.md
|
||||||
|
├── LICENSE
|
||||||
|
└── .gitignore
|
||||||
|
```
|
||||||
|
|
||||||
|
## 环境要求
|
||||||
|
|
||||||
|
- Python 3.8 及以上
|
||||||
|
- 目标服务器已安装 NVIDIA 驱动并可用 `nvidia-smi` 命令
|
||||||
|
- 监控机可通过 SSH 访问目标服务器(支持密码或密钥登录)
|
||||||
|
|
||||||
|
## 快速开始
|
||||||
|
|
||||||
|
### 1. 克隆仓库
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git clone https://github.com/your-username/GPUMonitor.git
|
||||||
|
cd GPUMonitor
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. 创建虚拟环境并安装依赖
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m venv venv
|
||||||
|
source venv/bin/activate # Linux / macOS
|
||||||
|
# venv\Scripts\activate # Windows (PowerShell/CMD)
|
||||||
|
|
||||||
|
pip install -r requirements.txt
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. 配置服务器列表
|
||||||
|
|
||||||
|
复制示例配置文件并重命名为本地配置(**请勿将 `config.yaml` 提交到仓库**):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cp config.example.yaml config.yaml
|
||||||
|
```
|
||||||
|
|
||||||
|
编辑 `config.yaml`,填入你的服务器信息:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
servers:
|
||||||
|
- alias: 'AIServer' # 服务器别名(前端展示用)
|
||||||
|
ip: '192.168.1.10' # 服务器 IP
|
||||||
|
port: 22 # SSH 端口
|
||||||
|
username: 'root' # SSH 用户名
|
||||||
|
password: 'your_password' # SSH 密码(建议改用 SSH Key)
|
||||||
|
# 可继续添加更多服务器 ...
|
||||||
|
|
||||||
|
settings:
|
||||||
|
interval: 5 # 采集间隔(秒)
|
||||||
|
timeout: 10 # SSH 连接超时(秒)
|
||||||
|
temperature_threshold: 80 # 温度告警阈值(°C)
|
||||||
|
```
|
||||||
|
|
||||||
|
> ⚠️ **安全提示**:配置文件中的密码为明文,请仅在受信任的本地环境使用,或改用 SSH 公钥认证。生产环境务必通过环境变量 `GPU_MONITOR_SECRET_KEY` 设置 Flask 会话密钥。
|
||||||
|
|
||||||
|
### 4. 运行
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python app.py
|
||||||
|
```
|
||||||
|
|
||||||
|
启动后访问浏览器:
|
||||||
|
|
||||||
|
```
|
||||||
|
http://<监控机IP>:8099
|
||||||
|
```
|
||||||
|
|
||||||
|
前端将自动通过 WebSocket 实时刷新 GPU 数据。
|
||||||
|
|
||||||
|
## 配置说明
|
||||||
|
|
||||||
|
| 配置项 | 说明 | 默认值 |
|
||||||
|
|--------|------|--------|
|
||||||
|
| `servers[].alias` | 服务器展示别名 | 必填 |
|
||||||
|
| `servers[].ip` | 服务器 IP 地址 | 必填 |
|
||||||
|
| `servers[].port` | SSH 端口 | 22 |
|
||||||
|
| `servers[].username` | SSH 用户名 | 必填 |
|
||||||
|
| `servers[].password` | SSH 密码 | 必填 |
|
||||||
|
| `settings.interval` | 采集间隔(秒) | 5 |
|
||||||
|
| `settings.timeout` | SSH 连接超时(秒) | 10 |
|
||||||
|
| `settings.temperature_threshold` | 温度告警阈值(°C) | 80 |
|
||||||
|
|
||||||
|
## 工作原理
|
||||||
|
|
||||||
|
1. `GPUScheduler` 使用 APScheduler 在后台按 `interval` 周期触发采集。
|
||||||
|
2. 每个采集周期遍历 `config.yaml` 中的服务器,由 `GPUCollector` 通过 paramiko 建立 SSH 连接。
|
||||||
|
3. 远程执行 `nvidia-smi --query-gpu=...` 并以 CSV 格式返回,解析为结构化数据。
|
||||||
|
4. 所有服务器结果汇总到 `last_data`,通过 SocketIO 的 `gpu_update` 事件推送给所有已连接的浏览器。
|
||||||
|
5. 前端接收数据后动态渲染服务器卡片,并维护每张 GPU 的利用率趋势图表。
|
||||||
|
|
||||||
|
## 环境变量
|
||||||
|
|
||||||
|
| 变量名 | 说明 |
|
||||||
|
|--------|------|
|
||||||
|
| `GPU_MONITOR_SECRET_KEY` | Flask `SECRET_KEY`,用于会话签名;未设置时每次启动随机生成 |
|
||||||
|
|
||||||
|
## 常见问题(FAQ)
|
||||||
|
|
||||||
|
**Q: 前端显示「无法连接服务器」?**
|
||||||
|
A: 检查目标服务器 IP / 端口 / 用户名 / 密码是否正确,以及监控机到目标机的网络与 SSH 权限。
|
||||||
|
|
||||||
|
**Q: 卡片显示但无数据 / 数据为空?**
|
||||||
|
A: 确认目标服务器已安装 NVIDIA 驱动且 `nvidia-smi` 在登录 Shell(`bash -l`)的 PATH 中可用。
|
||||||
|
|
||||||
|
**Q: 如何修改监听端口?**
|
||||||
|
A: 编辑 `app.py` 末尾 `socketio.run(app, host='0.0.0.0', port=8099)` 中的 `port`。
|
||||||
|
|
||||||
|
**Q: 支持 HTTPS / 反向代理吗?**
|
||||||
|
A: 生产环境建议在 Nginx 等反向代理后部署,并为 SocketIO 配置 `cors_allowed_origins` 与 WebSocket 转发。
|
||||||
|
|
||||||
|
## 安全建议
|
||||||
|
|
||||||
|
- 不要将包含真实 IP / 密码的 `config.yaml` 提交到任何公开仓库。
|
||||||
|
- 优先使用 SSH 公钥认证,避免明文密码。
|
||||||
|
- 通过反向代理启用 HTTPS,不要将服务直接暴露在公网 `0.0.0.0`。
|
||||||
|
- 设置强随机的 `GPU_MONITOR_SECRET_KEY` 环境变量。
|
||||||
|
|
||||||
|
## 贡献
|
||||||
|
|
||||||
|
欢迎提交 Issue 与 Pull Request!请参阅 [CONTRIBUTING.md](CONTRIBUTING.md)。
|
||||||
|
|
||||||
|
## 许可证
|
||||||
|
|
||||||
|
本项目基于 [MIT License](LICENSE) 开源。
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
from flask import Flask, render_template
|
||||||
|
from flask_socketio import SocketIO
|
||||||
|
from core.scheduler import GPUScheduler
|
||||||
|
import os
|
||||||
|
|
||||||
|
app = Flask(__name__)
|
||||||
|
# SECRET_KEY 从环境变量读取,避免硬编码密钥;未设置时使用随机值(生产环境务必配置)
|
||||||
|
app.config['SECRET_KEY'] = os.environ.get('GPU_MONITOR_SECRET_KEY', os.urandom(24).hex())
|
||||||
|
socketio = SocketIO(app, cors_allowed_origins="*")
|
||||||
|
|
||||||
|
# 初始化调度器
|
||||||
|
CONFIG_PATH = os.path.join(os.path.dirname(__file__), 'config.yaml')
|
||||||
|
gpu_scheduler = GPUScheduler(CONFIG_PATH, socketio=socketio)
|
||||||
|
|
||||||
|
@app.route('/')
|
||||||
|
def index():
|
||||||
|
"""主监控页面"""
|
||||||
|
return render_template('index.html')
|
||||||
|
|
||||||
|
@socketio.on('connect')
|
||||||
|
def handle_connect():
|
||||||
|
print("Client connected")
|
||||||
|
# 客户端连接时,立即发送一次当前缓存的数据
|
||||||
|
if gpu_scheduler.last_data:
|
||||||
|
socketio.emit('gpu_update', gpu_scheduler.last_data)
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
# 启动定时采集任务
|
||||||
|
if gpu_scheduler.start():
|
||||||
|
print("GPU Scheduler started successfully.")
|
||||||
|
else:
|
||||||
|
print("Failed to start GPU Scheduler. Please check config.yaml.")
|
||||||
|
|
||||||
|
# 运行 Flask 应用
|
||||||
|
# 注意:使用 socketio.run 而不是 app.run 以支持 WebSocket
|
||||||
|
socketio.run(app, host='0.0.0.0', port=8099, debug=True)
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
# GPU Monitor Configuration (示例文件)
|
||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
# 复制本文件为 config.yaml 并填入你自己的服务器信息
|
||||||
|
# 注意:config.yaml 已被 .gitignore 忽略,不会被提交到代码仓库
|
||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
|
||||||
|
# 目标服务器配置列表
|
||||||
|
servers:
|
||||||
|
- alias: 'AIServer'
|
||||||
|
ip: '192.168.1.10' # 目标服务器 IP,请替换为你的真实地址
|
||||||
|
port: 22 # SSH 端口
|
||||||
|
username: 'root' # SSH 用户名
|
||||||
|
password: 'your_ssh_password' # SSH 密码,建议使用 SSH Key 替代明文密码
|
||||||
|
- alias: 'NginxAgent'
|
||||||
|
ip: '192.168.1.11'
|
||||||
|
port: 22
|
||||||
|
username: 'root'
|
||||||
|
password: 'your_ssh_password'
|
||||||
|
|
||||||
|
# 全局监控设置
|
||||||
|
settings:
|
||||||
|
interval: 5 # 采集间隔 (秒)
|
||||||
|
timeout: 10 # SSH 连接超时 (秒)
|
||||||
|
temperature_threshold: 80 # 温度告警阈值 (Celsius)
|
||||||
@@ -0,0 +1,88 @@
|
|||||||
|
import paramiko
|
||||||
|
import logging
|
||||||
|
|
||||||
|
# 配置日志
|
||||||
|
logging.basicConfig(level=logging.INFO)
|
||||||
|
logger = logging.getLogger("GPUCollector")
|
||||||
|
|
||||||
|
class GPUCollector:
|
||||||
|
"""
|
||||||
|
负责通过 SSH 远程连接到服务器并采集 NVIDIA GPU 状态的类
|
||||||
|
"""
|
||||||
|
def __init__(self, server_config):
|
||||||
|
self.alias = server_config.get('alias', 'Unknown')
|
||||||
|
self.ip = server_config.get('ip')
|
||||||
|
self.port = server_config.get('port', 22)
|
||||||
|
self.username = server_config.get('username')
|
||||||
|
self.password = server_config.get('password')
|
||||||
|
self.timeout = 10
|
||||||
|
|
||||||
|
def fetch_gpu_data(self):
|
||||||
|
"""
|
||||||
|
执行远程 nvidia-smi 命令并解析结果
|
||||||
|
返回: List[Dict] 包含每张显卡的详细信息,失败则返回 None
|
||||||
|
"""
|
||||||
|
ssh = paramiko.SSHClient()
|
||||||
|
ssh.set_missing_host_key_policy(paramiko.AutoAddPolicy())
|
||||||
|
|
||||||
|
try:
|
||||||
|
# 建立连接
|
||||||
|
ssh.connect(
|
||||||
|
hostname=self.ip,
|
||||||
|
port=self.port,
|
||||||
|
username=self.username,
|
||||||
|
password=self.password,
|
||||||
|
timeout=self.timeout
|
||||||
|
)
|
||||||
|
|
||||||
|
# 使用 bash -l -c 强制加载登录环境 (Login Shell)
|
||||||
|
# 这样可以确保加载 /etc/profile 和 ~/.bash_profile,从而获取正确的 PATH
|
||||||
|
raw_cmd = (
|
||||||
|
"nvidia-smi --query-gpu=index,name,temperature.gpu,"
|
||||||
|
"utilization.gpu,utilization.memory,memory.total,"
|
||||||
|
"memory.used,memory.free --format=csv,noheader,nounits"
|
||||||
|
)
|
||||||
|
full_cmd = f'bash -l -c "{raw_cmd}"'
|
||||||
|
|
||||||
|
stdin, stdout, stderr = ssh.exec_command(full_cmd)
|
||||||
|
output = stdout.read().decode('utf-8').strip()
|
||||||
|
error = stderr.read().decode('utf-8').strip()
|
||||||
|
|
||||||
|
if error and not output:
|
||||||
|
logger.error(f"[{self.alias}] SSH Command Error: {error}")
|
||||||
|
return None
|
||||||
|
|
||||||
|
if not output:
|
||||||
|
logger.warning(f"[{self.alias}] No output received from nvidia-smi")
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 解析 CSV 数据
|
||||||
|
lines = output.split('\n')
|
||||||
|
gpu_list = []
|
||||||
|
for line in lines:
|
||||||
|
if not line: continue
|
||||||
|
parts = [p.strip() for p in line.split(',')]
|
||||||
|
if len(parts) == 8:
|
||||||
|
gpu_list.append({
|
||||||
|
"index": int(parts[0]),
|
||||||
|
"name": parts[1],
|
||||||
|
"temp": int(parts[2]),
|
||||||
|
"util": int(parts[3]),
|
||||||
|
"mem_util": int(parts[4]),
|
||||||
|
"mem_total": int(parts[5]),
|
||||||
|
"mem_used": int(parts[6]),
|
||||||
|
"mem_free": int(parts[7])
|
||||||
|
})
|
||||||
|
|
||||||
|
return gpu_list
|
||||||
|
|
||||||
|
except paramiko.AuthenticationException:
|
||||||
|
logger.error(f"[{self.alias}] SSH Authentication failed for {self.ip}")
|
||||||
|
except paramiko.SSHException as e:
|
||||||
|
logger.error(f"[{self.alias}] SSH Exception: {e}")
|
||||||
|
except Exception as e:
|
||||||
|
logger.error(f"[{self.alias}] Unexpected error: {e}")
|
||||||
|
finally:
|
||||||
|
ssh.close()
|
||||||
|
|
||||||
|
return None
|
||||||
@@ -0,0 +1,82 @@
|
|||||||
|
import yaml
|
||||||
|
import logging
|
||||||
|
from apscheduler.schedulers.background import BackgroundScheduler
|
||||||
|
from .collector import GPUCollector
|
||||||
|
|
||||||
|
# 配置日志
|
||||||
|
logging.basicConfig(level=logging.INFO)
|
||||||
|
logger = logging.getLogger("GPUScheduler")
|
||||||
|
|
||||||
|
class GPUScheduler:
|
||||||
|
"""
|
||||||
|
负责定时触发远程采集任务并将结果分发的调度类
|
||||||
|
"""
|
||||||
|
def __init__(self, config_path, socketio=None):
|
||||||
|
self.config_path = config_path
|
||||||
|
self.socketio = socketio # 可选:如果传入 socketio 实例,则直接推送数据到前端
|
||||||
|
self.scheduler = BackgroundScheduler()
|
||||||
|
self.last_data = {} # 存储各服务器最后一次采集的数据 {server_alias: [gpu_info]}
|
||||||
|
|
||||||
|
def load_config(self):
|
||||||
|
"""加载本地配置文件"""
|
||||||
|
try:
|
||||||
|
with open(self.config_path, 'r', encoding='utf-8') as f:
|
||||||
|
return yaml.safe_load(f)
|
||||||
|
except Exception as e:
|
||||||
|
logger.error(f"Failed to load config file: {e}")
|
||||||
|
return None
|
||||||
|
|
||||||
|
def collect_all_servers(self):
|
||||||
|
"""遍历所有服务器并采集数据"""
|
||||||
|
config = self.load_config()
|
||||||
|
if not config:
|
||||||
|
return
|
||||||
|
|
||||||
|
servers = config.get('servers', [])
|
||||||
|
all_results = {}
|
||||||
|
|
||||||
|
for s_conf in servers:
|
||||||
|
alias = s_conf.get('alias', 'Unknown')
|
||||||
|
logger.info(f"Collecting data from {alias}...")
|
||||||
|
|
||||||
|
collector = GPUCollector(s_conf)
|
||||||
|
data = collector.fetch_gpu_data()
|
||||||
|
|
||||||
|
if data is not None:
|
||||||
|
all_results[alias] = data
|
||||||
|
else:
|
||||||
|
all_results[alias] = None # 标记为采集失败
|
||||||
|
|
||||||
|
# 更新内存状态
|
||||||
|
self.last_data = all_results
|
||||||
|
|
||||||
|
# 如果配置了 socketio,则实时推送给前端
|
||||||
|
if self.socketio:
|
||||||
|
self.socketio.emit('gpu_update', all_results)
|
||||||
|
logger.info("GPU data broadcasted via SocketIO")
|
||||||
|
|
||||||
|
def start(self):
|
||||||
|
"""启动定时任务"""
|
||||||
|
config = self.load_config()
|
||||||
|
if not config:
|
||||||
|
logger.error("Cannot start scheduler: config file missing or invalid")
|
||||||
|
return False
|
||||||
|
|
||||||
|
interval = config.get('settings', {}).get('interval', 5)
|
||||||
|
|
||||||
|
# 添加定时任务
|
||||||
|
self.scheduler.add_job(
|
||||||
|
self.collect_all_servers,
|
||||||
|
'interval',
|
||||||
|
seconds=interval,
|
||||||
|
id='gpu_collection_job'
|
||||||
|
)
|
||||||
|
|
||||||
|
self.scheduler.start()
|
||||||
|
logger.info(f"Scheduler started. Interval: {interval}s")
|
||||||
|
return True
|
||||||
|
|
||||||
|
def stop(self):
|
||||||
|
"""停止调度器"""
|
||||||
|
self.scheduler.shutdown()
|
||||||
|
logger.info("Scheduler stopped")
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
Flask
|
||||||
|
Flask-SocketIO
|
||||||
|
paramiko
|
||||||
|
PyYAML
|
||||||
|
APScheduler
|
||||||
+201
@@ -0,0 +1,201 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="zh-CN">
|
||||||
|
<head>
|
||||||
|
<meta charset="UTF-8">
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||||
|
<title>A100 GPU 实时监控系统</title>
|
||||||
|
<script src="https://cdn.tailwindcss.com"></script>
|
||||||
|
<script src="https://cdnjs.cloudflare.com/ajax/libs/socket.io/4.7.2/socket.io.js"></script>
|
||||||
|
<script src="https://cdn.jsdelivr.net/npm/chart.js"></script>
|
||||||
|
<style>
|
||||||
|
.gpu-card {
|
||||||
|
transition: all 0.3s ease;
|
||||||
|
}
|
||||||
|
.gpu-card:hover {
|
||||||
|
transform: translateY(-5px);
|
||||||
|
}
|
||||||
|
.status-dot {
|
||||||
|
height: 10px;
|
||||||
|
width: 10px;
|
||||||
|
border-radius: 50%;
|
||||||
|
display: inline-block;
|
||||||
|
margin-right: 8px;
|
||||||
|
}
|
||||||
|
.status-online { background-color: #10B981; }
|
||||||
|
.status-offline { background-color: #EF4444; }
|
||||||
|
</style>
|
||||||
|
</head>
|
||||||
|
<body class="bg-slate-900 text-slate-100 min-h-screen p-8">
|
||||||
|
<div class="max-w-7xl mx-auto">
|
||||||
|
<header class="flex justify-between items-center mb-10">
|
||||||
|
<div>
|
||||||
|
<h1 class="text-3xl font-bold text-white">GPU Cluster Monitor</h1>
|
||||||
|
<p class="text-slate-400">实时监控 NVIDIA A100 状态</p>
|
||||||
|
</div>
|
||||||
|
<div id="connection-status" class="flex items-center bg-slate-800 px-4 py-2 rounded-full text-sm">
|
||||||
|
<span id="status-dot" class="status-dot status-offline"></span>
|
||||||
|
<span id="status-text">连接中...</span>
|
||||||
|
</div>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<!-- 服务器列表区域 -->
|
||||||
|
<div id="servers-container" class="grid grid-cols-1 md:grid-cols-2 lg:grid-cols-3 gap-6">
|
||||||
|
<!-- 动态生成服务器卡片 -->
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- GPU 详情模态框 (用于展示图表) -->
|
||||||
|
<div id="gpu-modal" class="fixed inset-0 bg-black/80 hidden flex items-center justify-center z-50 p-4">
|
||||||
|
<div class="bg-slate-800 rounded-2xl max-w-4xl w-full p-6 relative">
|
||||||
|
<button id="close-modal" class="absolute top-4 right-4 text-slate-400 hover:text-white text-2xl">×</button>
|
||||||
|
<h2 id="modal-title" class="text-2xl font-bold mb-6">GPU 详情</h2>
|
||||||
|
<div class="grid grid-cols-1 lg:grid-cols-2 gap-8">
|
||||||
|
<div id="modal-metrics" class="grid grid-cols-2 gap-4">
|
||||||
|
<!-- 实时数值 -->
|
||||||
|
</div>
|
||||||
|
<div class="h-64">
|
||||||
|
<canvas id="gpuChart"></canvas>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<script>
|
||||||
|
const socket = io();
|
||||||
|
let charts = {}; // 存储每个 GPU 的图表实例
|
||||||
|
|
||||||
|
// 更新连接状态
|
||||||
|
socket.on('connect', () => {
|
||||||
|
document.getElementById('status-dot').className = 'status-dot status-online';
|
||||||
|
document.getElementById('status-text').textContent = '连接正常';
|
||||||
|
});
|
||||||
|
|
||||||
|
socket.on('disconnect', () => {
|
||||||
|
document.getElementById('status-dot').className = 'status-dot status-offline';
|
||||||
|
document.getElementById('status-text').textContent = '连接中断';
|
||||||
|
});
|
||||||
|
|
||||||
|
// 处理接收到的 GPU 数据
|
||||||
|
socket.on('gpu_update', (data) => {
|
||||||
|
const container = document.getElementById('servers-container');
|
||||||
|
|
||||||
|
for (const [serverAlias, gpus] of Object.entries(data)) {
|
||||||
|
let serverCard = document.getElementById(`server-${serverAlias}`);
|
||||||
|
if (!serverCard) {
|
||||||
|
serverCard = createServerCard(serverAlias, gpus);
|
||||||
|
container.appendChild(serverCard);
|
||||||
|
}
|
||||||
|
updateServerCard(serverCard, serverAlias, gpus);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
function createServerCard(alias, gpus) {
|
||||||
|
const div = document.createElement('div');
|
||||||
|
div.id = `server-${alias}`;
|
||||||
|
div.className = 'gpu-card bg-slate-800 rounded-2xl p-6 border border-slate-700';
|
||||||
|
div.innerHTML = `
|
||||||
|
<div class="flex justify-between items-center mb-4">
|
||||||
|
<h3 class="text-xl font-semibold">${alias}</h3>
|
||||||
|
<span class="text-xs text-slate-400">A100 Cluster</span>
|
||||||
|
</div>
|
||||||
|
<div id="gpus-${alias}" class="space-y-4"></div>
|
||||||
|
`;
|
||||||
|
return div;
|
||||||
|
}
|
||||||
|
|
||||||
|
function updateServerCard(card, alias, gpus) {
|
||||||
|
const gpuContainer = document.getElementById(`gpus-${alias}`);
|
||||||
|
if (!gpus) {
|
||||||
|
gpuContainer.innerHTML = `
|
||||||
|
<div class="flex items-center justify-center p-4 bg-slate-900/50 rounded-lg text-red-400 text-sm">
|
||||||
|
<span class="status-dot status-offline"></span> 无法连接服务器
|
||||||
|
</div>
|
||||||
|
`;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
gpuContainer.innerHTML = '';
|
||||||
|
gpus.forEach((gpu, index) => {
|
||||||
|
const utilColor = gpu.util >= 80 ? 'text-red-400' : (gpu.util >= 50 ? 'text-yellow-400' : 'text-green-400');
|
||||||
|
gpuContainer.innerHTML += `
|
||||||
|
<div class="p-3 bg-slate-900/50 rounded-lg border border-slate-700 hover:bg-slate-700 cursor-pointer transition-colors"
|
||||||
|
onclick="openGpuDetail('${alias}', ${index})">
|
||||||
|
<div class="flex justify-between items-center mb-2">
|
||||||
|
<span class="text-sm font-medium text-slate-300">GPU ${gpu.index} [${gpu.name}]</span>
|
||||||
|
<span class="text-sm font-bold ${utilColor}">${gpu.util}%</span>
|
||||||
|
</div>
|
||||||
|
<div class="w-full bg-slate-700 h-2 rounded-full overflow-hidden">
|
||||||
|
<div class="bg-green-500 h-full transition-all duration-500" style="width: ${gpu.util}%"></div>
|
||||||
|
</div>
|
||||||
|
<div class="flex justify-between text-xs text-slate-400 mt-2">
|
||||||
|
<span>Temp: ${gpu.temp}°C</span>
|
||||||
|
<span>Mem: ${gpu.mem_used}/${gpu.mem_total} MiB</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
`;
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function openGpuDetail(serverAlias, gpuIndex) {
|
||||||
|
const modal = document.getElementById('gpu-modal');
|
||||||
|
const title = document.getElementById('modal-title');
|
||||||
|
const metrics = document.getElementById('modal-metrics');
|
||||||
|
const chartCanvas = document.getElementById('gpuChart');
|
||||||
|
|
||||||
|
modal.classList.remove('hidden');
|
||||||
|
title.textContent = `Server: ${serverAlias} - GPU ${gpuIndex}`;
|
||||||
|
|
||||||
|
metrics.innerHTML = `
|
||||||
|
<div class="p-4 bg-slate-900 rounded-lg text-center">
|
||||||
|
<span class="text-slate-400 text-xs block mb-1">GPU 利用率</span>
|
||||||
|
<span class="text-2xl font-bold text-green-400">实时更新中...</span>
|
||||||
|
</div>
|
||||||
|
<div class="p-4 bg-slate-900 rounded-lg text-center">
|
||||||
|
<span class="text-slate-400 text-xs block mb-1">GPU 温度</span>
|
||||||
|
<span class="text-2xl font-bold text-green-400">实时更新中...</span>
|
||||||
|
</div>
|
||||||
|
<div class="p-4 bg-slate-900 rounded-lg text-center">
|
||||||
|
<span class="text-slate-400 text-xs block mb-1">显存占用</span>
|
||||||
|
<span class="text-2xl font-bold text-green-400">实时更新中...</span>
|
||||||
|
</div>
|
||||||
|
<div class="p-4 bg-slate-900 rounded-lg text-center">
|
||||||
|
<span class="text-slate-400 text-xs block mb-1">GPU 频率</span>
|
||||||
|
<span class="text-green-400 text-2xl font-bold">实时更新中...</span>
|
||||||
|
</div>
|
||||||
|
`;
|
||||||
|
|
||||||
|
if (charts.gpuChart) {
|
||||||
|
charts.gpuChart.destroy();
|
||||||
|
}
|
||||||
|
|
||||||
|
const ctx = chartCanvas.getContext('2d');
|
||||||
|
Chart.defaults.color = '#94a3b8';
|
||||||
|
charts.gpuChart = new Chart(ctx, {
|
||||||
|
type: 'line',
|
||||||
|
data: {
|
||||||
|
labels: [],
|
||||||
|
datasets: [{
|
||||||
|
label: 'GPU Utilization (%)',
|
||||||
|
borderColor: '#10B981',
|
||||||
|
backgroundColor: 'rgba(16, 185, 129, 0.1)',
|
||||||
|
fill: true,
|
||||||
|
tension: 0.4,
|
||||||
|
data: []
|
||||||
|
}]
|
||||||
|
},
|
||||||
|
options: {
|
||||||
|
responsive: true,
|
||||||
|
maintainAspectRatio: false,
|
||||||
|
scales: {
|
||||||
|
y: { beginAtZero: true, max: 100 }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
document.getElementById('close-modal').onclick = () => {
|
||||||
|
document.getElementById('gpu-modal').classList.add('hidden');
|
||||||
|
};
|
||||||
|
</script>
|
||||||
|
</body>
|
||||||
|
</html>
|
||||||
在新工单中引用
屏蔽一个用户