by Mia'a AI Lab
sparkDash is a real-time web dashboard for one or more NVIDIA DGX Spark (GB10) machines in a single browser window. It streams GPU, CPU, unified memory, storage, network, and local LLM metrics — and lets you add, edit, reorder, or remove Sparks from the UI without restarts or code changes.
- Latest version changelog
- Features
- Full changelog
- Quick start
- Architecture
- Tech stack
- Repository layout
- REST API
- Configuration
- Security
- Scripts
- How it works
- Contributing
- License
- LLM decode benchmark — Run decode benchmark on each LLM panel; multi-select concurrency (
1–32, default 1, 2); levels run one-by-one with concurrent streams per level - Bench metrics — Server tok/s (engine counters, aligned with live Generation tok/s) and per-stream decode after first token; TTFT; last run restored after refresh/restart
- GB10 GPU metrics — host
gpu-memory.sh+ process list fallback when Docker /nvidia-smiapps are empty - Mobile dialogs — solid scrollable sheets for benchmark, Edit Spark, and Add Spark
- vLLM ops tiles — KV cache %, run/wait requests, TTFT p95, cumulative preemptions (vLLM only; info tooltips on each)
- Spark roles — Head / Worker / Standalone in Edit Spark; worker label + head picker; standalone LLM monitoring opt-in; role badges on Overview and Spark header
- Reliable shutdown — background remote shutdown so the UI gets a real response (no “Failed to fetch” when the host drops SSH)
Full history: CHANGELOG.md
| Area | What you get |
|---|---|
| Multi-unit | Any number of Sparks; each has a tabbed detail page plus a shared Overview |
| Live streaming | WebSocket metrics with configurable poll intervals |
| Local + remote | Host metrics via sysfs/proc/nvidia-smi; remotes over SSH (key or password) |
| LLM probe | Auto-detects llama.cpp, vLLM, or sglang; live tok/s per server |
| Decode benchmark | Multi-concurrency streaming decode tok/s (server + per-stream), persisted last run |
| vLLM health | KV cache %, run/wait queue, TTFT p95, preemptions from Prometheus /metrics |
| Multiple LLM ports | Monitor several LLM servers on different ports simultaneously — each gets its own panel with independent backend detection and metrics |
| GPU processes | See the top GPU processes by VRAM usage directly in the GPU panel, including process name and memory allocation |
| Spark uptime | System uptime displayed inline on each Spark header for at-a-glance availability |
| Power controls | Graceful shutdown (SSH host script) and Wake-on-LAN; batch actions on Overview |
| Spark roles | Head / Worker / Standalone — worker label + head link; standalone can disable LLM monitoring |
| Unified memory | GB10 128 GB LPDDR5X pool (~273 GB/s), GPU/CPU split, bandwidth via nvidia-smi dmon |
| Themes | Dark, light, cool white, OLED — neutral palettes, persisted in localStorage |
| Secrets | SSH passwords AES-256-GCM encrypted; never in sparks.json or API responses |
| Docker-first | Single privileged container for host metrics; prod and dev Compose files |
| Hot config | Add / edit / remove / reorder Sparks from the UI with no process restart |
git clone https://github.com/MiaAI-Lab/sparkDash.git
cd sparkDash
# Production (Docker)
docker compose up --build -d
# Or development (host, with hot reload)
npm install
npm run dev- Docker: open http://<host-ip>:5555 (arm64 image, auto-restart, host mounts for GPU/metrics access)
- Dev: Vite on http://localhost:5173 (proxies API/WS to Express)
For development with Docker (source-mounted, HMR):
docker compose -f docker-compose.dev.yml up --buildDesign principle: one Spark model, N instances. Every Spark is a record in config/sparks.json. The same SparkMonitor, SystemCollector, and LlmProbe code runs for all of them. Adding a unit is a config change, not a code change.
┌────────────────────── Docker container (sparkDash) ──────────────────────┐
│ Express (server/) │
│ ├─ config/sparks.json Spark registry (API read/write) │
│ ├─ SparkRegistry load/persist Sparks; change events │
│ ├─ SparkMonitor (per Spark) collector + LLM probe + rate baselines │
│ │ ├─ SystemCollector local sysfs/proc OR remote SSH │
│ │ └─ LlmProbe HTTP to host:LLM_PORT, backend autodetect │
│ ├─ REST /api/* │
│ └─ WebSocket /ws snapshot stream to browsers │
│ React SPA (src/) — Overview + per-Spark pages, themes, dialogs │
└────────────────────────────────────────────────────────────────────────────┘
│ SSH (key or sshpass) │ HTTP :8888
▼ ▼
remote Spark(s) each Spark’s LLM serverBrowser ←→ WebSocket /ws ←→ SparkMonitor.snapshot() ←→ collectors
Browser ←→ REST /api/* ←→ SparkRegistry + SparkMonitorPoll loops run in the background (even with no clients) so rate metrics — tokens/s, network bytes/s, disk I/O — stay correct.
| Layer | Stack |
|---|---|
| Frontend | React 19, TypeScript, Vite 8, Tailwind CSS v4 |
| Backend | Node.js (ESM), Express 5, ws |
| Platform | ARM64 — DGX Spark GB10 (Neoverse V2) |
| Deploy | Docker multi-stage (arm64), Compose |
| Secrets | AES-256-GCM SSH password store |
| Ports | 5555 dashboard/API; 5173 Vite (dev only) |
sparkDash/
├── src/ React + TypeScript SPA
│ ├── api/ REST client + shared types
│ ├── components/ Overview, Spark pages, dialogs, UI primitives
│ ├── hooks/ WebSocket snapshot, routing
│ └── theme / CSS Tailwind v4 + four themes
├── server/ Express + WebSocket (plain JS ESM)
│ ├── sparks/ SparkRegistry, SparkMonitor
│ ├── collectors/ SystemCollector, LlmProbe, ssh
│ ├── secretsStore.js Encrypted password persistence
│ └── validate.js Host/user validation (SSRF-minded)
├── config/ Runtime state (volume; secrets gitignored)
├── assets/ Screenshots
├── Dockerfile Production multi-stage arm64
├── docker-compose.yml Production
├── docker-compose.dev.yml
└── deploy.sh Rebuild / recreate helpers| Method | Path | Purpose |
|---|---|---|
| GET | /api/sparks |
List Sparks (passwords redacted) |
| POST | /api/sparks |
Add Spark and start its monitor |
| PATCH | /api/sparks/:id |
Update Spark (hot-swap config) |
| DELETE | /api/sparks/:id |
Remove Spark and drain monitor |
| PUT | /api/sparks/order |
Persist tab order |
| GET | /api/sparks/:id/metrics |
One-shot metrics snapshot |
| POST | /api/sparks/test |
Ephemeral SSH + LLM test (no persist) |
| POST | /api/sparks/:id/test |
Connectivity test (can save password) |
| PUT | /api/sparks/:id/password |
Save SSH password (works offline) |
| PUT | /api/sparks/:id/disabled-devices |
Hide storage devices (hot) |
| PUT | /api/sparks/:id/disabled-interfaces |
Hide network interfaces (hot) |
| PUT | /api/sparks/:id/llm-ports |
Replace all LLM ports (hot) |
| POST | /api/sparks/:id/llm-ports |
Add an LLM port (hot) |
| DELETE | /api/sparks/:id/llm-ports/:port |
Remove an LLM port (hot) |
| PUT | /api/sparks/:id/llm-port |
LLM port — backward-compat (hot) |
| GET | /api/settings |
Global settings |
| PUT | /api/settings |
Update global settings |
| WS | /ws |
Real-time metrics stream |
There is no authentication on the HTTP/WebSocket API. Run sparkDash only on a trusted network (or behind your own reverse proxy with auth).
Gear icon in the header, or GET/PUT /api/settings:
| Setting | Default | Description |
|---|---|---|
| Poll interval | 2000 ms | WebSocket broadcast interval (minimum 1000 ms) |
| Default LLM port | 8888 | Default for new Sparks |
| Auto-hide offline | false | Hide offline Sparks on Overview |
| Temperature unit | Celsius | Display GPU temperature in °C or °F |
Copy .env.example to .env if needed:
| Variable | Default | Description |
|---|---|---|
PORT |
5555 |
HTTP + WebSocket listen port |
LLM_PORT |
8888 |
Default LLM probe port |
POLL_INTERVAL_GPU |
2000 |
GPU poll (ms) |
POLL_INTERVAL_CPU |
2000 |
CPU / RAM poll (ms) |
POLL_INTERVAL_NETWORK |
2000 |
Network poll (ms) |
POLL_INTERVAL_STORAGE |
5000 |
Storage poll (ms) |
POLL_INTERVAL_LLM |
2000 |
LLM probe poll (ms) |
POLL_INTERVAL_BANDWIDTH |
2000 |
Memory bandwidth / dmon poll (ms) |
POLL_INTERVAL_LIVENESS |
5000 |
Online/SSH liveness check (ms) |
SPARKDASH_SECRETS_KEY |
(auto) | Passphrase or 64-char hex for secret encryption |
HOST_PROC_PATH |
/host/proc |
Host proc mount inside container |
HOST_SYS_PATH |
/host/sys |
Host sys mount |
HOST_ROOT_PATH |
/host/root |
Host root mount |
- Open the + tab.
- Set Name, LAN IP (required), optional CX7 IP, SSH user, and auth (key or password). Wake-on-LAN MAC is auto-read from enP7s7 when online (optional override in Edit).
- Test Connection for SSH + LLM reachability.
- Save — a tab appears and metrics start streaming.
- Shutdown (per Spark or Shutdown All on Overview) runs over SSH:
sudo -n /usr/local/bin/spark-shutdown
Install that script on each Spark and allow passwordless sudo for it only. - Wake / Wake All send a UDP magic packet (port 9). The MAC is taken from the enP7s7 interface automatically while the Spark is online (persisted as
detectedMacAddress). Optionally set a MAC override in Edit Spark. Broadcast is derived as/24from LAN IP, or255.255.255.255if LAN IP is missing. - Batch shutdown only targets online Sparks; offline nodes are skipped.
- Same trust model as the rest of the API: do not expose port 5555 beyond a trusted network — power actions are not separately authenticated.
Header theme control cycles:
| Theme | Notes |
|---|---|
| Dark (default) | Neutral grays, true black base, muted amber accent |
| Light | Warm paper whites |
| White | Cool neutral whites |
| OLED | True black for OLED panels |
Choice is stored in localStorage.
- SSH passwords are not stored in
sparks.jsonand are never returned by the API. - Passwords are encrypted with AES-256-GCM in
config/sparks-secrets.json(survives restarts). - Encryption key:
config/.secrets-key(auto-generated) orSPARKDASH_SECRETS_KEY. Do not delete the key file or encrypted secrets become unreadable. - Target validation rejects clearly unsafe IPv4 targets (link-local
169.254.0.0/16,0.0.0.0/8, multicast/reserved ≥ 224). Private, loopback, and public addresses are allowed so LAN and remote Sparks work. - SSH and HTTP probes use short timeouts (about 5 s SSH connect, 3 s HTTP) so a hung host cannot stall the poll loop.
- Prefer SSH keys over passwords.
- Treat the dashboard as LAN-trusted: the API is intentionally unauthenticated for ease of use on a private network. That includes power APIs (shutdown / Wake-on-LAN): anyone who can reach the dashboard can request fleet power actions.
| Command | Purpose |
|---|---|
npm run dev |
Vite (5173) + Express (5555) together |
npm run dev:server |
Express only (node --watch) |
npm run dev:client |
Vite only |
npm run build |
Production frontend → dist/ |
npm run typecheck |
tsc --noEmit |
npm start |
Production server (node server/index.js) |
npm run docker:up |
docker compose up -d |
npm run docker:prod |
Same as docker:up |
npm run docker:rebuild |
docker compose up --build -d |
npm run docker:dev |
Dev Compose |
npm run docker:dev:build |
Dev Compose with rebuild |
./deploy.sh |
Recreate container; --build, --frontend flags |
One SystemCollector path for both modes. When spark.isLocal is true, metrics come from host sysfs/proc and nvidia-smi (often via nsenter into the host namespace). Remote Sparks wrap the same commands in a shared sshExec() helper (key agent or sshpass).
Collectors catch errors and return zero/default metrics instead of crashing the loop. After sustained liveness failures, a Spark is marked offline; the UI shows stale or empty states rather than hard errors.
Name, IP, SSH credentials, LLM port, and device/interface filters update the running SparkMonitor without tearing down poll loops or losing rate baselines. Registry writes are atomic (temp file + rename).
Each configured LLM port gets its own LlmProbe instance running in parallel. Probes auto-detect backends:
- llama.cpp —
/slotsfor live decode rates; model from/props - vLLM / sglang —
/v1/models; sglang via/get_server_info, vLLM via Prometheus/metricscounters (scientific notation supported)
Rates are derived from per-probe cumulative counter diffs. Multiple ports can be added or removed at runtime without restarting the monitor.
Contributions are welcome. Conventions:
- Server: plain JavaScript ESM
- Client: TypeScript + React
- Prefer extending the shared Spark model over per-unit special cases
MIT — Copyright (c) 2026 Mia'a AI Lab
- Built for the NVIDIA DGX Spark (GB10) on ARM64
- Rebuilt from a legacy multi-unit dashboard with a single shared Spark model (no copy-pasted “Spark N” code paths)
- LLM probe behavior refined from production monitoring experience
