Files
R-clone-setup/CLAUDE.md
2026-08-18 11:31:53 +05:30

310 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# R-clone — Rclone→Synology Sync Monitoring Project
## Mission
The office deploys client machines (Linux, GPU) that process road-survey videos/images
and sync results to a Synology NAS via rclone. Nothing tracks whether syncs succeed.
**Goal: build a monitoring UI + alert system** for the fleet.
My stack: MERN (Node/Express/React/Mongo), FastAPI, intermediate Python, basic Linux.
## Production architecture (from office files — do NOT commit real creds)
Each client site machine runs (via a generated docker-compose):
- `algorithm3` — GPU/YOLO container, processes video → anomaly images + result CSVs
- `mysql`, `socket-server`, `local-gui-backend` — local site stack
- `watchtower` — auto-pulls new images every 300s from the private registry
- `rclone-synology-sync` — Alpine rclone image + `sync.sh` loop (this repo has exact copies)
### The sync loop (`sync-container/sync.sh` — byte-identical to production)
- Writes rclone.conf for an **SFTP remote** (`synodrive`) → Synology at a public
hostname, non-standard port, single shared user (creds in office compose only).
- Starts `rclone rcd --rc-web-gui` on **:5572** (RC API + built-in web GUI).
- Infinite loop: for each `SYNC_n=direction:local:remote` env var, run
`rclone copy --size-only` (up = local→NAS, down = NAS→local), then sleep
`SYNC_INTERVAL` seconds (**seconds — the "# minutes" comment in prod compose is wrong;
effective prod interval is 120s**).
- Logs to `/logs/`: `sync.log` (everything), `stats.log` (pretty per-round),
`errors.log` (grep " ERROR "), `completed.log` (grep "Copied|Moved"), `rc.log`.
### Production sync pairs (per site; ORG/SITENAME templated)
| # | Dir | What |
|---|-----|------|
| 1 | up | anomaly images → `Saudi_Video_Sync/SeekRight/Anomaly/<ORG>/<SITE>/TEST` |
| 2 | down | master sheets ← `ClientSync/ALGORITHM_DEPLOYMENT/Master_Sheets/<ORG>/<SITE>` |
| 3 | down | yolov8 weights ← `ClientSync/ALGORITHM_DEPLOYMENT/Models/yolov8` |
| 4 | down | NIGHT weights ← `ClientSync/ALGORITHM_DEPLOYMENT/Models/NIGHT` |
| 5 | up | result CSVs → `ClientSync/ALGORITHM_DEPLOYMENT/csv_files/<ORG>/<SITE>` |
**Monitoring focus = pairs 1 & 5** (client-produced files that must reach the NAS).
### Known production issues (verified, mention when relevant)
1. `RC_PASSWORD` never set in prod compose → **unauthenticated RC API** on host
network :5572 (full read/write to NAS + local disk from the LAN).
2. Plaintext creds in compose + git token in `generate_compose.sh` → flagged for rotation.
3. `--size-only` + `copy` (not `sync`): same-size edits never re-transfer; deletes
never propagate. "Synced" is weaker than it sounds.
4. **The built-in Web GUI shows nothing useful** — see below.
## KEY FINDING: rclone's built-in UI exists but can't see the syncs
`rclone rcd --rc-web-gui` serves a React GUI on :5572 (login admin/$RC_PASSWORD).
**Verified locally: while 5 pairs were actively copying, `POST /core/stats` returned
all zeros.** The `rclone copy` commands are separate OS processes; the rcd daemon only
reports jobs started through its own API. So the prod GUI is decorative.
Options for the real project:
- (A) Ship our own agent that tails logs + POSTs to a central collector (recommended;
zero change to sync behavior), or
- (B) Rewrite sync.sh to submit jobs via `rclone rc sync/copy _async=true` so rcd/GUI
sees them — bigger change, still no fleet view, history, or alerts.
- Adding `--use-json-log` to the rclone copy invocation would make log parsing trivial
(structured JSON per line) — one-flag change worth proposing.
## This repo = the production monitoring stack
```
docker-compose.yml mongo + collector — the permanent stack (port 4400)
monitor/collector/ FastAPI + pymongo + rclone; serves API and the built UI
monitor/agent/ stdlib-only sidecar deployed on each client machine
frontend/ React 18 + Vite dashboard → dist/ bind-mounted into collector
sync-container/ byte-identical prod sync.sh + Dockerfile (reference + used
by deploy/remote-machine builds) — never modify
deploy/ install bundles: agent-only/ (machine already runs prod
sync) and remote-machine/ (needs sync + agent)
send-to-friend/ agent-deploy.zip — the bundle actually on JAGAN-TEST-01
.backups/ pre-v2 versions of main.py/App.jsx/styles.css/.env
```
**The simulation lab was REMOVED 2026-08-10** (user decision: not needed for
production). `docker-compose.sim.yml`, `generator/`, `lan-forward.ps1` are in
git history; `client/`+`client2/`+`server-data/` were gitignored fake data and
are gone for good. Consequence: no local test fleet — alert paths can now only
be exercised against a real machine. The 2026-08-05/06 sections below describe
that lab; keep them as history, don't try to run them.
### Run / operate
```bash
cd /opt/R-clone/R-clone-setup
sudo docker-compose up -d --build # mongo + collector (permanent stack)
sudo docker-compose logs -f collector
sudo docker-compose down
```
**This host has docker-compose v1 (the hyphenated binary) only** — the v2
`docker compose` plugin is NOT installed and fails with
`unknown shorthand flag: 'd'`. Always write `docker-compose`.
### Environment (Linux Mint host — MOVED off Windows/WSL 2026-08-07)
The project now lives at **`/opt/R-clone/R-clone-setup`** on a Linux Mint box
(`vm3@vm3-mint`). Every `/mnt/c/...` path and WSL workaround in older notes is
dead — including `RCLONE_SFTP_SET_MODTIME=false`, which only existed because
`/mnt/c` rejects `utimes()`. Current realities:
- **Repo was root-owned**; `chown -R vm3:vm3 /opt/R-clone` ran 2026-08-07.
Without it, no edits, no `npm run build`, and git throws "dubious ownership".
- **docker needs `sudo`** — `vm3` is not in the `docker` group.
- **Ports 4000 and 8000 on this host are already taken** by unrelated apps
(a Node app in `~/seekright-video-calendar`, and another FastAPI service).
That is why the collector moved to **4400**. Don't move it back.
- **Backend vs frontend rebuild rules differ**: `main.py` is `COPY`d into the
image → a change needs `sudo docker-compose up -d --build collector`.
`frontend/dist` is bind-mounted → `npm run build` alone is enough, no restart.
- **Collector host LAN IP `192.168.1.201/22`** (so the /22 spans .0.3.255).
It is DHCP — it already changed once and broke every remote agent. Reserve it.
## Verified: RC API gives every UI metric (2026-08-05 experiment)
`monitor/watch_batch.py` submitted 30×5MB files via `POST /sync/copy _async=true`
(src `./client/batch_test`, throttled with `core/bwlimit`) and polled live:
- **scheduled** = `core/stats.totalTransfers` · **done** = `.transfers`
- **moving now** = `.transferring[]` (name, %, speed, per-file eta; max 4 shown = `--transfers`)
- **queued** = totalTransfers transfers len(transferring)
- **overall ETA/speed/bytes** = `.eta`, `.speed`, `.bytes/.totalBytes`
- **per-file duration** = `core/transferred[]` `completed_at started_at` (+ error field)
All 30 arrived; per-file report printed. Gotchas: per-file `eta` can be `None`
early; totalTransfers grows during discovery (not instantly 30); always pass
`{"group": "job/<jobid>"}` to scope stats to one job.
**Constraint**: RC only sees jobs submitted via RC. The prod loop's `rclone copy`
processes are invisible to it → for live per-file UI data in prod, sync.sh v2
should submit pairs via `rclone rc sync/copy _async=true` + poll, instead of
spawning `rclone copy`. Log parsing alone gives only aggregate 5s stats lines
+ completion events.
## Monitoring stack v1 — BUILT and running (2026-08-06)
`agent (per machine) → collector (FastAPI+Mongo) → React dashboard`, all in compose:
- `monitor/agent/agent.py` — stdlib-only sidecar; tails `/logs/sync.log` (ro mount),
regex-parses rounds/pairs/files/errors/progress, POSTs `/api/ingest` every 5s
(empty POST = heartbeat). Works against byte-identical prod sync.sh.
- `monitor/collector/main.py` — FastAPI + pymongo (**tz_aware=True — naive-vs-aware
datetime bug otherwise**). Collections: machines, pairs, events. Endpoints:
POST `/api/ingest`, GET `/api/overview` (fleet + computed alerts), GET
`/api/machines/{id}/events?type=&limit=`. Serves React build from `/app/static`.
Alert rules: offline (no heartbeat >30s, critical), pair last_status=fail
(serious), errors_1h>0 (warning).
- `frontend/` — React 18 + Vite. Fleet cards (status badge, stat tiles, pair rows,
live progress bar), alerts bar, click card → events panel with tabs
(Files/Errors/Rounds/Pair results). Polls every 4s. Dark theme, dataviz status
tokens (#0ca30c/#fab219/#ec835a/#d03b3b), icon+label never color alone.
Build: `cd frontend && npm run build` (dist/ is bind-mounted into the
collector — rebuild frontend = just `npm run build`, no docker rebuild).
- UI: **http://localhost:4400** · do not stop/restart containers without
asking first. (The RIYADH-01 second-machine demo died with the sim lab.)
## v1.1 — remote machines + RMM surface (2026-08-06)
- **Agent auth**: `AGENT_TOKEN` shared secret (in `.env`); collector rejects
ingest without `Authorization: Bearer <token>` (verified 401). Empty token =
auth off (never for remote use).
- **RMM endpoints**: GET `/api/health` (liveness), GET `/api/alerts`
(lightweight polling), CORS enabled (`CORS_ORIGINS` env). Integration points
for the user's future RMM: poll alerts, or embed /api/overview data.
- **`deploy/remote-machine/`**: self-contained bundle (compose + .env.example)
for any external machine (e.g. a friend's system): rclone-sync (1 up pair from
the algorithm's OUTPUT_DIR) + agent → user's collector. Remote machine needs a
network path to the collector — recommend Tailscale; COLLECTOR_URL then is the
tailnet IP. Bundle builds from ../../sync-container and ../../monitor/agent.
- Transfer visibility today: completed files ✓ (file_synced events), aggregate
live progress per pair ✓ (5s stats → progress bar), per-file in-flight +
pending queue ✗ — needs RC-submitted jobs (sync.sh v2, see RC section).
## v1.2 — server-side verification / audit feature (2026-08-06) — WORKING
The audit-team answer: "did the file actually reach the Synology?" proven from
the collector, no one logs into anything.
- Agent `WATCH_DIRS` env: mounts up-pair source dirs ro at the same container
paths as sync (`/sources/...`), ships file inventory (name/size/mtime) with
every ingest.
- Collector: own rclone remote (start.sh writes config; rclone installed via
apt in Dockerfile), background thread `rclone lsjson --recursive` on every
up-pair remote path every VERIFY_INTERVAL (45s local; use 300s+ in prod).
`pending` = local file not in server listing; `missing` = pending with mtime
older than MISSING_GRACE_S (600) → serious alert with file names.
- UI: 4th tile "pending → server" + amber `N pending` / red `N missing!` chips
per pair. Verified live: pending counts rise as generator drops files, fall
to 0 after each sync round.
- Friend's machine (compose seen 2026-08-06): logs at
`/home/testing/JAGAN/Prerequisites/Sync_logs`, SITENAME=TEST, standard 5
pairs. deploy/remote-machine bundle fits as-is; needs Tailscale (or LAN) to
reach the collector + AGENT_TOKEN. Nothing needed from his algorithm code —
agent reads only Sync_logs + source dirs.
## v2 — upload audit + windowed history (2026-08-07 → 10) — CURRENT
Prompted by: "admins audit UPLOADS, not downloads; need pending/in-progress/
completed and processed-counts per hour/day/3d/week/month, ~1 month of data."
Researched: rclone's native Prometheus metrics can't help (same single-rcd
blindness as the GUI + unscrapeable short-lived syncs); MFT dashboards use the
Completed/In-Progress/Queued/Failed taxonomy → adopted.
### Collector additions (`monitor/collector/main.py`)
- **Retention**: Mongo TTL index on `events.received_at`, `RETENTION_DAYS` env
(30). `_ensure_retention` falls back to `collMod` when the window changes.
- **GET /api/stats** — one aggregation pass over 30d computes per-machine AND
fleet rollups for windows 1h/24h/3d/7d/30d: uploaded, errors, failures,
rounds, bytes, avg_bps. Machines with zero activity still appear (that's the
signal admins need). ~15s poll from the UI.
- **GET /api/timeseries?range=&machine_id=** — epoch-aligned buckets
(5min/1h/3h/6h/1d per range), zero-filled, uploads + errors only.
- **GET /api/machines/{id}/events** — added `direction=`, `q=` (regex-escaped
search over file/message/pair), `skip/limit` paging, returns `total`.
- **DELETE /api/machines/{id}** (agent-token auth) — retire a machine; without
it a decommissioned machine alerts "offline" forever. Used to remove the
CONNTEST test entry 2026-08-10.
- `round_summary` now splits `files_up`/`files_down`/`bytes_up` (download bytes
are unknowable — files land outside WATCH_DIRS; never fake them).
- `_machine_view`: `up_pairs`/`down_pairs` split, `in_progress`, `failed_up`,
`uploads_1h` (direction-filtered). Alerts unchanged EXCEPT download-pair
failures now say "download pair failing" (still alerted, not shown as cards).
- **Verification defaults to the real NAS**: `DEFAULT_REMOTE` env (compose sets
`synoreal`), `VERIFY_ENABLED` true if either SYNOLOGY_HOST or
REAL_SYNOLOGY_HOST is set. Old behavior (default `synodrive` = the deleted
fake NAS) silently skipped verification for any site not in SITE_REMOTES.
### UI (v2.1 after a misstep — `frontend/src/App.jsx`)
First v2 cut replaced the machine cards with range-dependent tiles and dropped
download pair rows + the last-round tile. **User: "not understandable, it was
meaningful before."** Lesson recorded: the last-round-ago tile and the full
pair list ARE the at-a-glance value; window counters reading 0 look like a
broken dashboard. Current layout:
- **Cards (top, range-INDEPENDENT, restored)**: last round tile (amber when
>30min stale), uploaded (1h), queued→NAS (pending), errors (1h), live
progress bar, ALL 5 pair rows (up first), missing/pending chips.
- **"Upload history" section (below, driven by the range selector)**: KPI row
(uploaded / failed+errors / rounds / avg throughput+bytes), uploads-per-bucket
bar chart with hover tooltip + "view as table", per-machine comparison table
(the per-client speed audit). Chart empty state names the likely cause
(stopped sync container) instead of rendering all-zero bars.
- Palette validated (dataviz skill): uploads #3987e5 / errors #d03b3b on
#1a1a19 pass all six checks; separate stacked plots, never dual-axis.
### Ops facts (2026-08-09/10)
- JAGAN-TEST-01's sync container was stopped 2026-08-07 during load-test
cleanup and left stopped for 65h — the stalled-round alert caught it; the
"empty graph" complaint was exactly this. Agent online ≠ sync running.
- His Master_Sheets down pair fails exit 3: remote path
`.../Master_Sheets/TEST/TEST` doesn't exist on the NAS (ORG=SITE=TEST
substitution) — create the folder or fix his compose vars.
- Load-test cleanup order that WORKS: stop sync container → delete local
source files → `rclone purge` the NAS folder → start container. NAS-first
gets re-uploaded within one round (learned the hard way).
- `.env` trimmed to AGENT_TOKEN + REAL_SYNOLOGY_* + ORG/SITENAME.
send-to-friend/agent-deploy.zip rebuilt (old zip still had :8000).
## v2.1 — public exposure + one-line installer (2026-08-11)
- **Public URL**: https://rclone.seekright.com → Synology reverse proxy →
this VM :4400 (same pattern as the RMM hostnames). `PUBLIC_URL` in .env.
- **Viewer auth**: HTTP Basic middleware (VIEWER_USER/VIEWER_PASS in .env)
gates EVERYTHING except /api/ingest, /api/health, and requests carrying the
agent bearer token. Unset VIEWER_PASS = gate off. Browser prompts once.
- **One-line installer** (RMM deploy_agent.sh pattern, leak-lesson applied —
download is authenticated, live AGENT_TOKEN injected at serve time,
split-sentinel guard, exactly-once placeholder check):
```
curl -fsSL -u admin "https://rclone.seekright.com/deploy_agent.sh" -o /tmp/deploy_agent.sh
sudo bash /tmp/deploy_agent.sh <MACHINE_ID> [SITE] # then rm the file (holds live token)
```
Auto-discovers /logs + upload-source host paths from the machine's
rclone-synology-sync container (docker inspect), fetches agent.py from
GET /agent.py (single source of truth — no embedded copy to drift), writes
/opt/rclone-agent/, replaces any previous rclone-agent container, starts.
Re-run with no args = in-place upgrade keeping identity. MACHINE_ID is
REQUIRED on fresh installs (no hostname fallback — RMM's stale-entry lesson).
- Collector build context is now ./monitor (Dockerfile at collector/Dockerfile)
so the image carries agent/agent.py to serve.
## v2.2 — REAL fleet sync.sh has drifted from our copy (found 2026-08-11)
First real office machine (GCBOT-1 / IRBLAMBTHAM001, image from the private
registry) revealed the production sync.sh is NOT our byte-copy anymore:
- **Up-pairs are change-detected**: skipped unless file count or `du -sb` size
changed; logged as `SKIP: [↑ UP] <label> — unchanged (idle Ns)`. A forced
sync runs after FORCE_SYNC_AFTER=1800s idle (" reason: …" line follows
pair start). Down pairs unchanged.
- Rounds every **180s** (CHECK_INTERVAL), not 120.
- Consequences: up-pairs show activity at most every 30min when idle; a pair
can be absent from the dashboard until its first non-skipped run unless
skips are parsed. Agent v2 parses SKIP → `pair_skip` event; collector
records last_skip_at/last_idle_s and creates the pair as status "idle"
(NEVER overwrites fail — state files update before rclone runs, so a failed
up-pair is followed by skips until content changes). UI shows idle as ok.
- VERIFY_TIMEOUT env (default 300s) — 8.5k-file SFTP listings exceed 120s.
- MAX_INV_FILES now env-tunable (default 2000; GCBOT image tree is 8.5k).
- Jagan's machine runs OUR old sync.sh (bundle-built) — no SKIP lines there.
- TODO: replace sync-container/sync.sh with the real fleet version (get full
file from a machine: `docker exec rclone-synology-sync cat /sync.sh`).
## Roadmap for the actual deliverable
1. **Agent** (per client machine, sidecar container): tail `/logs/*.log` (or JSON log),
parse rounds/pairs/files/errors, POST events + 60s heartbeat to collector.
2. **Collector**: FastAPI + MongoDB. Models: Machine, SyncPair, Round, FileEvent, Alert.
3. **Server-side verification** (the feature logs can't provide): collector has its own
rclone SFTP remote to the NAS; `rclone lsjson` per site path to confirm reported
uploads actually arrived; alert if a file is still missing after a grace period.
Missed-heartbeat = dead machine alert (most important signal).
4. **UI**: React dashboard — fleet grid (last round, last heartbeat, error counts),
per-machine pair detail, file history, live progress later via RC API.
5. **Alerts**: dashboard + email first; Slack/Teams webhook later. Rules before channels.
## Conventions for this repo
- No real office credentials/hostnames in tracked files, with ONE deliberate
exception: `.env` holds the real `REAL_SYNOLOGY_*` values and IS committed
(see .gitignore header — private repo, creds needed on the VM after a pull).
Everything else uses placeholders like `<ORG>/<SITE>`.
- `sync-container/sync.sh` stays byte-identical to production — replica fidelity is
the point. Fixes go in compose env vars or new sidecar services, and proposed
prod changes get documented in this file instead.
- RMM integration is PARKED (user decision 2026-08-10, R-clone fixes first).
When resumed: this VM is also the production RMM server (26 field machines) —
read `/opt/rmm-backend/CLAUDE.md` before touching anything there. Plan agreed:
additive read-only proxy endpoint in rmm-backend + new rmm-ui page, developed
against a scratch uvicorn (`MONGODB_DB=rmm_test`, spare port) and a scratch
UI copy — NEVER by editing `/opt/rmm-ui/src` in place (vite hot-reloads
straight into production).