Files
R-clone-setup/CLAUDE.md
2026-08-18 11:31:53 +05:30

20 KiB
Raw Permalink Blame History

R-clone — Rclone→Synology Sync Monitoring Project

Mission

The office deploys client machines (Linux, GPU) that process road-survey videos/images and sync results to a Synology NAS via rclone. Nothing tracks whether syncs succeed. Goal: build a monitoring UI + alert system for the fleet. My stack: MERN (Node/Express/React/Mongo), FastAPI, intermediate Python, basic Linux.

Production architecture (from office files — do NOT commit real creds)

Each client site machine runs (via a generated docker-compose):

  • algorithm3 — GPU/YOLO container, processes video → anomaly images + result CSVs
  • mysql, socket-server, local-gui-backend — local site stack
  • watchtower — auto-pulls new images every 300s from the private registry
  • rclone-synology-sync — Alpine rclone image + sync.sh loop (this repo has exact copies)

The sync loop (sync-container/sync.sh — byte-identical to production)

  • Writes rclone.conf for an SFTP remote (synodrive) → Synology at a public hostname, non-standard port, single shared user (creds in office compose only).
  • Starts rclone rcd --rc-web-gui on :5572 (RC API + built-in web GUI).
  • Infinite loop: for each SYNC_n=direction:local:remote env var, run rclone copy --size-only (up = local→NAS, down = NAS→local), then sleep SYNC_INTERVAL seconds (seconds — the "# minutes" comment in prod compose is wrong; effective prod interval is 120s).
  • Logs to /logs/: sync.log (everything), stats.log (pretty per-round), errors.log (grep " ERROR "), completed.log (grep "Copied|Moved"), rc.log.

Production sync pairs (per site; ORG/SITENAME templated)

# Dir What
1 up anomaly images → Saudi_Video_Sync/SeekRight/Anomaly/<ORG>/<SITE>/TEST
2 down master sheets ← ClientSync/ALGORITHM_DEPLOYMENT/Master_Sheets/<ORG>/<SITE>
3 down yolov8 weights ← ClientSync/ALGORITHM_DEPLOYMENT/Models/yolov8
4 down NIGHT weights ← ClientSync/ALGORITHM_DEPLOYMENT/Models/NIGHT
5 up result CSVs → ClientSync/ALGORITHM_DEPLOYMENT/csv_files/<ORG>/<SITE>

Monitoring focus = pairs 1 & 5 (client-produced files that must reach the NAS).

Known production issues (verified, mention when relevant)

  1. RC_PASSWORD never set in prod compose → unauthenticated RC API on host network :5572 (full read/write to NAS + local disk from the LAN).
  2. Plaintext creds in compose + git token in generate_compose.sh → flagged for rotation.
  3. --size-only + copy (not sync): same-size edits never re-transfer; deletes never propagate. "Synced" is weaker than it sounds.
  4. The built-in Web GUI shows nothing useful — see below.

KEY FINDING: rclone's built-in UI exists but can't see the syncs

rclone rcd --rc-web-gui serves a React GUI on :5572 (login admin/$RC_PASSWORD). Verified locally: while 5 pairs were actively copying, POST /core/stats returned all zeros. The rclone copy commands are separate OS processes; the rcd daemon only reports jobs started through its own API. So the prod GUI is decorative. Options for the real project:

  • (A) Ship our own agent that tails logs + POSTs to a central collector (recommended; zero change to sync behavior), or
  • (B) Rewrite sync.sh to submit jobs via rclone rc sync/copy _async=true so rcd/GUI sees them — bigger change, still no fleet view, history, or alerts.
  • Adding --use-json-log to the rclone copy invocation would make log parsing trivial (structured JSON per line) — one-flag change worth proposing.

This repo = the production monitoring stack

docker-compose.yml   mongo + collector — the permanent stack (port 4400)
monitor/collector/   FastAPI + pymongo + rclone; serves API and the built UI
monitor/agent/       stdlib-only sidecar deployed on each client machine
frontend/            React 18 + Vite dashboard → dist/ bind-mounted into collector
sync-container/      byte-identical prod sync.sh + Dockerfile (reference + used
                     by deploy/remote-machine builds) — never modify
deploy/              install bundles: agent-only/ (machine already runs prod
                     sync) and remote-machine/ (needs sync + agent)
send-to-friend/      agent-deploy.zip — the bundle actually on JAGAN-TEST-01
.backups/            pre-v2 versions of main.py/App.jsx/styles.css/.env

The simulation lab was REMOVED 2026-08-10 (user decision: not needed for production). docker-compose.sim.yml, generator/, lan-forward.ps1 are in git history; client/+client2/+server-data/ were gitignored fake data and are gone for good. Consequence: no local test fleet — alert paths can now only be exercised against a real machine. The 2026-08-05/06 sections below describe that lab; keep them as history, don't try to run them.

Run / operate

cd /opt/R-clone/R-clone-setup
sudo docker-compose up -d --build          # mongo + collector (permanent stack)
sudo docker-compose logs -f collector
sudo docker-compose down

This host has docker-compose v1 (the hyphenated binary) only — the v2 docker compose plugin is NOT installed and fails with unknown shorthand flag: 'd'. Always write docker-compose.

Environment (Linux Mint host — MOVED off Windows/WSL 2026-08-07)

The project now lives at /opt/R-clone/R-clone-setup on a Linux Mint box (vm3@vm3-mint). Every /mnt/c/... path and WSL workaround in older notes is dead — including RCLONE_SFTP_SET_MODTIME=false, which only existed because /mnt/c rejects utimes(). Current realities:

  • Repo was root-owned; chown -R vm3:vm3 /opt/R-clone ran 2026-08-07. Without it, no edits, no npm run build, and git throws "dubious ownership".
  • docker needs sudovm3 is not in the docker group.
  • Ports 4000 and 8000 on this host are already taken by unrelated apps (a Node app in ~/seekright-video-calendar, and another FastAPI service). That is why the collector moved to 4400. Don't move it back.
  • Backend vs frontend rebuild rules differ: main.py is COPYd into the image → a change needs sudo docker-compose up -d --build collector. frontend/dist is bind-mounted → npm run build alone is enough, no restart.
  • Collector host LAN IP 192.168.1.201/22 (so the /22 spans .0.3.255). It is DHCP — it already changed once and broke every remote agent. Reserve it.

Verified: RC API gives every UI metric (2026-08-05 experiment)

monitor/watch_batch.py submitted 30×5MB files via POST /sync/copy _async=true (src ./client/batch_test, throttled with core/bwlimit) and polled live:

  • scheduled = core/stats.totalTransfers · done = .transfers
  • moving now = .transferring[] (name, %, speed, per-file eta; max 4 shown = --transfers)
  • queued = totalTransfers transfers len(transferring)
  • overall ETA/speed/bytes = .eta, .speed, .bytes/.totalBytes
  • per-file duration = core/transferred[] completed_at started_at (+ error field) All 30 arrived; per-file report printed. Gotchas: per-file eta can be None early; totalTransfers grows during discovery (not instantly 30); always pass {"group": "job/<jobid>"} to scope stats to one job. Constraint: RC only sees jobs submitted via RC. The prod loop's rclone copy processes are invisible to it → for live per-file UI data in prod, sync.sh v2 should submit pairs via rclone rc sync/copy _async=true + poll, instead of spawning rclone copy. Log parsing alone gives only aggregate 5s stats lines
  • completion events.

Monitoring stack v1 — BUILT and running (2026-08-06)

agent (per machine) → collector (FastAPI+Mongo) → React dashboard, all in compose:

  • monitor/agent/agent.py — stdlib-only sidecar; tails /logs/sync.log (ro mount), regex-parses rounds/pairs/files/errors/progress, POSTs /api/ingest every 5s (empty POST = heartbeat). Works against byte-identical prod sync.sh.
  • monitor/collector/main.py — FastAPI + pymongo (tz_aware=True — naive-vs-aware datetime bug otherwise). Collections: machines, pairs, events. Endpoints: POST /api/ingest, GET /api/overview (fleet + computed alerts), GET /api/machines/{id}/events?type=&limit=. Serves React build from /app/static. Alert rules: offline (no heartbeat >30s, critical), pair last_status=fail (serious), errors_1h>0 (warning).
  • frontend/ — React 18 + Vite. Fleet cards (status badge, stat tiles, pair rows, live progress bar), alerts bar, click card → events panel with tabs (Files/Errors/Rounds/Pair results). Polls every 4s. Dark theme, dataviz status tokens (#0ca30c/#fab219/#ec835a/#d03b3b), icon+label never color alone. Build: cd frontend && npm run build (dist/ is bind-mounted into the collector — rebuild frontend = just npm run build, no docker rebuild).
  • UI: http://localhost:4400 · do not stop/restart containers without asking first. (The RIYADH-01 second-machine demo died with the sim lab.)

v1.1 — remote machines + RMM surface (2026-08-06)

  • Agent auth: AGENT_TOKEN shared secret (in .env); collector rejects ingest without Authorization: Bearer <token> (verified 401). Empty token = auth off (never for remote use).
  • RMM endpoints: GET /api/health (liveness), GET /api/alerts (lightweight polling), CORS enabled (CORS_ORIGINS env). Integration points for the user's future RMM: poll alerts, or embed /api/overview data.
  • deploy/remote-machine/: self-contained bundle (compose + .env.example) for any external machine (e.g. a friend's system): rclone-sync (1 up pair from the algorithm's OUTPUT_DIR) + agent → user's collector. Remote machine needs a network path to the collector — recommend Tailscale; COLLECTOR_URL then is the tailnet IP. Bundle builds from ../../sync-container and ../../monitor/agent.
  • Transfer visibility today: completed files ✓ (file_synced events), aggregate live progress per pair ✓ (5s stats → progress bar), per-file in-flight + pending queue ✗ — needs RC-submitted jobs (sync.sh v2, see RC section).

v1.2 — server-side verification / audit feature (2026-08-06) — WORKING

The audit-team answer: "did the file actually reach the Synology?" proven from the collector, no one logs into anything.

  • Agent WATCH_DIRS env: mounts up-pair source dirs ro at the same container paths as sync (/sources/...), ships file inventory (name/size/mtime) with every ingest.
  • Collector: own rclone remote (start.sh writes config; rclone installed via apt in Dockerfile), background thread rclone lsjson --recursive on every up-pair remote path every VERIFY_INTERVAL (45s local; use 300s+ in prod). pending = local file not in server listing; missing = pending with mtime older than MISSING_GRACE_S (600) → serious alert with file names.
  • UI: 4th tile "pending → server" + amber N pending / red N missing! chips per pair. Verified live: pending counts rise as generator drops files, fall to 0 after each sync round.
  • Friend's machine (compose seen 2026-08-06): logs at /home/testing/JAGAN/Prerequisites/Sync_logs, SITENAME=TEST, standard 5 pairs. deploy/remote-machine bundle fits as-is; needs Tailscale (or LAN) to reach the collector + AGENT_TOKEN. Nothing needed from his algorithm code — agent reads only Sync_logs + source dirs.

v2 — upload audit + windowed history (2026-08-07 → 10) — CURRENT

Prompted by: "admins audit UPLOADS, not downloads; need pending/in-progress/ completed and processed-counts per hour/day/3d/week/month, ~1 month of data." Researched: rclone's native Prometheus metrics can't help (same single-rcd blindness as the GUI + unscrapeable short-lived syncs); MFT dashboards use the Completed/In-Progress/Queued/Failed taxonomy → adopted.

Collector additions (monitor/collector/main.py)

  • Retention: Mongo TTL index on events.received_at, RETENTION_DAYS env (30). _ensure_retention falls back to collMod when the window changes.
  • GET /api/stats — one aggregation pass over 30d computes per-machine AND fleet rollups for windows 1h/24h/3d/7d/30d: uploaded, errors, failures, rounds, bytes, avg_bps. Machines with zero activity still appear (that's the signal admins need). ~15s poll from the UI.
  • GET /api/timeseries?range=&machine_id= — epoch-aligned buckets (5min/1h/3h/6h/1d per range), zero-filled, uploads + errors only.
  • GET /api/machines/{id}/events — added direction=, q= (regex-escaped search over file/message/pair), skip/limit paging, returns total.
  • DELETE /api/machines/{id} (agent-token auth) — retire a machine; without it a decommissioned machine alerts "offline" forever. Used to remove the CONNTEST test entry 2026-08-10.
  • round_summary now splits files_up/files_down/bytes_up (download bytes are unknowable — files land outside WATCH_DIRS; never fake them).
  • _machine_view: up_pairs/down_pairs split, in_progress, failed_up, uploads_1h (direction-filtered). Alerts unchanged EXCEPT download-pair failures now say "download pair failing" (still alerted, not shown as cards).
  • Verification defaults to the real NAS: DEFAULT_REMOTE env (compose sets synoreal), VERIFY_ENABLED true if either SYNOLOGY_HOST or REAL_SYNOLOGY_HOST is set. Old behavior (default synodrive = the deleted fake NAS) silently skipped verification for any site not in SITE_REMOTES.

UI (v2.1 after a misstep — frontend/src/App.jsx)

First v2 cut replaced the machine cards with range-dependent tiles and dropped download pair rows + the last-round tile. User: "not understandable, it was meaningful before." Lesson recorded: the last-round-ago tile and the full pair list ARE the at-a-glance value; window counters reading 0 look like a broken dashboard. Current layout:

  • Cards (top, range-INDEPENDENT, restored): last round tile (amber when

    30min stale), uploaded (1h), queued→NAS (pending), errors (1h), live progress bar, ALL 5 pair rows (up first), missing/pending chips.

  • "Upload history" section (below, driven by the range selector): KPI row (uploaded / failed+errors / rounds / avg throughput+bytes), uploads-per-bucket bar chart with hover tooltip + "view as table", per-machine comparison table (the per-client speed audit). Chart empty state names the likely cause (stopped sync container) instead of rendering all-zero bars.
  • Palette validated (dataviz skill): uploads #3987e5 / errors #d03b3b on #1a1a19 pass all six checks; separate stacked plots, never dual-axis.

Ops facts (2026-08-09/10)

  • JAGAN-TEST-01's sync container was stopped 2026-08-07 during load-test cleanup and left stopped for 65h — the stalled-round alert caught it; the "empty graph" complaint was exactly this. Agent online ≠ sync running.
  • His Master_Sheets down pair fails exit 3: remote path .../Master_Sheets/TEST/TEST doesn't exist on the NAS (ORG=SITE=TEST substitution) — create the folder or fix his compose vars.
  • Load-test cleanup order that WORKS: stop sync container → delete local source files → rclone purge the NAS folder → start container. NAS-first gets re-uploaded within one round (learned the hard way).
  • .env trimmed to AGENT_TOKEN + REAL_SYNOLOGY_* + ORG/SITENAME. send-to-friend/agent-deploy.zip rebuilt (old zip still had :8000).

v2.1 — public exposure + one-line installer (2026-08-11)

  • Public URL: https://rclone.seekright.com → Synology reverse proxy → this VM :4400 (same pattern as the RMM hostnames). PUBLIC_URL in .env.
  • Viewer auth: HTTP Basic middleware (VIEWER_USER/VIEWER_PASS in .env) gates EVERYTHING except /api/ingest, /api/health, and requests carrying the agent bearer token. Unset VIEWER_PASS = gate off. Browser prompts once.
  • One-line installer (RMM deploy_agent.sh pattern, leak-lesson applied — download is authenticated, live AGENT_TOKEN injected at serve time, split-sentinel guard, exactly-once placeholder check):
    curl -fsSL -u admin "https://rclone.seekright.com/deploy_agent.sh" -o /tmp/deploy_agent.sh
    sudo bash /tmp/deploy_agent.sh <MACHINE_ID> [SITE]   # then rm the file (holds live token)
    
    Auto-discovers /logs + upload-source host paths from the machine's rclone-synology-sync container (docker inspect), fetches agent.py from GET /agent.py (single source of truth — no embedded copy to drift), writes /opt/rclone-agent/, replaces any previous rclone-agent container, starts. Re-run with no args = in-place upgrade keeping identity. MACHINE_ID is REQUIRED on fresh installs (no hostname fallback — RMM's stale-entry lesson).
  • Collector build context is now ./monitor (Dockerfile at collector/Dockerfile) so the image carries agent/agent.py to serve.

v2.2 — REAL fleet sync.sh has drifted from our copy (found 2026-08-11)

First real office machine (GCBOT-1 / IRBLAMBTHAM001, image from the private registry) revealed the production sync.sh is NOT our byte-copy anymore:

  • Up-pairs are change-detected: skipped unless file count or du -sb size changed; logged as SKIP: [↑ UP] <label> — unchanged (idle Ns). A forced sync runs after FORCE_SYNC_AFTER=1800s idle (" reason: …" line follows pair start). Down pairs unchanged.
  • Rounds every 180s (CHECK_INTERVAL), not 120.
  • Consequences: up-pairs show activity at most every 30min when idle; a pair can be absent from the dashboard until its first non-skipped run unless skips are parsed. Agent v2 parses SKIP → pair_skip event; collector records last_skip_at/last_idle_s and creates the pair as status "idle" (NEVER overwrites fail — state files update before rclone runs, so a failed up-pair is followed by skips until content changes). UI shows idle as ok.
  • VERIFY_TIMEOUT env (default 300s) — 8.5k-file SFTP listings exceed 120s.
  • MAX_INV_FILES now env-tunable (default 2000; GCBOT image tree is 8.5k).
  • Jagan's machine runs OUR old sync.sh (bundle-built) — no SKIP lines there.
  • TODO: replace sync-container/sync.sh with the real fleet version (get full file from a machine: docker exec rclone-synology-sync cat /sync.sh).

Roadmap for the actual deliverable

  1. Agent (per client machine, sidecar container): tail /logs/*.log (or JSON log), parse rounds/pairs/files/errors, POST events + 60s heartbeat to collector.
  2. Collector: FastAPI + MongoDB. Models: Machine, SyncPair, Round, FileEvent, Alert.
  3. Server-side verification (the feature logs can't provide): collector has its own rclone SFTP remote to the NAS; rclone lsjson per site path to confirm reported uploads actually arrived; alert if a file is still missing after a grace period. Missed-heartbeat = dead machine alert (most important signal).
  4. UI: React dashboard — fleet grid (last round, last heartbeat, error counts), per-machine pair detail, file history, live progress later via RC API.
  5. Alerts: dashboard + email first; Slack/Teams webhook later. Rules before channels.

Conventions for this repo

  • No real office credentials/hostnames in tracked files, with ONE deliberate exception: .env holds the real REAL_SYNOLOGY_* values and IS committed (see .gitignore header — private repo, creds needed on the VM after a pull). Everything else uses placeholders like <ORG>/<SITE>.
  • sync-container/sync.sh stays byte-identical to production — replica fidelity is the point. Fixes go in compose env vars or new sidecar services, and proposed prod changes get documented in this file instead.
  • RMM integration is PARKED (user decision 2026-08-10, R-clone fixes first). When resumed: this VM is also the production RMM server (26 field machines) — read /opt/rmm-backend/CLAUDE.md before touching anything there. Plan agreed: additive read-only proxy endpoint in rmm-backend + new rmm-ui page, developed against a scratch uvicorn (MONGODB_DB=rmm_test, spare port) and a scratch UI copy — NEVER by editing /opt/rmm-ui/src in place (vite hot-reloads straight into production).