# Vase Matcher A Flask web application that takes a photograph of a flower arrangement / vase and identifies which item in a fixed template catalogue it most closely resembles — judged independently by **eleven** different computer vision / deep learning signals, combined into one weighted verdict plus a set of independent diagnostic sections (color, shape, texture, flower count, vase identity) that never influence the score but explain *why* a match looks the way it does. This document covers the system from High-Level Design (what the pieces are and why they exist) down to Low-Level Design (exact algorithms, data contracts, file-by-file responsibilities). --- ## Table of Contents 1. [Overview](#1-overview) 2. [High-Level Design](#2-high-level-design) 3. [Feature Walkthrough](#3-feature-walkthrough) 4. [Low-Level Design](#4-low-level-design) 5. [API Reference](#5-api-reference) 6. [Configuration Reference](#6-configuration-reference) 7. [Environment & Setup](#7-environment--setup) 8. [Testing Tools](#8-testing-tools) 9. [Known Limitations & Design Tradeoffs](#9-known-limitations--design-tradeoffs) 10. [Glossary](#10-glossary) --- ## 1. Overview **Problem**: given a customer's photo of a vase/flower arrangement they received, determine which SKU in a small, fixed template catalogue it is, and explain the confidence of that determination in a way a human can sanity-check. **Approach**: no single CV method is reliable enough alone (classical keypoint matchers fail on repetitive floral texture; color alone ignores shape; deep matchers alone ignore the vase). So the app runs a stack of independent methods, combines four of them into a literal **weighted sum** for the primary verdict, aggregates the same four by **rank** (Borda count) as a cross-check, and surfaces several more signals (color, shape, texture, flower count, vase identity, an external AI vision check) purely as **explanatory diagnostics** that a human reviewer can use to understand *why* the weighted verdict landed where it did — without ever touching the score itself. ``` ┌─────────────────────────────┐ │ Upload a vase photo │ └──────────────┬───────────────┘ │ ┌──────────────▼───────────────┐ │ Which of N catalogue SKUs │ │ does this most resemble? │ └──────────────┬───────────────┘ │ ┌─────────────────────┼─────────────────────┐ │ │ │ ┌────────▼────────┐ ┌─────────▼─────────┐ ┌────────▼────────┐ │ SCORED (feeds │ │ DIAGNOSTIC ONLY │ │ ON-DEMAND │ │ the verdict) │ │ (never scored) │ │ (opt-in, slow) │ ├──────────────────┤ ├────────────────────┤ ├─────────────────┤ │ SIFT │ │ Shape matching │ │ Flower count │ │ ORB │ │ Texture matching │ │ (SAM3+YOLO-World)│ │ SuperPoint+Light- │ │ Color family grid │ │ Vase identity │ │ Glue ("SuperGlue")│ │ AI vision verify │ │ (DINOv2+CLIP) │ │ LoFTR │ │ │ │ │ │ Color space │ │ │ │ │ └──────────────────┘ └────────────────────┘ └─────────────────┘ ``` --- ## 2. High-Level Design ### 2.1 System Context ``` ┌──────────┐ HTTPS ┌─────────────────────────────┐ │ Browser │◄────────────────────►│ Flask app (app.py) │ │ (user) │ upload / results │ torch17_new (Python 3.8) │ └──────────┘ │ │ │ ┌─────────────────────────┐ │ │ │ pipeline/* (in-process) │ │ │ │ SIFT ORB SuperGlue LoFTR │ │ │ │ Color Shape Texture │ │ │ │ YOLO-World DINOv2 CLIP │ │ │ │ rembg (background removal)│ │ │ └─────────────────────────┘ │ └───────┬───────────────┬───────┘ │ │ subprocess │ │ HTTPS (optional) (SAM3 only) │ │ ▼ ▼ ┌────────────────────────┐ ┌──────────────────┐ │ sam3_worker.py │ │ External vision- │ │ sam2_env (Python 3.10) │ │ LLM endpoint │ │ transformers 5.5 + SAM3 │ │ (Cloudflare tunnel)│ │ facebook/sam3, 4-bit NF4 │ │ "AI verification" │ └────────────────────────┘ └──────────────────┘ │ ▼ ┌───────────────────┐ │ NVIDIA GPU (8GB) │ │ shared by both │ │ processes, │ │ never concurrently │ │ loaded (see §4.7) │ └───────────────────┘ ``` Everything runs on a **single 8GB consumer GPU** (GTX 1070, Pascal / compute capability 6.1). That one constraint shapes almost every architectural decision in this project: sequential (not concurrent) model execution, aggressive image downscaling, explicit model unload/reload cycles, 4-bit quantization for the newest model, and a cross-process split for the one dependency that couldn't share a Python environment with everything else. ### 2.2 Why Two Python Environments? `facebook/sam3` (via Hugging Face `transformers`) needs `transformers>=5.5.0`, which itself requires **Python ≥3.10**. The rest of the app (all the deep-learning matchers, rembg, etc.) lives in a `torch17_new` conda environment pinned to **Python 3.8** for compatibility with older CUDA/driver combinations already validated there. Upgrading that shared environment in place was judged too risky (could silently break DINOv2/CLIP/LightGlue/rembg, all already working). Instead: - SAM3 runs in a **separate conda env** (`sam2_env`, Python 3.10) that already had `transformers 5.5.0` and `bitsandbytes` installed. - The main Flask app talks to it via a **one-shot subprocess per request** (`sam3_worker.py`), passing a JSON request file and reading back a JSON response + PNG masks. See [§4.6](#46-sam3-cross-process-bridge) for the exact contract. - This keeps the two environments fully decoupled — SAM3 can be upgraded, reinstalled, or even moved to a different machine/GPU without touching the main app at all. ### 2.3 Core Design Principles | Principle | Why | |---|---| | **Sequential GPU usage, never concurrent** | An 8GB card can't hold multiple deep models' activations at once without fragmentation-driven CUDA OOM. Every heavy pipeline runs its methods one at a time, with `torch.cuda.empty_cache()` between them. | | **Aggressive, uniform downscaling** | Every image (upload *and* template) is capped to `MAX_IMAGE_DIM = 1600px` before touching any model — this single choke point is what keeps memory/time bounded regardless of how large the original photo was. | | **Opt-in for heavyweight features** | Flower counting and vase-identity comparison unload the core matching models first (to free GPU headroom) and are only triggered by explicit user action — they must never slow down the default "upload → match" path. | | **Scored vs. diagnostic separation** | Only 5 signals (SIFT, ORB, SuperGlue, LoFTR, Color) ever feed a number into the weighted verdict. Shape, texture, the color family grid, AI verification, flower count, and vase identity are **always** presented as separate, clearly-labeled sections that cannot move the score — this is a hard invariant enforced by keeping their computation entirely out of `_weighted_scores()`/`_overall_best()`. | | **Two independent "final answer" aggregates** | A literal weighted **sum** (`_weighted_scores`) and a rank-based **Borda count** (`_overall_best`) are computed from the same four raw method outputs but can legitimately disagree — the weighted sum is the primary displayed verdict; Borda is a sanity cross-check. | | **Soft-fail everything non-essential** | A failing color analysis, a stale AI-verification tunnel, a SAM3 subprocess crash — none of these ever take down the core match. Every optional section is wrapped in try/except and reported as its own `error` field. | | **Concept prompting over heuristic filtering** | SAM3 replaced an earlier SAM1-based approach that had to *guess* what counted as "a flower" from size/position/overlap heuristics. Prompting SAM3 directly with the word "flower" (or "vase") does that semantic work at the model level instead. | ### 2.4 Technology Stack | Layer | Technology | |---|---| | Backend web framework | Flask (development server, threaded) | | Classical CV | OpenCV (SIFT, ORB, FLANN matching, RANSAC homography, K-means, GLCM via scikit-image) | | Deep local features | SuperPoint + LightGlue (via the `lightglue` package) | | Deep dense matching | LoFTR (via `kornia.feature`) | | Background removal | `rembg` (BiRefNet-general-lite model), ONNX Runtime (CUDA + CPU fallback) | | Concept segmentation | SAM3 (`facebook/sam3`, Hugging Face `transformers`, 4-bit NF4 via `bitsandbytes`) | | Open-vocabulary detection | YOLO-World (`yolov8s-worldv2.pt`, via `ultralytics`) | | Embedding similarity | DINOv2 (`facebook/dinov2-base`) + CLIP (`openai/clip-vit-base-patch32`), via `transformers` | | External AI check | Third-party vision-LLM endpoint (Cloudflare tunnel), consumed over HTTPS | | Frontend | Vanilla JS (no framework), hand-rolled DOM building, CSS custom properties for theming | | Batch testing | Standalone Python scripts reusing the production pipeline (`testVaseMatcher.py`, `tester.py`) | --- ## 3. Feature Walkthrough Presented in the order they were built, which is also roughly the order of increasing sophistication: 1. **Core matching (SIFT / ORB / SuperGlue / LoFTR)** — four independent keypoint/dense matchers, each scored by RANSAC inlier count against every template, run strictly sequentially. 2. **Weighted final verdict + Borda overall-best** — two different ways of combining the four raw scores into one winning template. 3. **Color space section** — HS-histogram intersection, shown as its own section but also folded into the weighted verdict (`Color` is a 5th contributor there only). 4. **Background-removal size fix** — uploads over 15MB are downscaled + re-encoded rather than rejected outright. 5. **Color family grid** — dynamic K-means color-region discovery + side-by-side region-matching visualization, purely informational. 6. **Shape matching** — Hu-moment contour distance + canonical-silhouette IoU, its own section, never scored. 7. **Texture matching** — Local Binary Patterns + GLCM/Haralick features, its own section, never scored. 8. **AI verification** — sends the matched template photo + the user's original upload to an external vision-LLM endpoint for a free-text QC-style comparison, parsed into a structured Match/Confidence/ Discrepancies/Description card. 9. **Flower counting v1 (SAM1)** — Segment Anything (ViT-B) in automatic "segment everything" mode, filtered by size/position heuristics to guess which proposals were flowers. Superseded by step 11. 10. **YOLO-World cross-check + vase exclusion** — open-vocabulary detection used first to strip the vase/ribbon out of SAM1's flower count, and kept afterward as an independent second opinion shown side by side. 11. **Flower counting v2 (SAM3) + tuned recall** — SAM1 replaced with SAM3 concept-prompted segmentation (prompt: `"flower"`), which never proposes the vase/ribbon in the first place — no more heuristic filtering needed. (An intermediate step also tuned SAM1's own `points_stride`/`conf_thres`/`stability_score_thresh` for better recall before the SAM3 replacement landed; that tuning knowledge is preserved in this document for context even though the code path is gone.) 12. **Vase-identity comparison (DINOv2 + CLIP)** — crops the vase out of both photos using SAM3's precise mask (not just YOLO's bounding box), background-blacks-out everything else in the crop, and reports a Same/Uncertain/Different verdict from a weighted DINOv2+CLIP cosine similarity. 13. **Flower-count mismatch explanation** — a plain-English note ("your photo has 23 flowers vs. 9 in the template") surfaced wherever counts differ, explaining a likely contributor to a lower match score. 14. **`tester.py`** — a lean batch-accuracy script that also emits one combined report image per test photo (input segmentation, template segmentation, counts, vase comparison) using the real production pipeline. 15. **UI reorder** — the "Weighted final match" section moved to the very top of the results, so the final verdict is visible without scrolling. 16. **Auto-triggered "Flower check"** — a lightweight, CLIP-only (no DINOv2) flower-count + flower-similarity mini-panel shown beside the weighted card, fired automatically on every confident match (unlike the full flower-count feature, which stays behind its own button) — with a loading spinner while it computes. --- ## 4. Low-Level Design ### 4.1 Directory Structure ``` featureTransform/ ├── app.py Flask routes, logging setup, request glue ├── config.py All tunables; env-var capping; HF_TOKEN loading ├── sam3_worker.py Standalone SAM3 subprocess entrypoint (sam2_env) ├── tester.py Lean batch-accuracy + composite report images ├── testVaseMatcher.py Fuller batch-accuracy report (CSV, per-method) ├── .env HF_TOKEN=hf_... (gitignored, gated model access) │ ├── pipeline/ │ ├── engine.py Orchestrator: bootstrap, per-request pipeline, │ │ weighted/Borda aggregation, opt-in features │ ├── bg_removal.py rembg wrapper, disk caching, mask creation │ ├── classical.py SIFT / ORB extraction + FLANN matching │ ├── deep.py SuperPoint+LightGlue, LoFTR │ ├── color.py HS-histogram + dominant-color comparison │ ├── color_grid.py Dynamic K-means color-family grid │ ├── shape_match.py Hu moments + canonical-silhouette IoU │ ├── texture_match.py LBP + GLCM/Haralick │ ├── flower_count.py SAM3-backed per-instance flower counting │ ├── yolo_world.py Open-vocabulary flower/vase/ribbon detection │ ├── vase_compare.py DINOv2 + CLIP vase-identity comparison │ ├── sam3_client.py Subprocess bridge INTO sam3_worker.py │ ├── verify.py External AI-verification HTTP client │ └── utils.py Template listing, score annotation, compression │ ├── templates/ │ └── index.html Single-page UI (Jinja for the template gallery) │ ├── static/ │ ├── css/style.css │ └── js/main.js All client-side rendering logic │ ├── cache/ │ ├── templates_nobg/ Background-removed templates, cached by name │ └── uploads_nobg/ Background-removed uploads, cached by content hash │ ├── uploads// Per-request working files (originals, overlays, │ crops, family grids) -- swept after 6 hours ├── logs/app.log Rotating log, all pipeline.* + werkzeug output └── test_reports/ tester.py's composite report images ``` ### 4.2 Module Responsibilities #### `config.py` The single source of truth for every tunable constant in the system, plus two pieces of process-startup plumbing that **must** run before any heavy import: 1. Caps `OMP_NUM_THREADS`/`OPENBLAS_NUM_THREADS`/`MKL_NUM_THREADS`/ `NUMEXPR_NUM_THREADS`/`ORT_NUM_THREADS` to `min(4, cpu_count)` — done via `os.environ.setdefault` at import time, before `cv2`/`onnxruntime`/ `torch` are imported anywhere in the process. Left uncapped, each library grabs one thread per core, which starves everything else on a shared machine. 2. Sets `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` — the caching allocator's own recommended mitigation for fragmentation-driven CUDA OOM on a small card. It also reads `HF_TOKEN` once (env var first, then a tiny hand-rolled `.env` parser as fallback) so the SAM3 subprocess can be handed it explicitly without needing a `python-dotenv` dependency. #### `pipeline/engine.py` The orchestrator. Owns: - `bootstrap()` — precomputes every template's background-removed image and every method's template-side features/embeddings **once**, at startup, so a live request only ever has to process the single uploaded image. - `process_upload()` / `_process_upload_locked()` — the `/api/match` code path: background removal → sequential SIFT/ORB/SuperGlue/LoFTR → best-image annotation → color/shape/texture analysis → weighted + Borda aggregation → color family grid → response assembly. - `_weighted_scores()` / `_overall_best()` — the two aggregation algorithms (see [§4.4](#44-aggregation-layers)). - `count_flowers()` — the full opt-in flower-count + vase-comparison feature (SAM3 + YOLO-World + DINOv2/CLIP, renders visuals, returns everything for the "Count flowers" button). - `flower_summary()` — the lightweight, auto-triggered companion (SAM3 + CLIP only, no YOLO-World, no DINOv2, no rendered images) shown beside the weighted card on every confident match. - A single `threading.Lock` (`_pipeline_lock`) serializes every request's pipeline end-to-end, so at most one heavy CPU/GPU operation is ever in flight process-wide. #### `pipeline/bg_removal.py` Wraps `rembg` (BiRefNet-general-lite). Lazily creates an ONNX Runtime session (`CUDAExecutionProvider` first, falls back to CPU on failure — this fallback is "sticky" for the process's lifetime, a known limitation). Disk-caches results: uploads by content hash, templates by filename. `resize_max_dim()` is the single choke point enforcing `MAX_IMAGE_DIM` everywhere. `unload_session()` drops the session so `count_flowers()`/ `flower_summary()` can free its GPU memory before the heavier optional models load. #### `pipeline/classical.py` SIFT and ORB feature extraction + FLANN-based matching. Template descriptors are precomputed once and cached in-memory (`_template_features`). See [§4.3](#43-scoring-algorithms) for the exact scoring formula. #### `pipeline/deep.py` SuperPoint+LightGlue (labeled "SuperGlue" in the UI, matching an existing naming convention) and LoFTR. `unload_models()` drops all three model objects (SuperPoint, LightGlue, LoFTR) so the opt-in flower/vase features can reclaim their GPU memory; they lazily reload on the next `/api/match` call, identical to a fresh process start. #### `pipeline/color.py` Independent, non-scored-looking-but-actually-scored (folded into the weighted verdict via `Color`) comparison: Hue+Saturation 2D histogram (Value/brightness deliberately excluded for lighting invariance), compared via histogram intersection. Also computes dominant colors (K-means in a bounded sample) and per-color Lab-distance similarity pairs for the visual palette display. #### `pipeline/color_grid.py` Dynamic K-means color-region discovery in LAB space (lightness down-weighted so shadows of the same hue cluster together), rendered as a tile grid, region-matched by nearest average color, with an area-match-percentage per region and an overall area-weighted headline number. Purely visual — tied to whichever template the *weighted* verdict already picked. #### `pipeline/shape_match.py` Two complementary shape signals, averaged into one `match_pct`: - **Hu moments** (`cv2.matchShapes`, `CONTOURS_MATCH_I1`) — translation/ rotation/scale-invariant contour distance, converted to a similarity %. - **Canonical-silhouette IoU** — both masks cropped to their own bounding box, resized/centered into a fixed square canvas, then measured as direct pixel overlap. This also produces the side-by-side + overlay visualization images. #### `pipeline/texture_match.py` Two complementary texture signals, averaged into one `match_pct`: - **Local Binary Patterns** (uniform method), compared as a histogram intersection — catches fine, repetitive patterns (fabric weave, petal grain). - **GLCM/Haralick features** (contrast, homogeneity, energy, correlation), compared as a normalized relative difference, averaged across properties and across 4 angles/2 distances — catches coarser smooth-vs-rough material differences. #### `pipeline/flower_count.py` SAM3-backed per-instance flower counting. `count_flowers(bgr, mask, workdir, instances_raw=...)` accepts already-fetched SAM3 instances (so `engine.py` can batch multiple images/prompts into one subprocess call and hand each result set to this function separately), applies a light foreground-overlap sanity filter, then color-clusters the survivors (LAB K-means, capped at `SAM_MAX_KIND_CLUSTERS`) as a rough proxy for distinct flower "kinds". `union_mask()` OR-combines every instance mask into one — used to crop "just the flower material" for CLIP similarity. `render_instances()` draws the numbered, color-coded overlay. #### `pipeline/yolo_world.py` Open-vocabulary detection (`yolov8s-worldv2.pt`) prompted with `["flower", "vase", "ribbon", "bow"]`. Kept as an **independent second opinion** shown side by side with SAM3's per-instance count — it draws one box per contiguous flower region rather than per bloom, so it's a coarser, corroborating signal, not a replacement. #### `pipeline/vase_compare.py` - `crop_mask(bgr, mask_bool)` — crops the bounding box of any mask (padded by `VASE_CROP_PAD_FRAC`), blacking out every pixel the mask doesn't cover. Generic despite the module name — also used for the flower-only crop in `flower_summary()`. - `compare_vases()` — DINOv2 + CLIP cosine similarity, combined as a weighted average (`VASE_DINO_WEIGHT=0.6`, `VASE_CLIP_WEIGHT=0.4`), bucketed into `same`/`uncertain`/`different`. - `clip_similarity_pct()` — CLIP-only variant (never touches DINOv2), used by the lightweight auto-triggered flower check. - `unload_models()` — drops both model objects; a no-op for whichever one was never loaded (e.g. DINOv2 stays unloaded entirely on the CLIP-only path). #### `pipeline/sam3_client.py` The subprocess bridge (runs *in* the main app's Python 3.8 process). Saves each named input image to a temp PNG, writes a JSON request file describing every `(image, prompt, threshold)` job, invokes `sam3_worker.py` under `sam2_env`'s Python interpreter with `HF_TOKEN` passed through the subprocess environment, waits (bounded by `SAM3_TIMEOUT_SECONDS`), and parses the JSON response + PNG masks back into `{(image_key, prompt): [{"mask": bool ndarray, "score": float, "box": [...]}]}`. Always cleans up its temp directory, even on failure. #### `sam3_worker.py` Standalone script, **runs under `sam2_env`'s Python 3.10**, has no dependency on this app's `config`/`pipeline` modules. Loads `Sam3Model` + `Sam3Processor` 4-bit quantized (NF4 via `BitsAndBytesConfig`), processes every job in the request against the already-loaded model (one model load serves every job in a request), writes one PNG per detected instance plus a `response.json` manifest. #### `pipeline/verify.py` Thin HTTP client for the external AI-verification endpoint. Posts both images as multipart form data, parses the endpoint's free-text `result` field with a tolerant regex (`DISCREP\w*` rather than a literal spelling, since the LLM behind it doesn't always spell "DISCREPANCIES" consistently) into `{match, confidence, discrepancies, description}`. #### `pipeline/utils.py` Template file listing, score-annotation overlay text, PNG encoding, and `compress_image_bytes()` — downscales + re-encodes as JPEG only if an upload exceeds `COMPRESS_ABOVE_BYTES`, reusing `MAX_IMAGE_DIM` since the pipeline downsamples to that anyway. ### 4.3 Scoring Algorithms #### SIFT / ORB (`pipeline/classical.py`) ``` 1. Extract keypoints + descriptors (masked to foreground) 2. FLANN k=2 nearest-neighbor match against template descriptors 3. Lowe's ratio test: keep match if best.distance < 0.75 * secondBest.distance 4. RANSAC homography (5.0px threshold) across surviving matches 5. score = RANSAC inlier count (raw integer) confidence = inliers / good_matches * 100 (%) ``` #### SuperGlue = SuperPoint + LightGlue (`pipeline/deep.py`) Same shape as above, but SuperPoint (neural keypoints/descriptors) replaces SIFT/ORB's detector and LightGlue (learned matcher) replaces FLANN+ratio test. Matches filtered to inside the foreground mask before the same RANSAC-inlier-count scoring. #### LoFTR (`pipeline/deep.py`) Dense matcher — no keypoint detection step. Directly regresses pixel correspondences, keeps only ones above `LOFTR_CONFIDENCE_THRESHOLD` (0.5), filters to inside the mask, then the same RANSAC-inlier-count scoring. #### Color space (`pipeline/color.py`) ``` hist = HS 2D histogram (Hue×Saturation only, Value excluded), normalized to sum=1 match_pct = Σ min(hist_input[i], hist_template[i]) × 100 (histogram intersection) ``` #### Shape matching (`pipeline/shape_match.py`) ``` hu_similarity = max(0, 100 × (1 - matchShapes(contour_a, contour_b) / 1.5)) iou = |mask_a ∩ mask_b| / |mask_a ∪ mask_b| × 100 (canonical-aligned) match_pct = (hu_similarity + iou) / 2 ``` #### Texture matching (`pipeline/texture_match.py`) ``` lbp_similarity = Σ min(hist_a[i], hist_b[i]) × 100 (LBP histogram intersection) glcm_similarity = mean over 4 properties of max(0, 1 - |a-b|/max(|a|,|b|)) × 100 match_pct = (lbp_similarity + glcm_similarity) / 2 ``` #### Vase identity (`pipeline/vase_compare.py`) ``` cosine_pct(a, b) = max(0, min(1, a·b)) × 100 (a, b are L2-normalized embeddings) dino_pct = cosine_pct(DINOv2(crop_a), DINOv2(crop_b)) clip_pct = cosine_pct(CLIP(crop_a), CLIP(crop_b)) combined_pct = 0.6 × dino_pct + 0.4 × clip_pct verdict = "same" if combined_pct >= 75 "uncertain" if combined_pct >= 60 "different" otherwise ``` #### Flower CLIP similarity (`flower_summary`, CLIP-only variant) ``` flower_clip_pct = cosine_pct(CLIP(union_of_flower_crops_a), CLIP(union_of_flower_crops_b)) ``` ### 4.4 Aggregation Layers Two genuinely different "who won" computations exist side by side, and are allowed to disagree: **Weighted final match** (`engine._weighted_scores`) — a literal weighted **sum** of raw scores: ``` weighted_score(template) = Σ_method WEIGHT[method] × raw_score(method, template) WEIGHT = { LoFTR: 0.75, SuperGlue: 0.15, SIFT: 0.10, ORB: 0.10, Color: 0.7 } ``` Because this sums *raw* scores (inlier counts, which run much higher for LoFTR/SuperGlue than SIFT/ORB, plus Color's 0-100 percentage), a method with naturally larger magnitudes pulls harder on the total even at a similar or lower weight. That's the deliberate behavior of the formula as specified, not a bug — see the comment block in `config.py` above `METHOD_WEIGHTS`. **Overall best** (`engine._overall_best`, Borda count) — rank-based, and excludes `Color` entirely (only the original four methods participate): ``` for each method: rank all templates by that method's own raw score template at rank i (0-indexed) earns (n - i) points overall_best = template with the highest total points across all methods ``` Rank-based aggregation means no single method's raw-score magnitude can dominate the way it can in the weighted sum — this is why the two verdicts are kept, and shown, separately rather than collapsed into one number. ### 4.5 Request Lifecycles #### `POST /api/match` ``` Browser Flask (app.py) engine.py GPU │ upload image │ │ │ ├─────────────────────────► │ │ │ │ compress if >15MB │ │ │ ├──────────────────────────► │ │ │ │ │ bg_removal (rembg) │ │ │ ├────────────────────►│ │ │ │◄────────────────────┤ │ │ │ SIFT (sequential) │ │ │ ├────────────────────►│ │ │ │ ORB │ │ │ ├────────────────────►│ │ │ │ SuperGlue │ │ │ ├────────────────────►│ │ │ │ LoFTR │ │ │ ├────────────────────►│ │ │ │ Color / Shape / │ │ │ │ Texture (CPU) │ │ │ │ weighted + Borda │ │ │ │ color family grid │ │ │◄───────────────────────────┤ │ │◄───────────────────────── JSON result │ │ ``` SuperPoint/LightGlue/LoFTR **stay loaded** in memory after this request (for fast repeated matching) — they are only unloaded when an opt-in flower/vase feature explicitly needs the GPU headroom. #### `POST /api/count_flowers` (opt-in, "Count flowers" button) ``` Browser app.py engine.count_flowers() sam3_client → sam3_worker.py (sam2_env) YOLO-World DINOv2+CLIP │ click │ │ │ │ │ ├───────────────► │ │ │ │ │ ├────────────────────► │ │ │ │ │ │ unload SuperPoint/LightGlue/LoFTR + rembg session │ │ │ │ │ yolo_world.detect() ────────────────────────────────────────────►│ │ │ │ │ batch: input+flower, input+vase, [template+vase, template+flower] │ │ │ │ ├─────────────────────────────► │ │ │ │ │ (one SAM3 model load serves every job) │ │ │ │ │◄─────────────────────────────┤ │ │ │ │ │ cluster + render SAM instances (input & template) │ │ │ │ │ render YOLO-World boxes │ │ │ │ │ vase_compare (DINOv2+CLIP) ──────────────────────────────────────────────────►│ │ │ │◄─────────────────────────────────────────────────────────────────────────────┤ │ │ │ unload YOLO-World, DINOv2, CLIP │ │ │ │◄───────────────────┤ │ │ │◄─────────────── JSON + image URLs │ │ │ ``` #### `POST /api/flower_summary` (auto-triggered, beside the weighted card) ``` Browser app.py engine.flower_summary() sam3_client/worker CLIP only │ (fires automatically │ │ │ │ right after a confident match) │ │ │ ├────────────────────────────────────► │ │ │ │ unload matching models │ │ │ │ batch: input+flower, │ │ │ │ template+flower │ │ │ ├─────────────────────────────►│ │ │ │◄─────────────────────────────┤ │ │ │ count comparison (no render)│ │ │ │ union masks, crop, CLIP-only similarity ────────────►│ │ │◄──────────────────────────────────────────────────────┤ │ │ unload CLIP (DINOv2 never loaded) │ │◄──────────────────────────────────── │ ``` Deliberately skips YOLO-World, DINOv2, and any rendered overlay images — it exists purely to be fast enough to run unconditionally on every match. #### `POST /api/verify` (AI verification, auto-fires after a confident match) ``` Browser app.py verify.py External vision-LLM endpoint │ │ │ │ ├──────────────► │ │ │ ├────────────────► POST multipart (image1, image2) │ │ ├─────────────────────────►│ │ │ │◄─────────────────────────┤ │ │ │ parse MATCH/CONFIDENCE/ │ │ │ │ DISCREPANCIES/description│ │ │◄──────────────── │ │◄────────────── JSON (or 502 if endpoint unreachable) │ ``` Failures here (stale Cloudflare tunnel, timeout) are surfaced as a soft error in the UI, never a hard failure of the page. ### 4.6 SAM3 Cross-Process Bridge **Request** (JSON file, written by `sam3_client.py`, read by `sam3_worker.py`): ```json { "images": { "input": "/path/to/uploads//sam3_xxxx/input.png", "template": "/path/to/uploads//sam3_xxxx/template.png" }, "jobs": [ {"image": "input", "prompt": "flower", "threshold": 0.5}, {"image": "input", "prompt": "vase", "threshold": 0.3}, {"image": "template", "prompt": "vase", "threshold": 0.3} ], "output_dir": "/path/to/uploads//sam3_xxxx/out" } ``` **Response** (`response.json`, written by the worker): ```json { "results": [ {"image": "input", "prompt": "flower", "instances": [ {"mask_file": "input_flower_0_0.png", "score": 0.7486, "box": [x1, y1, x2, y2]}, {"mask_file": "input_flower_0_1.png", "score": 0.7006, "box": [...]} ]}, {"image": "input", "prompt": "vase", "instances": [ {"mask_file": "input_vase_1_0.png", "score": 0.9391, "box": [...]} ]} ], "error": null } ``` Each `mask_file` is a standalone grayscale PNG (0/255) at the original image's resolution — `sam3_client.py` reads it back with `cv2.imread(..., IMREAD_GRAYSCALE) > 127` to reconstruct a boolean mask. On any failure, `error` is a string (worker traceback included) and the process exits non-zero; `sam3_client.py` raises `Sam3Error` either way, which every caller treats as a soft failure. **Why 4-bit (NF4) quantization**: the unquantized checkpoint is ~3.2GB (`model.safetensors`). Loading it naively (8-bit, `device_map="cuda:0"`) hit a CUDA OOM during Transformers' own memory pre-flight ("caching allocator warmup") step on this 8GB card, because that warmup reserves capacity based on the *original* dtype size before quantization actually reduces it. Switching to `load_in_4bit=True` + `bnb_4bit_quant_type="nf4"` + `low_cpu_mem_usage=True` fixed this: ~700MB resident, ~1.9GB peak during inference, ~6s to load, ~2s per prompt. **Why `sam.predictor(...)` isn't called directly for automatic-mode tuning** (historical note, from the SAM1 era): Ultralytics' high-level `model.predict()`/`model()` call validates every kwarg against a fixed CLI-style config schema, which rejects automatic-mode-only parameters like `points_stride`/`conf_thres` with a `SyntaxError`. The fix at the time was calling `model.predictor(...)` directly (after a one-time warmup call to force predictor construction), bypassing that validation entirely since `BasePredictor.__call__` forwards arbitrary kwargs straight through. This code path no longer exists (SAM1 was fully replaced by SAM3), but the technique is documented here in case a similar situation recurs with another Ultralytics-wrapped model (e.g. YOLO-World). ### 4.7 GPU Memory Management Strategy ``` ┌───────────────────────────────────────────┐ │ Idle / after /api/match │ │ SuperPoint + LightGlue + LoFTR RESIDENT │ │ (~200-900MB, kept warm for fast matching) │ └──────────────────┬────────────────────────┘ │ user clicks "Count flowers" │ or a confident match fires │ the auto flower-check ▼ ┌───────────────────────────────────────────┐ │ deep.unload_models() │ │ bg_removal.unload_session() │ │ torch.cuda.empty_cache() │ │ → GPU now near-baseline (~200-900MB) │ └──────────────────┬────────────────────────┘ ▼ ┌───────────────────────────────────────────┐ │ SAM3 (separate PROCESS, sam2_env) │ │ ~700MB resident / ~1.9GB peak │ │ + optionally YOLO-World (~200MB) │ │ + optionally DINOv2 (~350MB) + CLIP (~600MB) │ └──────────────────┬────────────────────────┘ │ request finishes ▼ ┌───────────────────────────────────────────┐ │ yolo_world.unload_model() │ │ vase_compare.unload_models() │ │ (SAM3's own process has already exited — │ │ its GPU memory is freed by the OS/driver, │ │ nothing to unload on the main-process side) │ │ → back to near-baseline │ └───────────────────────────────────────────┘ ``` Every model getter (`deep.get_superpoint()`, `flower_count`'s SAM3 client, `vase_compare._get_dino()`/`_get_clip()`, `yolo_world.get_model()`) is a lazy singleton — `None` until first use, reset to `None` by the matching `unload_*()` function. This means "unload" is always safe to call even if the model was never loaded (checked via `is not None`), and every model transparently reloads itself on the next call that needs it, with no special-casing required anywhere else in the codebase. ### 4.8 Frontend Structure (`static/js/main.js`) Single IIFE, no framework, no build step. Key structure: - **Upload flow**: drag/drop or file picker → `fetch POST /api/match` → `renderResults(data)`. - **`renderResults(data)`** populates every section in DOM order (weighted card first, then overall-best, query images, method grid, flower count, color/shape/texture/family-grid sections), then fires two **auto-triggered, non-blocking** follow-ups: `runVerification()` (AI check) and `runFlowerSummary()` (the lightweight flower-check panel) — neither blocks the initial render, both show their own loading state and fail independently. - **`renderWeighted(data)`** builds the terracotta-gradient weighted card (input/template thumbnails, breakdown table, Strong/Moderate/Weak confidence pill derived purely from the *margin* between 1st and 2nd place — display-only, cannot feed back into the algorithm). - **Opt-in flower count**: `countFlowersBtn` click handler → `fetch POST /api/count_flowers` → `renderFlowerCount(data)`, which in turn calls `renderFlowerCountMismatch()` and `renderVaseComparison()`. - Every async section follows the same pattern: hide body / show spinner → fetch → on success show body, on failure show a scoped error message — never a page-level failure. --- ## 5. API Reference | Method | Path | Body | Returns | |---|---|---|---| | GET | `/` | — | Rendered HTML page (template gallery + upload form) | | GET | `/template_image/` | — | Raw template image bytes | | GET | `/uploads//` | — | Any file written into that request's working directory | | POST | `/api/match` | multipart `image` file | Full match result: per-method scores, weighted verdict, Borda overall-best, color/shape/texture/family-grid analysis | | POST | `/api/count_flowers` | JSON `{request_id, template}` | SAM3 + YOLO-World flower counts (with visuals), vase comparison, flower-count-mismatch note | | POST | `/api/flower_summary` | JSON `{request_id, template}` | Lightweight flower-count comparison + CLIP-only flower similarity (no images) | | POST | `/api/verify` | JSON `{request_id, template}` | External AI vision-LLM verdict: match/confidence/discrepancies/description | All POST endpoints return `{"error": "..."}` with a non-200 status on failure; the frontend treats every one of these as scoped/local, never page-fatal. --- ## 6. Configuration Reference All of the following live in `config.py`, grouped by feature: | Setting | Default | Purpose | |---|---|---| | `MAX_IMAGE_DIM` | 1600 | Every image downscaled to this before any model touches it | | `METHOD_WEIGHTS` | `{LoFTR: .75, SuperGlue: .15, SIFT: .10, ORB: .10, Color: .7}` | Weighted-sum verdict weights | | `SCORE_THRESHOLD` | per-method | "Confident match" badge cutoff (doesn't affect ranking) | | `MAX_CONTENT_LENGTH_BYTES` | 50MB | Hard reject ceiling | | `COMPRESS_ABOVE_BYTES` | 15MB | Soft threshold — downscale+recompress instead of reject | | `SHAPE_HU_DISTANCE_SCALE` | 1.5 | Divisor turning Hu-moment distance into a 0-100% similarity | | `TEXTURE_LBP_RADIUS` / `TEXTURE_GLCM_LEVELS` | 2 / 32 | Texture descriptor parameters | | `FAMILY_GRID_K` | 5 | Number of color families discovered per image | | `SAM3_FLOWER_THRESHOLD` / `SAM3_VASE_THRESHOLD` | 0.5 / 0.3 | Confidence cutoff per SAM3 concept prompt | | `SAM3_TIMEOUT_SECONDS` | 180 | Subprocess timeout for the SAM3 worker | | `VASE_DINO_WEIGHT` / `VASE_CLIP_WEIGHT` | 0.6 / 0.4 | Vase-comparison combined-score weights | | `VASE_SAME_THRESHOLD` / `VASE_UNCERTAIN_THRESHOLD` | 75 / 60 | Vase verdict cutoffs (untuned — no ground-truth calibration set exists yet) | | `YOLO_WORLD_CLASSES` | `["flower", "vase", "ribbon", "bow"]` | Open-vocabulary prompt classes | | `VERIFY_ENDPOINT_URL` | Cloudflare tunnel URL | Can go stale if the tunnel restarts — treated as a soft failure | --- ## 7. Environment & Setup ### Two conda environments | Env | Python | Purpose | Key packages | |---|---|---|---| | `torch17_new` | 3.8 | Main Flask app | `torch`, `opencv-python`, `kornia`, `lightglue`, `rembg`, `ultralytics`, `transformers==4.46.3` | | `sam2_env` | 3.10 | SAM3 subprocess only | `torch`, `transformers>=5.5.0`, `bitsandbytes`, `accelerate` | ### Model checkpoints / weights | Model | Source | Notes | |---|---|---| | SIFT / ORB | Built into OpenCV | No download needed | | SuperPoint + LightGlue | `lightglue` package | Auto-downloaded on first use | | LoFTR | `kornia.feature` | Auto-downloaded on first use | | rembg (BiRefNet-general-lite) | `rembg` package | Auto-downloaded on first use | | YOLO-World | `yolov8s-worldv2.pt` | Downloaded via `ultralytics`, cached locally | | DINOv2-base | `facebook/dinov2-base` | Hugging Face, public, no token needed | | CLIP ViT-B/32 | `openai/clip-vit-base-patch32` | Hugging Face, public, no token needed | | **SAM3** | `facebook/sam3` | **Gated** on Hugging Face — requires requesting access on the model page and an approved access token | ### `.env` file ``` HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx ``` Required only for the SAM3-backed features (flower counting, vase comparison, the auto-triggered flower check). Without it, those specific features fail softly with a clear "SAM3 unavailable" message — core matching is entirely unaffected. ### Running the app ```bash cd /media/suman/Backup_of_extra_/Sasi/featureTransform conda activate torch17_new python app.py # serves on http://0.0.0.0:5053 ``` --- ## 8. Testing Tools Two batch-test scripts exist, covering different needs: | | `testVaseMatcher.py` | `tester.py` | |---|---|---| | Output | Detailed CSV + per-method accuracy table | Console pass/fail + one composite report image per photo | | Coverage | All 4 methods + weighted + Borda + color pick, individually tallied | Weighted verdict only | | Flower/vase check | Numbers only (family-grid area-match %), no images | Full: SAM3 flower segmentation (input vs. template), counts, vase comparison — all rendered into one PNG | | Speed | Faster (no SAM3/CLIP round trip) | Slower (~15-25s/image extra for the flower/vase step); `--skip-flowers` for a fast pass | | Use case | Regression-testing every individual method's accuracy | Visual QA — "does this match make sense at a glance" | Both expect the same input layout: ``` / SKU_1/ photo1.jpg photo2.jpg ... SKU_2/ ... ``` where each subfolder name must exactly match an existing template name. **Common mistake**: pointing either script directly at a `SKU_X/` folder instead of its parent — the script silently finds 0 images, since it expects to `listdir()` a folder of *subfolders*, not a folder of images. ```bash python tester.py # ./flowers, full report images python tester.py /path/to/flowers --limit 3 python tester.py --skip-flowers # fast accuracy-only pass python testVaseMatcher.py /path/to/flowers --csv report.csv ``` --- ## 9. Known Limitations & Design Tradeoffs - **rembg's CUDA fallback is sticky for the process's lifetime.** A transient CUDA OOM during session creation permanently falls back that process to CPU-only background removal (much slower) until restarted. No auto-retry/self-healing exists yet. - **Vase-identity and flower-count-mismatch thresholds are untuned.** There is no ground-truth "same/different vase" or "acceptable flower count delta" labeled dataset — the cutoffs in `config.py` are reasonable defaults, not calibrated ones. - **The weighted verdict's literal-sum formula means raw-score magnitude matters, not just weight.** LoFTR/SuperGlue's inlier counts run far higher than SIFT/ORB's, so they can dominate the total even when their configured weight isn't proportionally larger. This is documented as intentional (matches the formula as originally specified), not a bug. - **YOLO-World's own "flower" count is coarse** — one box per contiguous flower region, not per bloom — so it will never numerically match SAM3's per-instance count; it's shown as an independent corroborating signal, not something expected to agree. - **The AI-verification endpoint is a Cloudflare tunnel URL** that can go stale if restarted on the other end; failures are always soft (shown as "AI verification unavailable"), never block the page. - **SAM3 requires a gated, manually-approved Hugging Face token.** The app cannot obtain this automatically — a human must request access at `huggingface.co/facebook/sam3` and provide a token with access granted. - **Everything assumes a single physical GPU with ~8GB VRAM** and Pascal or later compute capability. The sequential-execution and unload/reload discipline exists specifically because of that constraint — a larger card would allow (but doesn't require) more concurrency. --- ## 10. Glossary | Term | Meaning | |---|---| | **Inlier count** | Number of matched keypoints that agree with one consistent RANSAC-fit geometric transform between two images — the raw score for SIFT/ORB/SuperGlue/LoFTR | | **Weighted verdict** | The primary displayed "final answer" — literal weighted sum of 5 methods' raw scores | | **Borda count / Overall best** | A secondary, rank-based cross-check aggregate of 4 methods (excludes Color) | | **Concept segmentation** | SAM3's mode of operation: prompted with a plain-English word ("flower"), returns instance masks for exactly that concept, rather than segmenting everything indiscriminately | | **NF4** | "NormalFloat4" — the 4-bit quantization scheme used (via `bitsandbytes`) to fit SAM3 in this GPU's memory budget | | **Flower check** | The lightweight, auto-triggered (not opt-in), CLIP-only flower-count + flower-similarity panel shown beside the weighted card | | **Count flowers** | The full, opt-in, button-triggered feature: SAM3 + YOLO-World flower counts, rendered overlays, and DINOv2+CLIP vase comparison |