Files
Vase-Matcher/README.md
2026-08-04 17:03:26 +05:30

911 lines
54 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Vase Matcher
A Flask web application that takes a photograph of a flower arrangement /
vase and identifies which item in a fixed template catalogue it most
closely resembles — judged independently by **eleven** different computer
vision / deep learning signals, combined into one weighted verdict plus a
set of independent diagnostic sections (color, shape, texture, flower
count, vase identity) that never influence the score but explain *why* a
match looks the way it does.
This document covers the system from High-Level Design (what the pieces
are and why they exist) down to Low-Level Design (exact algorithms, data
contracts, file-by-file responsibilities).
---
## Table of Contents
1. [Overview](#1-overview)
2. [High-Level Design](#2-high-level-design)
3. [Feature Walkthrough](#3-feature-walkthrough)
4. [Low-Level Design](#4-low-level-design)
5. [API Reference](#5-api-reference)
6. [Configuration Reference](#6-configuration-reference)
7. [Environment & Setup](#7-environment--setup)
8. [Testing Tools](#8-testing-tools)
9. [Known Limitations & Design Tradeoffs](#9-known-limitations--design-tradeoffs)
10. [Glossary](#10-glossary)
---
## 1. Overview
**Problem**: given a customer's photo of a vase/flower arrangement they
received, determine which SKU in a small, fixed template catalogue it is,
and explain the confidence of that determination in a way a human can
sanity-check.
**Approach**: no single CV method is reliable enough alone (classical
keypoint matchers fail on repetitive floral texture; color alone ignores
shape; deep matchers alone ignore the vase). So the app runs a stack of
independent methods, combines four of them into a literal **weighted sum**
for the primary verdict, aggregates the same four by **rank** (Borda count)
as a cross-check, and surfaces several more signals (color, shape, texture,
flower count, vase identity, an external AI vision check) purely as
**explanatory diagnostics** that a human reviewer can use to understand
*why* the weighted verdict landed where it did — without ever touching the
score itself.
```
┌─────────────────────────────┐
│ Upload a vase photo │
└──────────────┬───────────────┘
┌──────────────▼───────────────┐
│ Which of N catalogue SKUs │
│ does this most resemble? │
└──────────────┬───────────────┘
┌─────────────────────┼─────────────────────┐
│ │ │
┌────────▼────────┐ ┌─────────▼─────────┐ ┌────────▼────────┐
│ SCORED (feeds │ │ DIAGNOSTIC ONLY │ │ ON-DEMAND │
│ the verdict) │ │ (never scored) │ │ (opt-in, slow) │
├──────────────────┤ ├────────────────────┤ ├─────────────────┤
│ SIFT │ │ Shape matching │ │ Flower count │
│ ORB │ │ Texture matching │ │ (SAM3+YOLO-World)│
│ SuperPoint+Light- │ │ Color family grid │ │ Vase identity │
│ Glue ("SuperGlue")│ │ AI vision verify │ │ (DINOv2+CLIP) │
│ LoFTR │ │ │ │ │
│ Color space │ │ │ │ │
└──────────────────┘ └────────────────────┘ └─────────────────┘
```
---
## 2. High-Level Design
### 2.1 System Context
```
┌──────────┐ HTTPS ┌─────────────────────────────┐
│ Browser │◄────────────────────►│ Flask app (app.py) │
│ (user) │ upload / results │ torch17_new (Python 3.8) │
└──────────┘ │ │
│ ┌─────────────────────────┐ │
│ │ pipeline/* (in-process) │ │
│ │ SIFT ORB SuperGlue LoFTR │ │
│ │ Color Shape Texture │ │
│ │ YOLO-World DINOv2 CLIP │ │
│ │ rembg (background removal)│ │
│ └─────────────────────────┘ │
└───────┬───────────────┬───────┘
│ │
subprocess │ │ HTTPS (optional)
(SAM3 only) │ │
▼ ▼
┌────────────────────────┐ ┌──────────────────┐
│ sam3_worker.py │ │ External vision- │
│ sam2_env (Python 3.10) │ │ LLM endpoint │
│ transformers 5.5 + SAM3 │ │ (Cloudflare tunnel)│
│ facebook/sam3, 4-bit NF4 │ │ "AI verification" │
└────────────────────────┘ └──────────────────┘
┌───────────────────┐
│ NVIDIA GPU (8GB) │
│ shared by both │
│ processes, │
│ never concurrently │
│ loaded (see §4.7) │
└───────────────────┘
```
Everything runs on a **single 8GB consumer GPU** (GTX 1070, Pascal /
compute capability 6.1). That one constraint shapes almost every
architectural decision in this project: sequential (not concurrent) model
execution, aggressive image downscaling, explicit model unload/reload
cycles, 4-bit quantization for the newest model, and a cross-process split
for the one dependency that couldn't share a Python environment with
everything else.
### 2.2 Why Two Python Environments?
`facebook/sam3` (via Hugging Face `transformers`) needs
`transformers>=5.5.0`, which itself requires **Python ≥3.10**. The rest of
the app (all the deep-learning matchers, rembg, etc.) lives in a
`torch17_new` conda environment pinned to **Python 3.8** for compatibility
with older CUDA/driver combinations already validated there. Upgrading that
shared environment in place was judged too risky (could silently break
DINOv2/CLIP/LightGlue/rembg, all already working). Instead:
- SAM3 runs in a **separate conda env** (`sam2_env`, Python 3.10) that
already had `transformers 5.5.0` and `bitsandbytes` installed.
- The main Flask app talks to it via a **one-shot subprocess per request**
(`sam3_worker.py`), passing a JSON request file and reading back a JSON
response + PNG masks. See [§4.6](#46-sam3-cross-process-bridge) for the
exact contract.
- This keeps the two environments fully decoupled — SAM3 can be upgraded,
reinstalled, or even moved to a different machine/GPU without touching
the main app at all.
### 2.3 Core Design Principles
| Principle | Why |
|---|---|
| **Sequential GPU usage, never concurrent** | An 8GB card can't hold multiple deep models' activations at once without fragmentation-driven CUDA OOM. Every heavy pipeline runs its methods one at a time, with `torch.cuda.empty_cache()` between them. |
| **Aggressive, uniform downscaling** | Every image (upload *and* template) is capped to `MAX_IMAGE_DIM = 1600px` before touching any model — this single choke point is what keeps memory/time bounded regardless of how large the original photo was. |
| **Opt-in for heavyweight features** | Flower counting and vase-identity comparison unload the core matching models first (to free GPU headroom) and are only triggered by explicit user action — they must never slow down the default "upload → match" path. |
| **Scored vs. diagnostic separation** | Only 5 signals (SIFT, ORB, SuperGlue, LoFTR, Color) ever feed a number into the weighted verdict. Shape, texture, the color family grid, AI verification, flower count, and vase identity are **always** presented as separate, clearly-labeled sections that cannot move the score — this is a hard invariant enforced by keeping their computation entirely out of `_weighted_scores()`/`_overall_best()`. |
| **Two independent "final answer" aggregates** | A literal weighted **sum** (`_weighted_scores`) and a rank-based **Borda count** (`_overall_best`) are computed from the same four raw method outputs but can legitimately disagree — the weighted sum is the primary displayed verdict; Borda is a sanity cross-check. |
| **Soft-fail everything non-essential** | A failing color analysis, a stale AI-verification tunnel, a SAM3 subprocess crash — none of these ever take down the core match. Every optional section is wrapped in try/except and reported as its own `error` field. |
| **Concept prompting over heuristic filtering** | SAM3 replaced an earlier SAM1-based approach that had to *guess* what counted as "a flower" from size/position/overlap heuristics. Prompting SAM3 directly with the word "flower" (or "vase") does that semantic work at the model level instead. |
### 2.4 Technology Stack
| Layer | Technology |
|---|---|
| Backend web framework | Flask (development server, threaded) |
| Classical CV | OpenCV (SIFT, ORB, FLANN matching, RANSAC homography, K-means, GLCM via scikit-image) |
| Deep local features | SuperPoint + LightGlue (via the `lightglue` package) |
| Deep dense matching | LoFTR (via `kornia.feature`) |
| Background removal | `rembg` (BiRefNet-general-lite model), ONNX Runtime (CUDA + CPU fallback) |
| Concept segmentation | SAM3 (`facebook/sam3`, Hugging Face `transformers`, 4-bit NF4 via `bitsandbytes`) |
| Open-vocabulary detection | YOLO-World (`yolov8s-worldv2.pt`, via `ultralytics`) |
| Embedding similarity | DINOv2 (`facebook/dinov2-base`) + CLIP (`openai/clip-vit-base-patch32`), via `transformers` |
| External AI check | Third-party vision-LLM endpoint (Cloudflare tunnel), consumed over HTTPS |
| Frontend | Vanilla JS (no framework), hand-rolled DOM building, CSS custom properties for theming |
| Batch testing | Standalone Python scripts reusing the production pipeline (`testVaseMatcher.py`, `tester.py`) |
---
## 3. Feature Walkthrough
Presented in the order they were built, which is also roughly the order of
increasing sophistication:
1. **Core matching (SIFT / ORB / SuperGlue / LoFTR)** — four independent
keypoint/dense matchers, each scored by RANSAC inlier count against
every template, run strictly sequentially.
2. **Weighted final verdict + Borda overall-best** — two different ways of
combining the four raw scores into one winning template.
3. **Color space section** — HS-histogram intersection, shown as its own
section but also folded into the weighted verdict (`Color` is a 5th
contributor there only).
4. **Background-removal size fix** — uploads over 15MB are downscaled +
re-encoded rather than rejected outright.
5. **Color family grid** — dynamic K-means color-region discovery +
side-by-side region-matching visualization, purely informational.
6. **Shape matching** — Hu-moment contour distance + canonical-silhouette
IoU, its own section, never scored.
7. **Texture matching** — Local Binary Patterns + GLCM/Haralick features,
its own section, never scored.
8. **AI verification** — sends the matched template photo + the user's
original upload to an external vision-LLM endpoint for a free-text
QC-style comparison, parsed into a structured Match/Confidence/
Discrepancies/Description card.
9. **Flower counting v1 (SAM1)** — Segment Anything (ViT-B) in automatic
"segment everything" mode, filtered by size/position heuristics to
guess which proposals were flowers. Superseded by step 11.
10. **YOLO-World cross-check + vase exclusion** — open-vocabulary detection
used first to strip the vase/ribbon out of SAM1's flower count, and
kept afterward as an independent second opinion shown side by side.
11. **Flower counting v2 (SAM3) + tuned recall** — SAM1 replaced with SAM3
concept-prompted segmentation (prompt: `"flower"`), which never
proposes the vase/ribbon in the first place — no more heuristic
filtering needed. (An intermediate step also tuned SAM1's own
`points_stride`/`conf_thres`/`stability_score_thresh` for better
recall before the SAM3 replacement landed; that tuning knowledge is
preserved in this document for context even though the code path is
gone.)
12. **Vase-identity comparison (DINOv2 + CLIP)** — crops the vase out of
both photos using SAM3's precise mask (not just YOLO's bounding box),
background-blacks-out everything else in the crop, and reports a
Same/Uncertain/Different verdict from a weighted DINOv2+CLIP cosine
similarity.
13. **Flower-count mismatch explanation** — a plain-English note ("your
photo has 23 flowers vs. 9 in the template") surfaced wherever counts
differ, explaining a likely contributor to a lower match score.
14. **`tester.py`** — a lean batch-accuracy script that also emits one
combined report image per test photo (input segmentation, template
segmentation, counts, vase comparison) using the real production
pipeline.
15. **UI reorder** — the "Weighted final match" section moved to the very
top of the results, so the final verdict is visible without scrolling.
16. **Auto-triggered "Flower check"** — a lightweight, CLIP-only (no
DINOv2) flower-count + flower-similarity mini-panel shown beside the
weighted card, fired automatically on every confident match (unlike
the full flower-count feature, which stays behind its own button) —
with a loading spinner while it computes.
---
## 4. Low-Level Design
### 4.1 Directory Structure
```
featureTransform/
├── app.py Flask routes, logging setup, request glue
├── config.py All tunables; env-var capping; HF_TOKEN loading
├── sam3_worker.py Standalone SAM3 subprocess entrypoint (sam2_env)
├── tester.py Lean batch-accuracy + composite report images
├── testVaseMatcher.py Fuller batch-accuracy report (CSV, per-method)
├── .env HF_TOKEN=hf_... (gitignored, gated model access)
├── pipeline/
│ ├── engine.py Orchestrator: bootstrap, per-request pipeline,
│ │ weighted/Borda aggregation, opt-in features
│ ├── bg_removal.py rembg wrapper, disk caching, mask creation
│ ├── classical.py SIFT / ORB extraction + FLANN matching
│ ├── deep.py SuperPoint+LightGlue, LoFTR
│ ├── color.py HS-histogram + dominant-color comparison
│ ├── color_grid.py Dynamic K-means color-family grid
│ ├── shape_match.py Hu moments + canonical-silhouette IoU
│ ├── texture_match.py LBP + GLCM/Haralick
│ ├── flower_count.py SAM3-backed per-instance flower counting
│ ├── yolo_world.py Open-vocabulary flower/vase/ribbon detection
│ ├── vase_compare.py DINOv2 + CLIP vase-identity comparison
│ ├── sam3_client.py Subprocess bridge INTO sam3_worker.py
│ ├── verify.py External AI-verification HTTP client
│ └── utils.py Template listing, score annotation, compression
├── templates/
│ └── index.html Single-page UI (Jinja for the template gallery)
├── static/
│ ├── css/style.css
│ └── js/main.js All client-side rendering logic
├── cache/
│ ├── templates_nobg/ Background-removed templates, cached by name
│ └── uploads_nobg/ Background-removed uploads, cached by content hash
├── uploads/<request_id>/ Per-request working files (originals, overlays,
│ crops, family grids) -- swept after 6 hours
├── logs/app.log Rotating log, all pipeline.* + werkzeug output
└── test_reports/ tester.py's composite report images
```
### 4.2 Module Responsibilities
#### `config.py`
The single source of truth for every tunable constant in the system, plus
two pieces of process-startup plumbing that **must** run before any heavy
import:
1. Caps `OMP_NUM_THREADS`/`OPENBLAS_NUM_THREADS`/`MKL_NUM_THREADS`/
`NUMEXPR_NUM_THREADS`/`ORT_NUM_THREADS` to `min(4, cpu_count)` — done via
`os.environ.setdefault` at import time, before `cv2`/`onnxruntime`/
`torch` are imported anywhere in the process. Left uncapped, each
library grabs one thread per core, which starves everything else on a
shared machine.
2. Sets `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` — the caching
allocator's own recommended mitigation for fragmentation-driven CUDA
OOM on a small card.
It also reads `HF_TOKEN` once (env var first, then a tiny hand-rolled
`.env` parser as fallback) so the SAM3 subprocess can be handed it
explicitly without needing a `python-dotenv` dependency.
#### `pipeline/engine.py`
The orchestrator. Owns:
- `bootstrap()` — precomputes every template's background-removed image
and every method's template-side features/embeddings **once**, at
startup, so a live request only ever has to process the single uploaded
image.
- `process_upload()` / `_process_upload_locked()` — the `/api/match` code
path: background removal → sequential SIFT/ORB/SuperGlue/LoFTR →
best-image annotation → color/shape/texture analysis → weighted +
Borda aggregation → color family grid → response assembly.
- `_weighted_scores()` / `_overall_best()` — the two aggregation
algorithms (see [§4.4](#44-aggregation-layers)).
- `count_flowers()` — the full opt-in flower-count + vase-comparison
feature (SAM3 + YOLO-World + DINOv2/CLIP, renders visuals, returns
everything for the "Count flowers" button).
- `flower_summary()` — the lightweight, auto-triggered companion (SAM3 +
CLIP only, no YOLO-World, no DINOv2, no rendered images) shown beside
the weighted card on every confident match.
- A single `threading.Lock` (`_pipeline_lock`) serializes every request's
pipeline end-to-end, so at most one heavy CPU/GPU operation is ever in
flight process-wide.
#### `pipeline/bg_removal.py`
Wraps `rembg` (BiRefNet-general-lite). Lazily creates an ONNX Runtime
session (`CUDAExecutionProvider` first, falls back to CPU on failure —
this fallback is "sticky" for the process's lifetime, a known limitation).
Disk-caches results: uploads by content hash, templates by filename.
`resize_max_dim()` is the single choke point enforcing `MAX_IMAGE_DIM`
everywhere. `unload_session()` drops the session so `count_flowers()`/
`flower_summary()` can free its GPU memory before the heavier optional
models load.
#### `pipeline/classical.py`
SIFT and ORB feature extraction + FLANN-based matching. Template
descriptors are precomputed once and cached in-memory (`_template_features`).
See [§4.3](#43-scoring-algorithms) for the exact scoring formula.
#### `pipeline/deep.py`
SuperPoint+LightGlue (labeled "SuperGlue" in the UI, matching an existing
naming convention) and LoFTR. `unload_models()` drops all three model
objects (SuperPoint, LightGlue, LoFTR) so the opt-in flower/vase features
can reclaim their GPU memory; they lazily reload on the next `/api/match`
call, identical to a fresh process start.
#### `pipeline/color.py`
Independent, non-scored-looking-but-actually-scored (folded into the
weighted verdict via `Color`) comparison: Hue+Saturation 2D histogram
(Value/brightness deliberately excluded for lighting invariance),
compared via histogram intersection. Also computes dominant colors
(K-means in a bounded sample) and per-color Lab-distance similarity pairs
for the visual palette display.
#### `pipeline/color_grid.py`
Dynamic K-means color-region discovery in LAB space (lightness
down-weighted so shadows of the same hue cluster together), rendered as a
tile grid, region-matched by nearest average color, with an
area-match-percentage per region and an overall area-weighted headline
number. Purely visual — tied to whichever template the *weighted* verdict
already picked.
#### `pipeline/shape_match.py`
Two complementary shape signals, averaged into one `match_pct`:
- **Hu moments** (`cv2.matchShapes`, `CONTOURS_MATCH_I1`) — translation/
rotation/scale-invariant contour distance, converted to a similarity %.
- **Canonical-silhouette IoU** — both masks cropped to their own bounding
box, resized/centered into a fixed square canvas, then measured as
direct pixel overlap. This also produces the side-by-side + overlay
visualization images.
#### `pipeline/texture_match.py`
Two complementary texture signals, averaged into one `match_pct`:
- **Local Binary Patterns** (uniform method), compared as a histogram
intersection — catches fine, repetitive patterns (fabric weave, petal
grain).
- **GLCM/Haralick features** (contrast, homogeneity, energy, correlation),
compared as a normalized relative difference, averaged across
properties and across 4 angles/2 distances — catches coarser
smooth-vs-rough material differences.
#### `pipeline/flower_count.py`
SAM3-backed per-instance flower counting. `count_flowers(bgr, mask,
workdir, instances_raw=...)` accepts already-fetched SAM3 instances (so
`engine.py` can batch multiple images/prompts into one subprocess call
and hand each result set to this function separately), applies a light
foreground-overlap sanity filter, then color-clusters the survivors (LAB
K-means, capped at `SAM_MAX_KIND_CLUSTERS`) as a rough proxy for distinct
flower "kinds". `union_mask()` OR-combines every instance mask into one —
used to crop "just the flower material" for CLIP similarity.
`render_instances()` draws the numbered, color-coded overlay.
#### `pipeline/yolo_world.py`
Open-vocabulary detection (`yolov8s-worldv2.pt`) prompted with
`["flower", "vase", "ribbon", "bow"]`. Kept as an **independent second
opinion** shown side by side with SAM3's per-instance count — it draws
one box per contiguous flower region rather than per bloom, so it's a
coarser, corroborating signal, not a replacement.
#### `pipeline/vase_compare.py`
- `crop_mask(bgr, mask_bool)` — crops the bounding box of any mask (padded
by `VASE_CROP_PAD_FRAC`), blacking out every pixel the mask doesn't
cover. Generic despite the module name — also used for the flower-only
crop in `flower_summary()`.
- `compare_vases()` — DINOv2 + CLIP cosine similarity, combined as a
weighted average (`VASE_DINO_WEIGHT=0.6`, `VASE_CLIP_WEIGHT=0.4`),
bucketed into `same`/`uncertain`/`different`.
- `clip_similarity_pct()` — CLIP-only variant (never touches DINOv2),
used by the lightweight auto-triggered flower check.
- `unload_models()` — drops both model objects; a no-op for whichever one
was never loaded (e.g. DINOv2 stays unloaded entirely on the
CLIP-only path).
#### `pipeline/sam3_client.py`
The subprocess bridge (runs *in* the main app's Python 3.8 process). Saves
each named input image to a temp PNG, writes a JSON request file
describing every `(image, prompt, threshold)` job, invokes
`sam3_worker.py` under `sam2_env`'s Python interpreter with `HF_TOKEN`
passed through the subprocess environment, waits (bounded by
`SAM3_TIMEOUT_SECONDS`), and parses the JSON response + PNG masks back
into `{(image_key, prompt): [{"mask": bool ndarray, "score": float, "box": [...]}]}`.
Always cleans up its temp directory, even on failure.
#### `sam3_worker.py`
Standalone script, **runs under `sam2_env`'s Python 3.10**, has no
dependency on this app's `config`/`pipeline` modules. Loads `Sam3Model` +
`Sam3Processor` 4-bit quantized (NF4 via `BitsAndBytesConfig`), processes
every job in the request against the already-loaded model (one model load
serves every job in a request), writes one PNG per detected instance plus
a `response.json` manifest.
#### `pipeline/verify.py`
Thin HTTP client for the external AI-verification endpoint. Posts both
images as multipart form data, parses the endpoint's free-text `result`
field with a tolerant regex (`DISCREP\w*` rather than a literal spelling,
since the LLM behind it doesn't always spell "DISCREPANCIES" consistently)
into `{match, confidence, discrepancies, description}`.
#### `pipeline/utils.py`
Template file listing, score-annotation overlay text, PNG encoding, and
`compress_image_bytes()` — downscales + re-encodes as JPEG only if an
upload exceeds `COMPRESS_ABOVE_BYTES`, reusing `MAX_IMAGE_DIM` since the
pipeline downsamples to that anyway.
### 4.3 Scoring Algorithms
#### SIFT / ORB (`pipeline/classical.py`)
```
1. Extract keypoints + descriptors (masked to foreground)
2. FLANN k=2 nearest-neighbor match against template descriptors
3. Lowe's ratio test: keep match if best.distance < 0.75 * secondBest.distance
4. RANSAC homography (5.0px threshold) across surviving matches
5. score = RANSAC inlier count (raw integer)
confidence = inliers / good_matches * 100 (%)
```
#### SuperGlue = SuperPoint + LightGlue (`pipeline/deep.py`)
Same shape as above, but SuperPoint (neural keypoints/descriptors) replaces
SIFT/ORB's detector and LightGlue (learned matcher) replaces FLANN+ratio
test. Matches filtered to inside the foreground mask before the same
RANSAC-inlier-count scoring.
#### LoFTR (`pipeline/deep.py`)
Dense matcher — no keypoint detection step. Directly regresses pixel
correspondences, keeps only ones above `LOFTR_CONFIDENCE_THRESHOLD` (0.5),
filters to inside the mask, then the same RANSAC-inlier-count scoring.
#### Color space (`pipeline/color.py`)
```
hist = HS 2D histogram (Hue×Saturation only, Value excluded), normalized to sum=1
match_pct = Σ min(hist_input[i], hist_template[i]) × 100 (histogram intersection)
```
#### Shape matching (`pipeline/shape_match.py`)
```
hu_similarity = max(0, 100 × (1 - matchShapes(contour_a, contour_b) / 1.5))
iou = |mask_a ∩ mask_b| / |mask_a mask_b| × 100 (canonical-aligned)
match_pct = (hu_similarity + iou) / 2
```
#### Texture matching (`pipeline/texture_match.py`)
```
lbp_similarity = Σ min(hist_a[i], hist_b[i]) × 100 (LBP histogram intersection)
glcm_similarity = mean over 4 properties of max(0, 1 - |a-b|/max(|a|,|b|)) × 100
match_pct = (lbp_similarity + glcm_similarity) / 2
```
#### Vase identity (`pipeline/vase_compare.py`)
```
cosine_pct(a, b) = max(0, min(1, a·b)) × 100 (a, b are L2-normalized embeddings)
dino_pct = cosine_pct(DINOv2(crop_a), DINOv2(crop_b))
clip_pct = cosine_pct(CLIP(crop_a), CLIP(crop_b))
combined_pct = 0.6 × dino_pct + 0.4 × clip_pct
verdict = "same" if combined_pct >= 75
"uncertain" if combined_pct >= 60
"different" otherwise
```
#### Flower CLIP similarity (`flower_summary`, CLIP-only variant)
```
flower_clip_pct = cosine_pct(CLIP(union_of_flower_crops_a), CLIP(union_of_flower_crops_b))
```
### 4.4 Aggregation Layers
Two genuinely different "who won" computations exist side by side, and are
allowed to disagree:
**Weighted final match** (`engine._weighted_scores`) — a literal weighted
**sum** of raw scores:
```
weighted_score(template) = Σ_method WEIGHT[method] × raw_score(method, template)
WEIGHT = { LoFTR: 0.75, SuperGlue: 0.15, SIFT: 0.10, ORB: 0.10, Color: 0.7 }
```
Because this sums *raw* scores (inlier counts, which run much higher for
LoFTR/SuperGlue than SIFT/ORB, plus Color's 0-100 percentage), a method
with naturally larger magnitudes pulls harder on the total even at a
similar or lower weight. That's the deliberate behavior of the formula as
specified, not a bug — see the comment block in `config.py` above
`METHOD_WEIGHTS`.
**Overall best** (`engine._overall_best`, Borda count) — rank-based, and
excludes `Color` entirely (only the original four methods participate):
```
for each method:
rank all templates by that method's own raw score
template at rank i (0-indexed) earns (n - i) points
overall_best = template with the highest total points across all methods
```
Rank-based aggregation means no single method's raw-score magnitude can
dominate the way it can in the weighted sum — this is why the two verdicts
are kept, and shown, separately rather than collapsed into one number.
### 4.5 Request Lifecycles
#### `POST /api/match`
```
Browser Flask (app.py) engine.py GPU
│ upload image │ │ │
├─────────────────────────► │ │
│ │ compress if >15MB │ │
│ ├──────────────────────────► │ │
│ │ │ bg_removal (rembg) │
│ │ ├────────────────────►│
│ │ │◄────────────────────┤
│ │ │ SIFT (sequential) │
│ │ ├────────────────────►│
│ │ │ ORB │
│ │ ├────────────────────►│
│ │ │ SuperGlue │
│ │ ├────────────────────►│
│ │ │ LoFTR │
│ │ ├────────────────────►│
│ │ │ Color / Shape / │
│ │ │ Texture (CPU) │
│ │ │ weighted + Borda │
│ │ │ color family grid │
│ │◄───────────────────────────┤ │
│◄───────────────────────── JSON result │ │
```
SuperPoint/LightGlue/LoFTR **stay loaded** in memory after this request
(for fast repeated matching) — they are only unloaded when an opt-in
flower/vase feature explicitly needs the GPU headroom.
#### `POST /api/count_flowers` (opt-in, "Count flowers" button)
```
Browser app.py engine.count_flowers() sam3_client → sam3_worker.py (sam2_env) YOLO-World DINOv2+CLIP
│ click │ │ │ │ │
├───────────────► │ │ │ │
│ ├────────────────────► │ │ │
│ │ │ unload SuperPoint/LightGlue/LoFTR + rembg session │ │
│ │ │ yolo_world.detect() ────────────────────────────────────────────►│ │
│ │ │ batch: input+flower, input+vase, [template+vase, template+flower] │ │
│ │ ├─────────────────────────────► │ │
│ │ │ (one SAM3 model load serves every job) │ │
│ │ │◄─────────────────────────────┤ │ │
│ │ │ cluster + render SAM instances (input & template) │ │
│ │ │ render YOLO-World boxes │ │
│ │ │ vase_compare (DINOv2+CLIP) ──────────────────────────────────────────────────►│
│ │ │◄─────────────────────────────────────────────────────────────────────────────┤
│ │ │ unload YOLO-World, DINOv2, CLIP │ │
│ │◄───────────────────┤ │ │
│◄─────────────── JSON + image URLs │ │ │
```
#### `POST /api/flower_summary` (auto-triggered, beside the weighted card)
```
Browser app.py engine.flower_summary() sam3_client/worker CLIP only
│ (fires automatically │ │ │
│ right after a confident match) │ │ │
├────────────────────────────────────► │ │
│ │ unload matching models │ │
│ │ batch: input+flower, │ │
│ │ template+flower │ │
│ ├─────────────────────────────►│ │
│ │◄─────────────────────────────┤ │
│ │ count comparison (no render)│ │
│ │ union masks, crop, CLIP-only similarity ────────────►│
│ │◄──────────────────────────────────────────────────────┤
│ │ unload CLIP (DINOv2 never loaded) │
│◄──────────────────────────────────── │
```
Deliberately skips YOLO-World, DINOv2, and any rendered overlay images — it
exists purely to be fast enough to run unconditionally on every match.
#### `POST /api/verify` (AI verification, auto-fires after a confident match)
```
Browser app.py verify.py External vision-LLM endpoint
│ │ │ │
├──────────────► │ │
│ ├────────────────► POST multipart (image1, image2)
│ │ ├─────────────────────────►│
│ │ │◄─────────────────────────┤
│ │ │ parse MATCH/CONFIDENCE/ │
│ │ │ DISCREPANCIES/description│
│ │◄──────────────── │
│◄────────────── JSON (or 502 if endpoint unreachable) │
```
Failures here (stale Cloudflare tunnel, timeout) are surfaced as a soft
error in the UI, never a hard failure of the page.
### 4.6 SAM3 Cross-Process Bridge
**Request** (JSON file, written by `sam3_client.py`, read by `sam3_worker.py`):
```json
{
"images": {
"input": "/path/to/uploads/<request_id>/sam3_xxxx/input.png",
"template": "/path/to/uploads/<request_id>/sam3_xxxx/template.png"
},
"jobs": [
{"image": "input", "prompt": "flower", "threshold": 0.5},
{"image": "input", "prompt": "vase", "threshold": 0.3},
{"image": "template", "prompt": "vase", "threshold": 0.3}
],
"output_dir": "/path/to/uploads/<request_id>/sam3_xxxx/out"
}
```
**Response** (`response.json`, written by the worker):
```json
{
"results": [
{"image": "input", "prompt": "flower", "instances": [
{"mask_file": "input_flower_0_0.png", "score": 0.7486, "box": [x1, y1, x2, y2]},
{"mask_file": "input_flower_0_1.png", "score": 0.7006, "box": [...]}
]},
{"image": "input", "prompt": "vase", "instances": [
{"mask_file": "input_vase_1_0.png", "score": 0.9391, "box": [...]}
]}
],
"error": null
}
```
Each `mask_file` is a standalone grayscale PNG (0/255) at the original
image's resolution — `sam3_client.py` reads it back with
`cv2.imread(..., IMREAD_GRAYSCALE) > 127` to reconstruct a boolean mask.
On any failure, `error` is a string (worker traceback included) and the
process exits non-zero; `sam3_client.py` raises `Sam3Error` either way,
which every caller treats as a soft failure.
**Why 4-bit (NF4) quantization**: the unquantized checkpoint is ~3.2GB
(`model.safetensors`). Loading it naively (8-bit, `device_map="cuda:0"`)
hit a CUDA OOM during Transformers' own memory pre-flight ("caching
allocator warmup") step on this 8GB card, because that warmup reserves
capacity based on the *original* dtype size before quantization actually
reduces it. Switching to `load_in_4bit=True` + `bnb_4bit_quant_type="nf4"`
+ `low_cpu_mem_usage=True` fixed this: ~700MB resident, ~1.9GB peak during
inference, ~6s to load, ~2s per prompt.
**Why `sam.predictor(...)` isn't called directly for automatic-mode
tuning** (historical note, from the SAM1 era): Ultralytics' high-level
`model.predict()`/`model()` call validates every kwarg against a fixed
CLI-style config schema, which rejects automatic-mode-only parameters like
`points_stride`/`conf_thres` with a `SyntaxError`. The fix at the time was
calling `model.predictor(...)` directly (after a one-time warmup call to
force predictor construction), bypassing that validation entirely since
`BasePredictor.__call__` forwards arbitrary kwargs straight through. This
code path no longer exists (SAM1 was fully replaced by SAM3), but the
technique is documented here in case a similar situation recurs with
another Ultralytics-wrapped model (e.g. YOLO-World).
### 4.7 GPU Memory Management Strategy
```
┌───────────────────────────────────────────┐
│ Idle / after /api/match │
│ SuperPoint + LightGlue + LoFTR RESIDENT │
│ (~200-900MB, kept warm for fast matching) │
└──────────────────┬────────────────────────┘
│ user clicks "Count flowers"
│ or a confident match fires
│ the auto flower-check
┌───────────────────────────────────────────┐
│ deep.unload_models() │
│ bg_removal.unload_session() │
│ torch.cuda.empty_cache() │
│ → GPU now near-baseline (~200-900MB) │
└──────────────────┬────────────────────────┘
┌───────────────────────────────────────────┐
│ SAM3 (separate PROCESS, sam2_env) │
│ ~700MB resident / ~1.9GB peak │
│ + optionally YOLO-World (~200MB) │
│ + optionally DINOv2 (~350MB) + CLIP (~600MB) │
└──────────────────┬────────────────────────┘
│ request finishes
┌───────────────────────────────────────────┐
│ yolo_world.unload_model() │
│ vase_compare.unload_models() │
│ (SAM3's own process has already exited — │
│ its GPU memory is freed by the OS/driver, │
│ nothing to unload on the main-process side) │
│ → back to near-baseline │
└───────────────────────────────────────────┘
```
Every model getter (`deep.get_superpoint()`, `flower_count`'s SAM3 client,
`vase_compare._get_dino()`/`_get_clip()`, `yolo_world.get_model()`) is a
lazy singleton — `None` until first use, reset to `None` by the matching
`unload_*()` function. This means "unload" is always safe to call even if
the model was never loaded (checked via `is not None`), and every model
transparently reloads itself on the next call that needs it, with no
special-casing required anywhere else in the codebase.
### 4.8 Frontend Structure (`static/js/main.js`)
Single IIFE, no framework, no build step. Key structure:
- **Upload flow**: drag/drop or file picker → `fetch POST /api/match`
`renderResults(data)`.
- **`renderResults(data)`** populates every section in DOM order (weighted
card first, then overall-best, query images, method grid, flower count,
color/shape/texture/family-grid sections), then fires two
**auto-triggered, non-blocking** follow-ups: `runVerification()` (AI
check) and `runFlowerSummary()` (the lightweight flower-check panel) —
neither blocks the initial render, both show their own loading state and
fail independently.
- **`renderWeighted(data)`** builds the terracotta-gradient weighted card
(input/template thumbnails, breakdown table, Strong/Moderate/Weak
confidence pill derived purely from the *margin* between 1st and 2nd
place — display-only, cannot feed back into the algorithm).
- **Opt-in flower count**: `countFlowersBtn` click handler → `fetch POST
/api/count_flowers` → `renderFlowerCount(data)`, which in turn calls
`renderFlowerCountMismatch()` and `renderVaseComparison()`.
- Every async section follows the same pattern: hide body / show spinner →
fetch → on success show body, on failure show a scoped error message —
never a page-level failure.
---
## 5. API Reference
| Method | Path | Body | Returns |
|---|---|---|---|
| GET | `/` | — | Rendered HTML page (template gallery + upload form) |
| GET | `/template_image/<filename>` | — | Raw template image bytes |
| GET | `/uploads/<request_id>/<filename>` | — | Any file written into that request's working directory |
| POST | `/api/match` | multipart `image` file | Full match result: per-method scores, weighted verdict, Borda overall-best, color/shape/texture/family-grid analysis |
| POST | `/api/count_flowers` | JSON `{request_id, template}` | SAM3 + YOLO-World flower counts (with visuals), vase comparison, flower-count-mismatch note |
| POST | `/api/flower_summary` | JSON `{request_id, template}` | Lightweight flower-count comparison + CLIP-only flower similarity (no images) |
| POST | `/api/verify` | JSON `{request_id, template}` | External AI vision-LLM verdict: match/confidence/discrepancies/description |
All POST endpoints return `{"error": "..."}` with a non-200 status on
failure; the frontend treats every one of these as scoped/local, never
page-fatal.
---
## 6. Configuration Reference
All of the following live in `config.py`, grouped by feature:
| Setting | Default | Purpose |
|---|---|---|
| `MAX_IMAGE_DIM` | 1600 | Every image downscaled to this before any model touches it |
| `METHOD_WEIGHTS` | `{LoFTR: .75, SuperGlue: .15, SIFT: .10, ORB: .10, Color: .7}` | Weighted-sum verdict weights |
| `SCORE_THRESHOLD` | per-method | "Confident match" badge cutoff (doesn't affect ranking) |
| `MAX_CONTENT_LENGTH_BYTES` | 50MB | Hard reject ceiling |
| `COMPRESS_ABOVE_BYTES` | 15MB | Soft threshold — downscale+recompress instead of reject |
| `SHAPE_HU_DISTANCE_SCALE` | 1.5 | Divisor turning Hu-moment distance into a 0-100% similarity |
| `TEXTURE_LBP_RADIUS` / `TEXTURE_GLCM_LEVELS` | 2 / 32 | Texture descriptor parameters |
| `FAMILY_GRID_K` | 5 | Number of color families discovered per image |
| `SAM3_FLOWER_THRESHOLD` / `SAM3_VASE_THRESHOLD` | 0.5 / 0.3 | Confidence cutoff per SAM3 concept prompt |
| `SAM3_TIMEOUT_SECONDS` | 180 | Subprocess timeout for the SAM3 worker |
| `VASE_DINO_WEIGHT` / `VASE_CLIP_WEIGHT` | 0.6 / 0.4 | Vase-comparison combined-score weights |
| `VASE_SAME_THRESHOLD` / `VASE_UNCERTAIN_THRESHOLD` | 75 / 60 | Vase verdict cutoffs (untuned — no ground-truth calibration set exists yet) |
| `YOLO_WORLD_CLASSES` | `["flower", "vase", "ribbon", "bow"]` | Open-vocabulary prompt classes |
| `VERIFY_ENDPOINT_URL` | Cloudflare tunnel URL | Can go stale if the tunnel restarts — treated as a soft failure |
---
## 7. Environment & Setup
### Two conda environments
| Env | Python | Purpose | Key packages |
|---|---|---|---|
| `torch17_new` | 3.8 | Main Flask app | `torch`, `opencv-python`, `kornia`, `lightglue`, `rembg`, `ultralytics`, `transformers==4.46.3` |
| `sam2_env` | 3.10 | SAM3 subprocess only | `torch`, `transformers>=5.5.0`, `bitsandbytes`, `accelerate` |
### Model checkpoints / weights
| Model | Source | Notes |
|---|---|---|
| SIFT / ORB | Built into OpenCV | No download needed |
| SuperPoint + LightGlue | `lightglue` package | Auto-downloaded on first use |
| LoFTR | `kornia.feature` | Auto-downloaded on first use |
| rembg (BiRefNet-general-lite) | `rembg` package | Auto-downloaded on first use |
| YOLO-World | `yolov8s-worldv2.pt` | Downloaded via `ultralytics`, cached locally |
| DINOv2-base | `facebook/dinov2-base` | Hugging Face, public, no token needed |
| CLIP ViT-B/32 | `openai/clip-vit-base-patch32` | Hugging Face, public, no token needed |
| **SAM3** | `facebook/sam3` | **Gated** on Hugging Face — requires requesting access on the model page and an approved access token |
### `.env` file
```
HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
```
Required only for the SAM3-backed features (flower counting, vase
comparison, the auto-triggered flower check). Without it, those specific
features fail softly with a clear "SAM3 unavailable" message — core
matching is entirely unaffected.
### Running the app
```bash
cd /media/suman/Backup_of_extra_/Sasi/featureTransform
conda activate torch17_new
python app.py
# serves on http://0.0.0.0:5053
```
---
## 8. Testing Tools
Two batch-test scripts exist, covering different needs:
| | `testVaseMatcher.py` | `tester.py` |
|---|---|---|
| Output | Detailed CSV + per-method accuracy table | Console pass/fail + one composite report image per photo |
| Coverage | All 4 methods + weighted + Borda + color pick, individually tallied | Weighted verdict only |
| Flower/vase check | Numbers only (family-grid area-match %), no images | Full: SAM3 flower segmentation (input vs. template), counts, vase comparison — all rendered into one PNG |
| Speed | Faster (no SAM3/CLIP round trip) | Slower (~15-25s/image extra for the flower/vase step); `--skip-flowers` for a fast pass |
| Use case | Regression-testing every individual method's accuracy | Visual QA — "does this match make sense at a glance" |
Both expect the same input layout:
```
<test_dir>/
SKU_1/ photo1.jpg photo2.jpg ...
SKU_2/ ...
```
where each subfolder name must exactly match an existing template name.
**Common mistake**: pointing either script directly at a `SKU_X/` folder
instead of its parent — the script silently finds 0 images, since it
expects to `listdir()` a folder of *subfolders*, not a folder of images.
```bash
python tester.py # ./flowers, full report images
python tester.py /path/to/flowers --limit 3
python tester.py --skip-flowers # fast accuracy-only pass
python testVaseMatcher.py /path/to/flowers --csv report.csv
```
---
## 9. Known Limitations & Design Tradeoffs
- **rembg's CUDA fallback is sticky for the process's lifetime.** A
transient CUDA OOM during session creation permanently falls back that
process to CPU-only background removal (much slower) until restarted.
No auto-retry/self-healing exists yet.
- **Vase-identity and flower-count-mismatch thresholds are untuned.**
There is no ground-truth "same/different vase" or "acceptable flower
count delta" labeled dataset — the cutoffs in `config.py` are
reasonable defaults, not calibrated ones.
- **The weighted verdict's literal-sum formula means raw-score magnitude
matters, not just weight.** LoFTR/SuperGlue's inlier counts run far
higher than SIFT/ORB's, so they can dominate the total even when their
configured weight isn't proportionally larger. This is documented as
intentional (matches the formula as originally specified), not a bug.
- **YOLO-World's own "flower" count is coarse** — one box per contiguous
flower region, not per bloom — so it will never numerically match SAM3's
per-instance count; it's shown as an independent corroborating signal,
not something expected to agree.
- **The AI-verification endpoint is a Cloudflare tunnel URL** that can go
stale if restarted on the other end; failures are always soft (shown as
"AI verification unavailable"), never block the page.
- **SAM3 requires a gated, manually-approved Hugging Face token.** The app
cannot obtain this automatically — a human must request access at
`huggingface.co/facebook/sam3` and provide a token with access granted.
- **Everything assumes a single physical GPU with ~8GB VRAM** and Pascal
or later compute capability. The sequential-execution and
unload/reload discipline exists specifically because of that constraint
— a larger card would allow (but doesn't require) more concurrency.
---
## 10. Glossary
| Term | Meaning |
|---|---|
| **Inlier count** | Number of matched keypoints that agree with one consistent RANSAC-fit geometric transform between two images — the raw score for SIFT/ORB/SuperGlue/LoFTR |
| **Weighted verdict** | The primary displayed "final answer" — literal weighted sum of 5 methods' raw scores |
| **Borda count / Overall best** | A secondary, rank-based cross-check aggregate of 4 methods (excludes Color) |
| **Concept segmentation** | SAM3's mode of operation: prompted with a plain-English word ("flower"), returns instance masks for exactly that concept, rather than segmenting everything indiscriminately |
| **NF4** | "NormalFloat4" — the 4-bit quantization scheme used (via `bitsandbytes`) to fit SAM3 in this GPU's memory budget |
| **Flower check** | The lightweight, auto-triggered (not opt-in), CLIP-only flower-count + flower-similarity panel shown beside the weighted card |
| **Count flowers** | The full, opt-in, button-triggered feature: SAM3 + YOLO-World flower counts, rendered overlays, and DINOv2+CLIP vase comparison |