Vase Matcher
A Flask web application that takes a photograph of a flower arrangement / vase and identifies which item in a fixed template catalogue it most closely resembles — judged independently by eleven different computer vision / deep learning signals, combined into one weighted verdict plus a set of independent diagnostic sections (color, shape, texture, flower count, vase identity) that never influence the score but explain why a match looks the way it does.
This document covers the system from High-Level Design (what the pieces are and why they exist) down to Low-Level Design (exact algorithms, data contracts, file-by-file responsibilities).
Table of Contents
- Overview
- High-Level Design
- Feature Walkthrough
- Low-Level Design
- API Reference
- Configuration Reference
- Environment & Setup
- Testing Tools
- Known Limitations & Design Tradeoffs
- Glossary
1. Overview
Problem: given a customer's photo of a vase/flower arrangement they received, determine which SKU in a small, fixed template catalogue it is, and explain the confidence of that determination in a way a human can sanity-check.
Approach: no single CV method is reliable enough alone (classical keypoint matchers fail on repetitive floral texture; color alone ignores shape; deep matchers alone ignore the vase). So the app runs a stack of independent methods, combines four of them into a literal weighted sum for the primary verdict, aggregates the same four by rank (Borda count) as a cross-check, and surfaces several more signals (color, shape, texture, flower count, vase identity, an external AI vision check) purely as explanatory diagnostics that a human reviewer can use to understand why the weighted verdict landed where it did — without ever touching the score itself.
┌─────────────────────────────┐
│ Upload a vase photo │
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ Which of N catalogue SKUs │
│ does this most resemble? │
└──────────────┬───────────────┘
│
┌─────────────────────┼─────────────────────┐
│ │ │
┌────────▼────────┐ ┌─────────▼─────────┐ ┌────────▼────────┐
│ SCORED (feeds │ │ DIAGNOSTIC ONLY │ │ ON-DEMAND │
│ the verdict) │ │ (never scored) │ │ (opt-in, slow) │
├──────────────────┤ ├────────────────────┤ ├─────────────────┤
│ SIFT │ │ Shape matching │ │ Flower count │
│ ORB │ │ Texture matching │ │ (SAM3+YOLO-World)│
│ SuperPoint+Light- │ │ Color family grid │ │ Vase identity │
│ Glue ("SuperGlue")│ │ AI vision verify │ │ (DINOv2+CLIP) │
│ LoFTR │ │ │ │ │
│ Color space │ │ │ │ │
└──────────────────┘ └────────────────────┘ └─────────────────┘
2. High-Level Design
2.1 System Context
┌──────────┐ HTTPS ┌─────────────────────────────┐
│ Browser │◄────────────────────►│ Flask app (app.py) │
│ (user) │ upload / results │ torch17_new (Python 3.8) │
└──────────┘ │ │
│ ┌─────────────────────────┐ │
│ │ pipeline/* (in-process) │ │
│ │ SIFT ORB SuperGlue LoFTR │ │
│ │ Color Shape Texture │ │
│ │ YOLO-World DINOv2 CLIP │ │
│ │ rembg (background removal)│ │
│ └─────────────────────────┘ │
└───────┬───────────────┬───────┘
│ │
subprocess │ │ HTTPS (optional)
(SAM3 only) │ │
▼ ▼
┌────────────────────────┐ ┌──────────────────┐
│ sam3_worker.py │ │ External vision- │
│ sam2_env (Python 3.10) │ │ LLM endpoint │
│ transformers 5.5 + SAM3 │ │ (Cloudflare tunnel)│
│ facebook/sam3, 4-bit NF4 │ │ "AI verification" │
└────────────────────────┘ └──────────────────┘
│
▼
┌───────────────────┐
│ NVIDIA GPU (8GB) │
│ shared by both │
│ processes, │
│ never concurrently │
│ loaded (see §4.7) │
└───────────────────┘
Everything runs on a single 8GB consumer GPU (GTX 1070, Pascal / compute capability 6.1). That one constraint shapes almost every architectural decision in this project: sequential (not concurrent) model execution, aggressive image downscaling, explicit model unload/reload cycles, 4-bit quantization for the newest model, and a cross-process split for the one dependency that couldn't share a Python environment with everything else.
2.2 Why Two Python Environments?
facebook/sam3 (via Hugging Face transformers) needs
transformers>=5.5.0, which itself requires Python ≥3.10. The rest of
the app (all the deep-learning matchers, rembg, etc.) lives in a
torch17_new conda environment pinned to Python 3.8 for compatibility
with older CUDA/driver combinations already validated there. Upgrading that
shared environment in place was judged too risky (could silently break
DINOv2/CLIP/LightGlue/rembg, all already working). Instead:
- SAM3 runs in a separate conda env (
sam2_env, Python 3.10) that already hadtransformers 5.5.0andbitsandbytesinstalled. - The main Flask app talks to it via a one-shot subprocess per request
(
sam3_worker.py), passing a JSON request file and reading back a JSON response + PNG masks. See §4.6 for the exact contract. - This keeps the two environments fully decoupled — SAM3 can be upgraded, reinstalled, or even moved to a different machine/GPU without touching the main app at all.
2.3 Core Design Principles
| Principle | Why |
|---|---|
| Sequential GPU usage, never concurrent | An 8GB card can't hold multiple deep models' activations at once without fragmentation-driven CUDA OOM. Every heavy pipeline runs its methods one at a time, with torch.cuda.empty_cache() between them. |
| Aggressive, uniform downscaling | Every image (upload and template) is capped to MAX_IMAGE_DIM = 1600px before touching any model — this single choke point is what keeps memory/time bounded regardless of how large the original photo was. |
| Opt-in for heavyweight features | Flower counting and vase-identity comparison unload the core matching models first (to free GPU headroom) and are only triggered by explicit user action — they must never slow down the default "upload → match" path. |
| Scored vs. diagnostic separation | Only 5 signals (SIFT, ORB, SuperGlue, LoFTR, Color) ever feed a number into the weighted verdict. Shape, texture, the color family grid, AI verification, flower count, and vase identity are always presented as separate, clearly-labeled sections that cannot move the score — this is a hard invariant enforced by keeping their computation entirely out of _weighted_scores()/_overall_best(). |
| Two independent "final answer" aggregates | A literal weighted sum (_weighted_scores) and a rank-based Borda count (_overall_best) are computed from the same four raw method outputs but can legitimately disagree — the weighted sum is the primary displayed verdict; Borda is a sanity cross-check. |
| Soft-fail everything non-essential | A failing color analysis, a stale AI-verification tunnel, a SAM3 subprocess crash — none of these ever take down the core match. Every optional section is wrapped in try/except and reported as its own error field. |
| Concept prompting over heuristic filtering | SAM3 replaced an earlier SAM1-based approach that had to guess what counted as "a flower" from size/position/overlap heuristics. Prompting SAM3 directly with the word "flower" (or "vase") does that semantic work at the model level instead. |
2.4 Technology Stack
| Layer | Technology |
|---|---|
| Backend web framework | Flask (development server, threaded) |
| Classical CV | OpenCV (SIFT, ORB, FLANN matching, RANSAC homography, K-means, GLCM via scikit-image) |
| Deep local features | SuperPoint + LightGlue (via the lightglue package) |
| Deep dense matching | LoFTR (via kornia.feature) |
| Background removal | rembg (BiRefNet-general-lite model), ONNX Runtime (CUDA + CPU fallback) |
| Concept segmentation | SAM3 (facebook/sam3, Hugging Face transformers, 4-bit NF4 via bitsandbytes) |
| Open-vocabulary detection | YOLO-World (yolov8s-worldv2.pt, via ultralytics) |
| Embedding similarity | DINOv2 (facebook/dinov2-base) + CLIP (openai/clip-vit-base-patch32), via transformers |
| External AI check | Third-party vision-LLM endpoint (Cloudflare tunnel), consumed over HTTPS |
| Frontend | Vanilla JS (no framework), hand-rolled DOM building, CSS custom properties for theming |
| Batch testing | Standalone Python scripts reusing the production pipeline (testVaseMatcher.py, tester.py) |
3. Feature Walkthrough
Presented in the order they were built, which is also roughly the order of increasing sophistication:
- Core matching (SIFT / ORB / SuperGlue / LoFTR) — four independent keypoint/dense matchers, each scored by RANSAC inlier count against every template, run strictly sequentially.
- Weighted final verdict + Borda overall-best — two different ways of combining the four raw scores into one winning template.
- Color space section — HS-histogram intersection, shown as its own
section but also folded into the weighted verdict (
Coloris a 5th contributor there only). - Background-removal size fix — uploads over 15MB are downscaled + re-encoded rather than rejected outright.
- Color family grid — dynamic K-means color-region discovery + side-by-side region-matching visualization, purely informational.
- Shape matching — Hu-moment contour distance + canonical-silhouette IoU, its own section, never scored.
- Texture matching — Local Binary Patterns + GLCM/Haralick features, its own section, never scored.
- AI verification — sends the matched template photo + the user's original upload to an external vision-LLM endpoint for a free-text QC-style comparison, parsed into a structured Match/Confidence/ Discrepancies/Description card.
- Flower counting v1 (SAM1) — Segment Anything (ViT-B) in automatic "segment everything" mode, filtered by size/position heuristics to guess which proposals were flowers. Superseded by step 11.
- YOLO-World cross-check + vase exclusion — open-vocabulary detection used first to strip the vase/ribbon out of SAM1's flower count, and kept afterward as an independent second opinion shown side by side.
- Flower counting v2 (SAM3) + tuned recall — SAM1 replaced with SAM3
concept-prompted segmentation (prompt:
"flower"), which never proposes the vase/ribbon in the first place — no more heuristic filtering needed. (An intermediate step also tuned SAM1's ownpoints_stride/conf_thres/stability_score_threshfor better recall before the SAM3 replacement landed; that tuning knowledge is preserved in this document for context even though the code path is gone.) - Vase-identity comparison (DINOv2 + CLIP) — crops the vase out of both photos using SAM3's precise mask (not just YOLO's bounding box), background-blacks-out everything else in the crop, and reports a Same/Uncertain/Different verdict from a weighted DINOv2+CLIP cosine similarity.
- Flower-count mismatch explanation — a plain-English note ("your photo has 23 flowers vs. 9 in the template") surfaced wherever counts differ, explaining a likely contributor to a lower match score.
tester.py— a lean batch-accuracy script that also emits one combined report image per test photo (input segmentation, template segmentation, counts, vase comparison) using the real production pipeline.- UI reorder — the "Weighted final match" section moved to the very top of the results, so the final verdict is visible without scrolling.
- Auto-triggered "Flower check" — a lightweight, CLIP-only (no DINOv2) flower-count + flower-similarity mini-panel shown beside the weighted card, fired automatically on every confident match (unlike the full flower-count feature, which stays behind its own button) — with a loading spinner while it computes.
4. Low-Level Design
4.1 Directory Structure
featureTransform/
├── app.py Flask routes, logging setup, request glue
├── config.py All tunables; env-var capping; HF_TOKEN loading
├── sam3_worker.py Standalone SAM3 subprocess entrypoint (sam2_env)
├── tester.py Lean batch-accuracy + composite report images
├── testVaseMatcher.py Fuller batch-accuracy report (CSV, per-method)
├── .env HF_TOKEN=hf_... (gitignored, gated model access)
│
├── pipeline/
│ ├── engine.py Orchestrator: bootstrap, per-request pipeline,
│ │ weighted/Borda aggregation, opt-in features
│ ├── bg_removal.py rembg wrapper, disk caching, mask creation
│ ├── classical.py SIFT / ORB extraction + FLANN matching
│ ├── deep.py SuperPoint+LightGlue, LoFTR
│ ├── color.py HS-histogram + dominant-color comparison
│ ├── color_grid.py Dynamic K-means color-family grid
│ ├── shape_match.py Hu moments + canonical-silhouette IoU
│ ├── texture_match.py LBP + GLCM/Haralick
│ ├── flower_count.py SAM3-backed per-instance flower counting
│ ├── yolo_world.py Open-vocabulary flower/vase/ribbon detection
│ ├── vase_compare.py DINOv2 + CLIP vase-identity comparison
│ ├── sam3_client.py Subprocess bridge INTO sam3_worker.py
│ ├── verify.py External AI-verification HTTP client
│ └── utils.py Template listing, score annotation, compression
│
├── templates/
│ └── index.html Single-page UI (Jinja for the template gallery)
│
├── static/
│ ├── css/style.css
│ └── js/main.js All client-side rendering logic
│
├── cache/
│ ├── templates_nobg/ Background-removed templates, cached by name
│ └── uploads_nobg/ Background-removed uploads, cached by content hash
│
├── uploads/<request_id>/ Per-request working files (originals, overlays,
│ crops, family grids) -- swept after 6 hours
├── logs/app.log Rotating log, all pipeline.* + werkzeug output
└── test_reports/ tester.py's composite report images
4.2 Module Responsibilities
config.py
The single source of truth for every tunable constant in the system, plus two pieces of process-startup plumbing that must run before any heavy import:
- Caps
OMP_NUM_THREADS/OPENBLAS_NUM_THREADS/MKL_NUM_THREADS/NUMEXPR_NUM_THREADS/ORT_NUM_THREADStomin(4, cpu_count)— done viaos.environ.setdefaultat import time, beforecv2/onnxruntime/torchare imported anywhere in the process. Left uncapped, each library grabs one thread per core, which starves everything else on a shared machine. - Sets
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True— the caching allocator's own recommended mitigation for fragmentation-driven CUDA OOM on a small card.
It also reads HF_TOKEN once (env var first, then a tiny hand-rolled
.env parser as fallback) so the SAM3 subprocess can be handed it
explicitly without needing a python-dotenv dependency.
pipeline/engine.py
The orchestrator. Owns:
bootstrap()— precomputes every template's background-removed image and every method's template-side features/embeddings once, at startup, so a live request only ever has to process the single uploaded image.process_upload()/_process_upload_locked()— the/api/matchcode path: background removal → sequential SIFT/ORB/SuperGlue/LoFTR → best-image annotation → color/shape/texture analysis → weighted + Borda aggregation → color family grid → response assembly._weighted_scores()/_overall_best()— the two aggregation algorithms (see §4.4).count_flowers()— the full opt-in flower-count + vase-comparison feature (SAM3 + YOLO-World + DINOv2/CLIP, renders visuals, returns everything for the "Count flowers" button).flower_summary()— the lightweight, auto-triggered companion (SAM3 + CLIP only, no YOLO-World, no DINOv2, no rendered images) shown beside the weighted card on every confident match.- A single
threading.Lock(_pipeline_lock) serializes every request's pipeline end-to-end, so at most one heavy CPU/GPU operation is ever in flight process-wide.
pipeline/bg_removal.py
Wraps rembg (BiRefNet-general-lite). Lazily creates an ONNX Runtime
session (CUDAExecutionProvider first, falls back to CPU on failure —
this fallback is "sticky" for the process's lifetime, a known limitation).
Disk-caches results: uploads by content hash, templates by filename.
resize_max_dim() is the single choke point enforcing MAX_IMAGE_DIM
everywhere. unload_session() drops the session so count_flowers()/
flower_summary() can free its GPU memory before the heavier optional
models load.
pipeline/classical.py
SIFT and ORB feature extraction + FLANN-based matching. Template
descriptors are precomputed once and cached in-memory (_template_features).
See §4.3 for the exact scoring formula.
pipeline/deep.py
SuperPoint+LightGlue (labeled "SuperGlue" in the UI, matching an existing
naming convention) and LoFTR. unload_models() drops all three model
objects (SuperPoint, LightGlue, LoFTR) so the opt-in flower/vase features
can reclaim their GPU memory; they lazily reload on the next /api/match
call, identical to a fresh process start.
pipeline/color.py
Independent, non-scored-looking-but-actually-scored (folded into the
weighted verdict via Color) comparison: Hue+Saturation 2D histogram
(Value/brightness deliberately excluded for lighting invariance),
compared via histogram intersection. Also computes dominant colors
(K-means in a bounded sample) and per-color Lab-distance similarity pairs
for the visual palette display.
pipeline/color_grid.py
Dynamic K-means color-region discovery in LAB space (lightness down-weighted so shadows of the same hue cluster together), rendered as a tile grid, region-matched by nearest average color, with an area-match-percentage per region and an overall area-weighted headline number. Purely visual — tied to whichever template the weighted verdict already picked.
pipeline/shape_match.py
Two complementary shape signals, averaged into one match_pct:
- Hu moments (
cv2.matchShapes,CONTOURS_MATCH_I1) — translation/ rotation/scale-invariant contour distance, converted to a similarity %. - Canonical-silhouette IoU — both masks cropped to their own bounding box, resized/centered into a fixed square canvas, then measured as direct pixel overlap. This also produces the side-by-side + overlay visualization images.
pipeline/texture_match.py
Two complementary texture signals, averaged into one match_pct:
- Local Binary Patterns (uniform method), compared as a histogram intersection — catches fine, repetitive patterns (fabric weave, petal grain).
- GLCM/Haralick features (contrast, homogeneity, energy, correlation), compared as a normalized relative difference, averaged across properties and across 4 angles/2 distances — catches coarser smooth-vs-rough material differences.
pipeline/flower_count.py
SAM3-backed per-instance flower counting. count_flowers(bgr, mask, workdir, instances_raw=...) accepts already-fetched SAM3 instances (so
engine.py can batch multiple images/prompts into one subprocess call
and hand each result set to this function separately), applies a light
foreground-overlap sanity filter, then color-clusters the survivors (LAB
K-means, capped at SAM_MAX_KIND_CLUSTERS) as a rough proxy for distinct
flower "kinds". union_mask() OR-combines every instance mask into one —
used to crop "just the flower material" for CLIP similarity.
render_instances() draws the numbered, color-coded overlay.
pipeline/yolo_world.py
Open-vocabulary detection (yolov8s-worldv2.pt) prompted with
["flower", "vase", "ribbon", "bow"]. Kept as an independent second
opinion shown side by side with SAM3's per-instance count — it draws
one box per contiguous flower region rather than per bloom, so it's a
coarser, corroborating signal, not a replacement.
pipeline/vase_compare.py
crop_mask(bgr, mask_bool)— crops the bounding box of any mask (padded byVASE_CROP_PAD_FRAC), blacking out every pixel the mask doesn't cover. Generic despite the module name — also used for the flower-only crop inflower_summary().compare_vases()— DINOv2 + CLIP cosine similarity, combined as a weighted average (VASE_DINO_WEIGHT=0.6,VASE_CLIP_WEIGHT=0.4), bucketed intosame/uncertain/different.clip_similarity_pct()— CLIP-only variant (never touches DINOv2), used by the lightweight auto-triggered flower check.unload_models()— drops both model objects; a no-op for whichever one was never loaded (e.g. DINOv2 stays unloaded entirely on the CLIP-only path).
pipeline/sam3_client.py
The subprocess bridge (runs in the main app's Python 3.8 process). Saves
each named input image to a temp PNG, writes a JSON request file
describing every (image, prompt, threshold) job, invokes
sam3_worker.py under sam2_env's Python interpreter with HF_TOKEN
passed through the subprocess environment, waits (bounded by
SAM3_TIMEOUT_SECONDS), and parses the JSON response + PNG masks back
into {(image_key, prompt): [{"mask": bool ndarray, "score": float, "box": [...]}]}.
Always cleans up its temp directory, even on failure.
sam3_worker.py
Standalone script, runs under sam2_env's Python 3.10, has no
dependency on this app's config/pipeline modules. Loads Sam3Model +
Sam3Processor 4-bit quantized (NF4 via BitsAndBytesConfig), processes
every job in the request against the already-loaded model (one model load
serves every job in a request), writes one PNG per detected instance plus
a response.json manifest.
pipeline/verify.py
Thin HTTP client for the external AI-verification endpoint. Posts both
images as multipart form data, parses the endpoint's free-text result
field with a tolerant regex (DISCREP\w* rather than a literal spelling,
since the LLM behind it doesn't always spell "DISCREPANCIES" consistently)
into {match, confidence, discrepancies, description}.
pipeline/utils.py
Template file listing, score-annotation overlay text, PNG encoding, and
compress_image_bytes() — downscales + re-encodes as JPEG only if an
upload exceeds COMPRESS_ABOVE_BYTES, reusing MAX_IMAGE_DIM since the
pipeline downsamples to that anyway.
4.3 Scoring Algorithms
SIFT / ORB (pipeline/classical.py)
1. Extract keypoints + descriptors (masked to foreground)
2. FLANN k=2 nearest-neighbor match against template descriptors
3. Lowe's ratio test: keep match if best.distance < 0.75 * secondBest.distance
4. RANSAC homography (5.0px threshold) across surviving matches
5. score = RANSAC inlier count (raw integer)
confidence = inliers / good_matches * 100 (%)
SuperGlue = SuperPoint + LightGlue (pipeline/deep.py)
Same shape as above, but SuperPoint (neural keypoints/descriptors) replaces SIFT/ORB's detector and LightGlue (learned matcher) replaces FLANN+ratio test. Matches filtered to inside the foreground mask before the same RANSAC-inlier-count scoring.
LoFTR (pipeline/deep.py)
Dense matcher — no keypoint detection step. Directly regresses pixel
correspondences, keeps only ones above LOFTR_CONFIDENCE_THRESHOLD (0.5),
filters to inside the mask, then the same RANSAC-inlier-count scoring.
Color space (pipeline/color.py)
hist = HS 2D histogram (Hue×Saturation only, Value excluded), normalized to sum=1
match_pct = Σ min(hist_input[i], hist_template[i]) × 100 (histogram intersection)
Shape matching (pipeline/shape_match.py)
hu_similarity = max(0, 100 × (1 - matchShapes(contour_a, contour_b) / 1.5))
iou = |mask_a ∩ mask_b| / |mask_a ∪ mask_b| × 100 (canonical-aligned)
match_pct = (hu_similarity + iou) / 2
Texture matching (pipeline/texture_match.py)
lbp_similarity = Σ min(hist_a[i], hist_b[i]) × 100 (LBP histogram intersection)
glcm_similarity = mean over 4 properties of max(0, 1 - |a-b|/max(|a|,|b|)) × 100
match_pct = (lbp_similarity + glcm_similarity) / 2
Vase identity (pipeline/vase_compare.py)
cosine_pct(a, b) = max(0, min(1, a·b)) × 100 (a, b are L2-normalized embeddings)
dino_pct = cosine_pct(DINOv2(crop_a), DINOv2(crop_b))
clip_pct = cosine_pct(CLIP(crop_a), CLIP(crop_b))
combined_pct = 0.6 × dino_pct + 0.4 × clip_pct
verdict = "same" if combined_pct >= 75
"uncertain" if combined_pct >= 60
"different" otherwise
Flower CLIP similarity (flower_summary, CLIP-only variant)
flower_clip_pct = cosine_pct(CLIP(union_of_flower_crops_a), CLIP(union_of_flower_crops_b))
4.4 Aggregation Layers
Two genuinely different "who won" computations exist side by side, and are allowed to disagree:
Weighted final match (engine._weighted_scores) — a literal weighted
sum of raw scores:
weighted_score(template) = Σ_method WEIGHT[method] × raw_score(method, template)
WEIGHT = { LoFTR: 0.75, SuperGlue: 0.15, SIFT: 0.10, ORB: 0.10, Color: 0.7 }
Because this sums raw scores (inlier counts, which run much higher for
LoFTR/SuperGlue than SIFT/ORB, plus Color's 0-100 percentage), a method
with naturally larger magnitudes pulls harder on the total even at a
similar or lower weight. That's the deliberate behavior of the formula as
specified, not a bug — see the comment block in config.py above
METHOD_WEIGHTS.
Overall best (engine._overall_best, Borda count) — rank-based, and
excludes Color entirely (only the original four methods participate):
for each method:
rank all templates by that method's own raw score
template at rank i (0-indexed) earns (n - i) points
overall_best = template with the highest total points across all methods
Rank-based aggregation means no single method's raw-score magnitude can dominate the way it can in the weighted sum — this is why the two verdicts are kept, and shown, separately rather than collapsed into one number.
4.5 Request Lifecycles
POST /api/match
Browser Flask (app.py) engine.py GPU
│ upload image │ │ │
├─────────────────────────► │ │
│ │ compress if >15MB │ │
│ ├──────────────────────────► │ │
│ │ │ bg_removal (rembg) │
│ │ ├────────────────────►│
│ │ │◄────────────────────┤
│ │ │ SIFT (sequential) │
│ │ ├────────────────────►│
│ │ │ ORB │
│ │ ├────────────────────►│
│ │ │ SuperGlue │
│ │ ├────────────────────►│
│ │ │ LoFTR │
│ │ ├────────────────────►│
│ │ │ Color / Shape / │
│ │ │ Texture (CPU) │
│ │ │ weighted + Borda │
│ │ │ color family grid │
│ │◄───────────────────────────┤ │
│◄───────────────────────── JSON result │ │
SuperPoint/LightGlue/LoFTR stay loaded in memory after this request (for fast repeated matching) — they are only unloaded when an opt-in flower/vase feature explicitly needs the GPU headroom.
POST /api/count_flowers (opt-in, "Count flowers" button)
Browser app.py engine.count_flowers() sam3_client → sam3_worker.py (sam2_env) YOLO-World DINOv2+CLIP
│ click │ │ │ │ │
├───────────────► │ │ │ │
│ ├────────────────────► │ │ │
│ │ │ unload SuperPoint/LightGlue/LoFTR + rembg session │ │
│ │ │ yolo_world.detect() ────────────────────────────────────────────►│ │
│ │ │ batch: input+flower, input+vase, [template+vase, template+flower] │ │
│ │ ├─────────────────────────────► │ │
│ │ │ (one SAM3 model load serves every job) │ │
│ │ │◄─────────────────────────────┤ │ │
│ │ │ cluster + render SAM instances (input & template) │ │
│ │ │ render YOLO-World boxes │ │
│ │ │ vase_compare (DINOv2+CLIP) ──────────────────────────────────────────────────►│
│ │ │◄─────────────────────────────────────────────────────────────────────────────┤
│ │ │ unload YOLO-World, DINOv2, CLIP │ │
│ │◄───────────────────┤ │ │
│◄─────────────── JSON + image URLs │ │ │
POST /api/flower_summary (auto-triggered, beside the weighted card)
Browser app.py engine.flower_summary() sam3_client/worker CLIP only
│ (fires automatically │ │ │
│ right after a confident match) │ │ │
├────────────────────────────────────► │ │
│ │ unload matching models │ │
│ │ batch: input+flower, │ │
│ │ template+flower │ │
│ ├─────────────────────────────►│ │
│ │◄─────────────────────────────┤ │
│ │ count comparison (no render)│ │
│ │ union masks, crop, CLIP-only similarity ────────────►│
│ │◄──────────────────────────────────────────────────────┤
│ │ unload CLIP (DINOv2 never loaded) │
│◄──────────────────────────────────── │
Deliberately skips YOLO-World, DINOv2, and any rendered overlay images — it exists purely to be fast enough to run unconditionally on every match.
POST /api/verify (AI verification, auto-fires after a confident match)
Browser app.py verify.py External vision-LLM endpoint
│ │ │ │
├──────────────► │ │
│ ├────────────────► POST multipart (image1, image2)
│ │ ├─────────────────────────►│
│ │ │◄─────────────────────────┤
│ │ │ parse MATCH/CONFIDENCE/ │
│ │ │ DISCREPANCIES/description│
│ │◄──────────────── │
│◄────────────── JSON (or 502 if endpoint unreachable) │
Failures here (stale Cloudflare tunnel, timeout) are surfaced as a soft error in the UI, never a hard failure of the page.
4.6 SAM3 Cross-Process Bridge
Request (JSON file, written by sam3_client.py, read by sam3_worker.py):
{
"images": {
"input": "/path/to/uploads/<request_id>/sam3_xxxx/input.png",
"template": "/path/to/uploads/<request_id>/sam3_xxxx/template.png"
},
"jobs": [
{"image": "input", "prompt": "flower", "threshold": 0.5},
{"image": "input", "prompt": "vase", "threshold": 0.3},
{"image": "template", "prompt": "vase", "threshold": 0.3}
],
"output_dir": "/path/to/uploads/<request_id>/sam3_xxxx/out"
}
Response (response.json, written by the worker):
{
"results": [
{"image": "input", "prompt": "flower", "instances": [
{"mask_file": "input_flower_0_0.png", "score": 0.7486, "box": [x1, y1, x2, y2]},
{"mask_file": "input_flower_0_1.png", "score": 0.7006, "box": [...]}
]},
{"image": "input", "prompt": "vase", "instances": [
{"mask_file": "input_vase_1_0.png", "score": 0.9391, "box": [...]}
]}
],
"error": null
}
Each mask_file is a standalone grayscale PNG (0/255) at the original
image's resolution — sam3_client.py reads it back with
cv2.imread(..., IMREAD_GRAYSCALE) > 127 to reconstruct a boolean mask.
On any failure, error is a string (worker traceback included) and the
process exits non-zero; sam3_client.py raises Sam3Error either way,
which every caller treats as a soft failure.
Why 4-bit (NF4) quantization: the unquantized checkpoint is ~3.2GB
(model.safetensors). Loading it naively (8-bit, device_map="cuda:0")
hit a CUDA OOM during Transformers' own memory pre-flight ("caching
allocator warmup") step on this 8GB card, because that warmup reserves
capacity based on the original dtype size before quantization actually
reduces it. Switching to load_in_4bit=True + bnb_4bit_quant_type="nf4"
low_cpu_mem_usage=Truefixed this: ~700MB resident, ~1.9GB peak during inference, ~6s to load, ~2s per prompt.
Why sam.predictor(...) isn't called directly for automatic-mode
tuning (historical note, from the SAM1 era): Ultralytics' high-level
model.predict()/model() call validates every kwarg against a fixed
CLI-style config schema, which rejects automatic-mode-only parameters like
points_stride/conf_thres with a SyntaxError. The fix at the time was
calling model.predictor(...) directly (after a one-time warmup call to
force predictor construction), bypassing that validation entirely since
BasePredictor.__call__ forwards arbitrary kwargs straight through. This
code path no longer exists (SAM1 was fully replaced by SAM3), but the
technique is documented here in case a similar situation recurs with
another Ultralytics-wrapped model (e.g. YOLO-World).
4.7 GPU Memory Management Strategy
┌───────────────────────────────────────────┐
│ Idle / after /api/match │
│ SuperPoint + LightGlue + LoFTR RESIDENT │
│ (~200-900MB, kept warm for fast matching) │
└──────────────────┬────────────────────────┘
│ user clicks "Count flowers"
│ or a confident match fires
│ the auto flower-check
▼
┌───────────────────────────────────────────┐
│ deep.unload_models() │
│ bg_removal.unload_session() │
│ torch.cuda.empty_cache() │
│ → GPU now near-baseline (~200-900MB) │
└──────────────────┬────────────────────────┘
▼
┌───────────────────────────────────────────┐
│ SAM3 (separate PROCESS, sam2_env) │
│ ~700MB resident / ~1.9GB peak │
│ + optionally YOLO-World (~200MB) │
│ + optionally DINOv2 (~350MB) + CLIP (~600MB) │
└──────────────────┬────────────────────────┘
│ request finishes
▼
┌───────────────────────────────────────────┐
│ yolo_world.unload_model() │
│ vase_compare.unload_models() │
│ (SAM3's own process has already exited — │
│ its GPU memory is freed by the OS/driver, │
│ nothing to unload on the main-process side) │
│ → back to near-baseline │
└───────────────────────────────────────────┘
Every model getter (deep.get_superpoint(), flower_count's SAM3 client,
vase_compare._get_dino()/_get_clip(), yolo_world.get_model()) is a
lazy singleton — None until first use, reset to None by the matching
unload_*() function. This means "unload" is always safe to call even if
the model was never loaded (checked via is not None), and every model
transparently reloads itself on the next call that needs it, with no
special-casing required anywhere else in the codebase.
4.8 Frontend Structure (static/js/main.js)
Single IIFE, no framework, no build step. Key structure:
- Upload flow: drag/drop or file picker →
fetch POST /api/match→renderResults(data). renderResults(data)populates every section in DOM order (weighted card first, then overall-best, query images, method grid, flower count, color/shape/texture/family-grid sections), then fires two auto-triggered, non-blocking follow-ups:runVerification()(AI check) andrunFlowerSummary()(the lightweight flower-check panel) — neither blocks the initial render, both show their own loading state and fail independently.renderWeighted(data)builds the terracotta-gradient weighted card (input/template thumbnails, breakdown table, Strong/Moderate/Weak confidence pill derived purely from the margin between 1st and 2nd place — display-only, cannot feed back into the algorithm).- Opt-in flower count:
countFlowersBtnclick handler →fetch POST /api/count_flowers→renderFlowerCount(data), which in turn callsrenderFlowerCountMismatch()andrenderVaseComparison(). - Every async section follows the same pattern: hide body / show spinner → fetch → on success show body, on failure show a scoped error message — never a page-level failure.
5. API Reference
| Method | Path | Body | Returns |
|---|---|---|---|
| GET | / |
— | Rendered HTML page (template gallery + upload form) |
| GET | /template_image/<filename> |
— | Raw template image bytes |
| GET | /uploads/<request_id>/<filename> |
— | Any file written into that request's working directory |
| POST | /api/match |
multipart image file |
Full match result: per-method scores, weighted verdict, Borda overall-best, color/shape/texture/family-grid analysis |
| POST | /api/count_flowers |
JSON {request_id, template} |
SAM3 + YOLO-World flower counts (with visuals), vase comparison, flower-count-mismatch note |
| POST | /api/flower_summary |
JSON {request_id, template} |
Lightweight flower-count comparison + CLIP-only flower similarity (no images) |
| POST | /api/verify |
JSON {request_id, template} |
External AI vision-LLM verdict: match/confidence/discrepancies/description |
All POST endpoints return {"error": "..."} with a non-200 status on
failure; the frontend treats every one of these as scoped/local, never
page-fatal.
6. Configuration Reference
All of the following live in config.py, grouped by feature:
| Setting | Default | Purpose |
|---|---|---|
MAX_IMAGE_DIM |
1600 | Every image downscaled to this before any model touches it |
METHOD_WEIGHTS |
{LoFTR: .75, SuperGlue: .15, SIFT: .10, ORB: .10, Color: .7} |
Weighted-sum verdict weights |
SCORE_THRESHOLD |
per-method | "Confident match" badge cutoff (doesn't affect ranking) |
MAX_CONTENT_LENGTH_BYTES |
50MB | Hard reject ceiling |
COMPRESS_ABOVE_BYTES |
15MB | Soft threshold — downscale+recompress instead of reject |
SHAPE_HU_DISTANCE_SCALE |
1.5 | Divisor turning Hu-moment distance into a 0-100% similarity |
TEXTURE_LBP_RADIUS / TEXTURE_GLCM_LEVELS |
2 / 32 | Texture descriptor parameters |
FAMILY_GRID_K |
5 | Number of color families discovered per image |
SAM3_FLOWER_THRESHOLD / SAM3_VASE_THRESHOLD |
0.5 / 0.3 | Confidence cutoff per SAM3 concept prompt |
SAM3_TIMEOUT_SECONDS |
180 | Subprocess timeout for the SAM3 worker |
VASE_DINO_WEIGHT / VASE_CLIP_WEIGHT |
0.6 / 0.4 | Vase-comparison combined-score weights |
VASE_SAME_THRESHOLD / VASE_UNCERTAIN_THRESHOLD |
75 / 60 | Vase verdict cutoffs (untuned — no ground-truth calibration set exists yet) |
YOLO_WORLD_CLASSES |
["flower", "vase", "ribbon", "bow"] |
Open-vocabulary prompt classes |
VERIFY_ENDPOINT_URL |
Cloudflare tunnel URL | Can go stale if the tunnel restarts — treated as a soft failure |
7. Environment & Setup
Two conda environments
| Env | Python | Purpose | Key packages |
|---|---|---|---|
torch17_new |
3.8 | Main Flask app | torch, opencv-python, kornia, lightglue, rembg, ultralytics, transformers==4.46.3 |
sam2_env |
3.10 | SAM3 subprocess only | torch, transformers>=5.5.0, bitsandbytes, accelerate |
Model checkpoints / weights
| Model | Source | Notes |
|---|---|---|
| SIFT / ORB | Built into OpenCV | No download needed |
| SuperPoint + LightGlue | lightglue package |
Auto-downloaded on first use |
| LoFTR | kornia.feature |
Auto-downloaded on first use |
| rembg (BiRefNet-general-lite) | rembg package |
Auto-downloaded on first use |
| YOLO-World | yolov8s-worldv2.pt |
Downloaded via ultralytics, cached locally |
| DINOv2-base | facebook/dinov2-base |
Hugging Face, public, no token needed |
| CLIP ViT-B/32 | openai/clip-vit-base-patch32 |
Hugging Face, public, no token needed |
| SAM3 | facebook/sam3 |
Gated on Hugging Face — requires requesting access on the model page and an approved access token |
.env file
HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
Required only for the SAM3-backed features (flower counting, vase comparison, the auto-triggered flower check). Without it, those specific features fail softly with a clear "SAM3 unavailable" message — core matching is entirely unaffected.
Running the app
cd /media/suman/Backup_of_extra_/Sasi/featureTransform
conda activate torch17_new
python app.py
# serves on http://0.0.0.0:5053
8. Testing Tools
Two batch-test scripts exist, covering different needs:
testVaseMatcher.py |
tester.py |
|
|---|---|---|
| Output | Detailed CSV + per-method accuracy table | Console pass/fail + one composite report image per photo |
| Coverage | All 4 methods + weighted + Borda + color pick, individually tallied | Weighted verdict only |
| Flower/vase check | Numbers only (family-grid area-match %), no images | Full: SAM3 flower segmentation (input vs. template), counts, vase comparison — all rendered into one PNG |
| Speed | Faster (no SAM3/CLIP round trip) | Slower (~15-25s/image extra for the flower/vase step); --skip-flowers for a fast pass |
| Use case | Regression-testing every individual method's accuracy | Visual QA — "does this match make sense at a glance" |
Both expect the same input layout:
<test_dir>/
SKU_1/ photo1.jpg photo2.jpg ...
SKU_2/ ...
where each subfolder name must exactly match an existing template name.
Common mistake: pointing either script directly at a SKU_X/ folder
instead of its parent — the script silently finds 0 images, since it
expects to listdir() a folder of subfolders, not a folder of images.
python tester.py # ./flowers, full report images
python tester.py /path/to/flowers --limit 3
python tester.py --skip-flowers # fast accuracy-only pass
python testVaseMatcher.py /path/to/flowers --csv report.csv
9. Known Limitations & Design Tradeoffs
- rembg's CUDA fallback is sticky for the process's lifetime. A transient CUDA OOM during session creation permanently falls back that process to CPU-only background removal (much slower) until restarted. No auto-retry/self-healing exists yet.
- Vase-identity and flower-count-mismatch thresholds are untuned.
There is no ground-truth "same/different vase" or "acceptable flower
count delta" labeled dataset — the cutoffs in
config.pyare reasonable defaults, not calibrated ones. - The weighted verdict's literal-sum formula means raw-score magnitude matters, not just weight. LoFTR/SuperGlue's inlier counts run far higher than SIFT/ORB's, so they can dominate the total even when their configured weight isn't proportionally larger. This is documented as intentional (matches the formula as originally specified), not a bug.
- YOLO-World's own "flower" count is coarse — one box per contiguous flower region, not per bloom — so it will never numerically match SAM3's per-instance count; it's shown as an independent corroborating signal, not something expected to agree.
- The AI-verification endpoint is a Cloudflare tunnel URL that can go stale if restarted on the other end; failures are always soft (shown as "AI verification unavailable"), never block the page.
- SAM3 requires a gated, manually-approved Hugging Face token. The app
cannot obtain this automatically — a human must request access at
huggingface.co/facebook/sam3and provide a token with access granted. - Everything assumes a single physical GPU with ~8GB VRAM and Pascal or later compute capability. The sequential-execution and unload/reload discipline exists specifically because of that constraint — a larger card would allow (but doesn't require) more concurrency.
10. Glossary
| Term | Meaning |
|---|---|
| Inlier count | Number of matched keypoints that agree with one consistent RANSAC-fit geometric transform between two images — the raw score for SIFT/ORB/SuperGlue/LoFTR |
| Weighted verdict | The primary displayed "final answer" — literal weighted sum of 5 methods' raw scores |
| Borda count / Overall best | A secondary, rank-based cross-check aggregate of 4 methods (excludes Color) |
| Concept segmentation | SAM3's mode of operation: prompted with a plain-English word ("flower"), returns instance masks for exactly that concept, rather than segmenting everything indiscriminately |
| NF4 | "NormalFloat4" — the 4-bit quantization scheme used (via bitsandbytes) to fit SAM3 in this GPU's memory budget |
| Flower check | The lightweight, auto-triggered (not opt-in), CLIP-only flower-count + flower-similarity panel shown beside the weighted card |
| Count flowers | The full, opt-in, button-triggered feature: SAM3 + YOLO-World flower counts, rendered overlays, and DINOv2+CLIP vase comparison |