You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何查找图像中各视觉词汇的相对位置?(已构建20维视觉词字典)

Hey there! Let's break down how to find the relative positions of each visual word in your image, since you already have your 20-word visual dictionary and frequency array ready.

The frequency array is an aggregated statistic—it doesn't hold any location information on its own. So you'll need to go back to the raw feature point data generated when you mapped image features to your visual dictionary. Here's how to do it step by step:

1. Retrace & Save Raw Feature Records

When you extract local features (like SIFT, ORB, or SURF) from your target image and match each feature to a visual word in your dictionary, don't just count occurrences. Instead, save two key pieces of data for every feature point:

  • The visual word ID (the index of the matched word in your 20-word dictionary, e.g., 0 to 19)
  • The pixel coordinates (x, y) of the feature point in the image

Here's a Python-style pseudocode example to illustrate:

# Assume you've already extracted feature descriptors and keypoints from the image
# And your pre-trained visual dictionary is a k-means model with 20 clusters
visual_word_ids = kmeans.predict(descriptors)  # Each feature maps to a word ID (0-19)

# Initialize a dictionary to store positions per visual word
word_positions = {word_id: [] for word_id in range(20)}

# Map each feature's position to its corresponding visual word
for idx, word_id in enumerate(visual_word_ids):
    x_coord = keypoints[idx].pt[0]
    y_coord = keypoints[idx].pt[1]
    word_positions[word_id].append((x_coord, y_coord))

Now word_positions contains all pixel locations where each visual word appears in the image.

2. Calculate Relative Positions (Customize Based on Your Needs)

If you specifically need relative positions (not raw pixel coordinates), you can compute them using the raw positions:

  • Relative to image center: First calculate the image's center (center_x, center_y), then compute the offset of each position from this center.
  • Relative to other visual words: For example, calculate the average distance between all positions of word A and word B, or pairwise relative coordinates.

Example of calculating positions relative to the image center:

img_height, img_width = image.shape[:2]
center_x = img_width / 2
center_y = img_height / 2

# Compute relative positions for each visual word
relative_word_positions = {}
for word_id, positions in word_positions.items():
    relative_coords = [(x - center_x, y - center_y) for (x, y) in positions]
    relative_word_positions[word_id] = relative_coords

3. What If You Lost the Raw Feature Data?

If you didn't save the original feature point positions earlier, your only option is to re-run the feature extraction and visual word matching pipeline on the target image—this time making sure to record each feature's position alongside its matched visual word. There's no way to recover position data from the frequency array alone.

Quick Notes
  • A single visual word can correspond to multiple feature points, so each word will have a list of positions, not a single location.
  • If you want to summarize "where" a visual word lives in the image (e.g., which quadrant it's mostly in), you can run statistics on its position list—like taking the average coordinate, or clustering the positions to find a core region.

内容的提问来源于stack exchange,提问作者S.EB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:45:07