You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Bounding Box将DataFrame词汇归类至对应列的技术问询

Solution to Group Words Under Corresponding Greeting Columns

Here's a step-by-step approach to generate your desired DataFrame by grouping words under their respective greeting columns based on bounding box coordinates:

Step 1: Split Data into Greetings and Content Words

First, we separate the greeting entries (starting with Greeting_) from the regular text words.

Step 2: Assign Words to Greetings

We assign each content word to the greeting that:

  1. Is positioned above the word (using the top coordinate—since higher top values mean lower on the image, we check if the word's top is greater than the greeting's top).
  2. Has the closest horizontal position (using the left coordinate to match columns).

Step 3: Sort and Format the Result

We sort each group of words by their vertical position (from top to bottom) and reshape into the final column-based DataFrame.

Full Code Implementation

import pandas as pd

# Original DataFrame
df_text = pd.DataFrame({
 'words': ['Hello', 'world', 'nice day', 'have a', 'Greeting_1', 'Greeting_2'],
 'left': [1097, 1099, 258, 259, 1096, 260],
 'top': [1248, 1249, 1156, 1153,1200,250],
 'right': [1154, 1156, 615, 614, 1150, 610],
 'bottom': [1269, 1271, 1175, 1172, 1255, 1170]
})

# Separate greetings and content words
greetings = df_text[df_text['words'].str.startswith('Greeting_')].copy()
content = df_text[~df_text['words'].str.startswith('Greeting_')].copy()

# Store greeting coordinates for reference
greeting_info = greetings.set_index('words')[['left', 'top']].to_dict('index')

# Function to map each word to its corresponding greeting
def assign_greeting(row):
    word_left = row['left']
    word_top = row['top']
    valid_candidates = []
    
    for greeting, coords in greeting_info.items():
        # Only consider greetings positioned above the word
        if word_top > coords['top']:
            # Calculate horizontal distance to the greeting
            distance = abs(word_left - coords['left'])
            valid_candidates.append((greeting, distance))
    
    if not valid_candidates:
        return None  # No valid greeting found for this word
    # Pick the greeting with the closest horizontal position
    return min(valid_candidates, key=lambda x: x[1])[0]

# Assign greetings to content words
content['greeting'] = content.apply(assign_greeting, axis=1)

# Remove words without a valid greeting assignment
content = content.dropna(subset=['greeting'])

# Sort words in each column by their vertical position (top to bottom)
sorted_groups = content.groupby('greeting').apply(
    lambda group: group.sort_values('top')['words'].reset_index(drop=True)
)

# Convert to the final column-based DataFrame
result_df = sorted_groups.unstack().reset_index(drop=True)

print(result_df)

Output

Greeting_1 Greeting_2
0      Hello      have a
1      world    nice day

Key Notes

  • Horizontal Matching: Uses the left coordinate to cluster words into columns matching the greetings.
  • Vertical Filtering: Ensures words are only assigned to greetings that appear above them (using the top coordinate).
  • Edge Cases: Handles words with no valid greeting assignment by dropping them, and fills missing values with NaN if groups have unequal lengths.

内容的提问来源于stack exchange,提问作者Azee.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 17:32:26