基于Bounding Box将DataFrame词汇归类至对应列的技术问询
Solution to Group Words Under Corresponding Greeting Columns
Here's a step-by-step approach to generate your desired DataFrame by grouping words under their respective greeting columns based on bounding box coordinates:
Step 1: Split Data into Greetings and Content Words
First, we separate the greeting entries (starting with Greeting_) from the regular text words.
Step 2: Assign Words to Greetings
We assign each content word to the greeting that:
- Is positioned above the word (using the
topcoordinate—since highertopvalues mean lower on the image, we check if the word'stopis greater than the greeting'stop). - Has the closest horizontal position (using the
leftcoordinate to match columns).
Step 3: Sort and Format the Result
We sort each group of words by their vertical position (from top to bottom) and reshape into the final column-based DataFrame.
Full Code Implementation
import pandas as pd # Original DataFrame df_text = pd.DataFrame({ 'words': ['Hello', 'world', 'nice day', 'have a', 'Greeting_1', 'Greeting_2'], 'left': [1097, 1099, 258, 259, 1096, 260], 'top': [1248, 1249, 1156, 1153,1200,250], 'right': [1154, 1156, 615, 614, 1150, 610], 'bottom': [1269, 1271, 1175, 1172, 1255, 1170] }) # Separate greetings and content words greetings = df_text[df_text['words'].str.startswith('Greeting_')].copy() content = df_text[~df_text['words'].str.startswith('Greeting_')].copy() # Store greeting coordinates for reference greeting_info = greetings.set_index('words')[['left', 'top']].to_dict('index') # Function to map each word to its corresponding greeting def assign_greeting(row): word_left = row['left'] word_top = row['top'] valid_candidates = [] for greeting, coords in greeting_info.items(): # Only consider greetings positioned above the word if word_top > coords['top']: # Calculate horizontal distance to the greeting distance = abs(word_left - coords['left']) valid_candidates.append((greeting, distance)) if not valid_candidates: return None # No valid greeting found for this word # Pick the greeting with the closest horizontal position return min(valid_candidates, key=lambda x: x[1])[0] # Assign greetings to content words content['greeting'] = content.apply(assign_greeting, axis=1) # Remove words without a valid greeting assignment content = content.dropna(subset=['greeting']) # Sort words in each column by their vertical position (top to bottom) sorted_groups = content.groupby('greeting').apply( lambda group: group.sort_values('top')['words'].reset_index(drop=True) ) # Convert to the final column-based DataFrame result_df = sorted_groups.unstack().reset_index(drop=True) print(result_df)
Output
Greeting_1 Greeting_2 0 Hello have a 1 world nice day
Key Notes
- Horizontal Matching: Uses the
leftcoordinate to cluster words into columns matching the greetings. - Vertical Filtering: Ensures words are only assigned to greetings that appear above them (using the
topcoordinate). - Edge Cases: Handles words with no valid greeting assignment by dropping them, and fills missing values with
NaNif groups have unequal lengths.
内容的提问来源于stack exchange,提问作者Azee.
相关产品推荐
相关产品推荐

