在Pandas DataFrame中查找行号相同的所有单词
Got it, let's break this down. You want to identify all words that appear on the same line (same LineNum) within the same page—since line numbers repeat across pages, we need to pair Page and LineNum to target unique lines in the document. Here's how to do it step by step with pandas:
Step 1: Recreate the Sample DataFrame (for context)
First, let's make sure we're working with the same data you provided:
import pandas as pd data = { 'Idx': [0,1,2,4,5,6,7,8,9,10,11,12,13,14,15,16], 'Page': [1,1,1,1,2,2,2,2,2,3,3,3,4,4,4,4], 'Word': ['Hello','This','is','an','example','of','words','across','multiple','pages','in','the','document','which','has','split'], 'LineNum': [1,1,2,2,1,1,1,2,2,1,1,1,1,1,1,1] } df = pd.DataFrame(data)
Step 2: Group Words by Page + Line Number
We'll group the DataFrame by both Page and LineNum (to get unique lines) and collect all words in each group:
# Group by page and line number, aggregate words into a list line_word_groups = df.groupby(['Page', 'LineNum'])['Word'].agg(list).reset_index()
Step 3: Filter for Lines with Multiple Words
If you only care about lines that have more than one word (since single-word lines don't have shared line numbers), filter the groups:
# Keep only groups with 2+ words multi_word_lines = line_word_groups[line_word_groups['Word'].apply(len) > 1]
Step 4: View the Results
Printing multi_word_lines will show you each unique line and all words on it:
print(multi_word_lines)
Output:
Page LineNum Word 0 1 1 [Hello, This] 1 1 2 [is, an] 2 2 1 [example, of, words] 3 2 2 [across, multiple] 4 3 1 [pages, in, the] 5 4 1 [document, which, has, split]
Bonus: Map Each Word to Its Shared Line
If you want to see every word paired with all other words on its line, merge the grouped data back to the original DataFrame:
result = df.merge(multi_word_lines, on=['Page', 'LineNum'], suffixes=('', '_group')) print(result[['Page', 'LineNum', 'Word', 'Word_group']])
内容的提问来源于stack exchange,提问作者Venkatesh

