You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas DataFrame中查找行号相同的所有单词

Solution to Find Words Sharing the Same Line Number per Page

Got it, let's break this down. You want to identify all words that appear on the same line (same LineNum) within the same page—since line numbers repeat across pages, we need to pair Page and LineNum to target unique lines in the document. Here's how to do it step by step with pandas:

Step 1: Recreate the Sample DataFrame (for context)

First, let's make sure we're working with the same data you provided:

import pandas as pd

data = {
    'Idx': [0,1,2,4,5,6,7,8,9,10,11,12,13,14,15,16],
    'Page': [1,1,1,1,2,2,2,2,2,3,3,3,4,4,4,4],
    'Word': ['Hello','This','is','an','example','of','words','across','multiple','pages','in','the','document','which','has','split'],
    'LineNum': [1,1,2,2,1,1,1,2,2,1,1,1,1,1,1,1]
}
df = pd.DataFrame(data)

Step 2: Group Words by Page + Line Number

We'll group the DataFrame by both Page and LineNum (to get unique lines) and collect all words in each group:

# Group by page and line number, aggregate words into a list
line_word_groups = df.groupby(['Page', 'LineNum'])['Word'].agg(list).reset_index()

Step 3: Filter for Lines with Multiple Words

If you only care about lines that have more than one word (since single-word lines don't have shared line numbers), filter the groups:

# Keep only groups with 2+ words
multi_word_lines = line_word_groups[line_word_groups['Word'].apply(len) > 1]

Step 4: View the Results

Printing multi_word_lines will show you each unique line and all words on it:

print(multi_word_lines)

Output:

Page  LineNum          Word
0     1        1  [Hello, This]
1     1        2        [is, an]
2     2        1  [example, of, words]
3     2        2  [across, multiple]
4     3        1    [pages, in, the]
5     4        1  [document, which, has, split]

Bonus: Map Each Word to Its Shared Line

If you want to see every word paired with all other words on its line, merge the grouped data back to the original DataFrame:

result = df.merge(multi_word_lines, on=['Page', 'LineNum'], suffixes=('', '_group'))
print(result[['Page', 'LineNum', 'Word', 'Word_group']])

内容的提问来源于stack exchange,提问作者Venkatesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:15:16