You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效遍历按行存单词的大型Pandas DataFrame并提取句子?

高效提取Pandas DataFrame中的完整句子

针对大型DataFrame,不要手动遍历每行,用Pandas的矢量化操作和分组功能可以高效完成句子提取,步骤如下:

1. 标记句子边界并生成分组ID

句子的结束标志是col2列不为空的行,我们可以通过这个特征给每个完整句子分配唯一ID:

import pandas as pd

# 示例数据
d = {'col1': ['This', 'is', 'a', 'simple', 'sentence',
              'This', 'is', 'another', 'sentence',
              'This', 'is', 'the', 'third', 'sentence', 
              'Is', 'this', 'a', 'sentence', 'too'],
     'col2': ['', '', '', '', '!',
              '', '', '', '.',
              '', '', '', '', '...',
              '', '', '', '', '?']}
df = pd.DataFrame(data=d)

# 标记句子结尾行
df['is_sentence_end'] = df['col2'] != ''
# 生成句子分组ID:每个完整句子的所有行将拥有相同ID
df['sentence_id'] = df['is_sentence_end'].shift(fill_value=True).cumsum()

2. 提取单个句子

根据sentence_id筛选即可得到对应句子的DataFrame片段:

# 提取第一个句子(对应原数据0-4行)
first_sentence = df[df['sentence_id'] == 1]
print(first_sentence)

# 提取第四个句子(对应原数据14-18行)
fourth_sentence = df[df['sentence_id'] == 4]
print(fourth_sentence)

3. 批量提取所有句子

如果需要一次性获取所有句子,可以用groupby直接分组:

# 将所有句子存入列表,每个元素是一个句子的DataFrame
all_sentences = [group for _, group in df.groupby('sentence_id')]

# 访问第二个句子
print(all_sentences[1])

效率说明

这种方法全程使用Pandas内置的矢量化运算和分组机制,避免了低效的逐行遍历,在处理大型数据集时性能优势明显。

内容的提问来源于stack exchange,提问作者striatum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 02:50:09