如何高效遍历按行存单词的大型Pandas DataFrame并提取句子?
高效提取Pandas DataFrame中的完整句子
针对大型DataFrame,不要手动遍历每行,用Pandas的矢量化操作和分组功能可以高效完成句子提取,步骤如下:
1. 标记句子边界并生成分组ID
句子的结束标志是col2列不为空的行,我们可以通过这个特征给每个完整句子分配唯一ID:
import pandas as pd # 示例数据 d = {'col1': ['This', 'is', 'a', 'simple', 'sentence', 'This', 'is', 'another', 'sentence', 'This', 'is', 'the', 'third', 'sentence', 'Is', 'this', 'a', 'sentence', 'too'], 'col2': ['', '', '', '', '!', '', '', '', '.', '', '', '', '', '...', '', '', '', '', '?']} df = pd.DataFrame(data=d) # 标记句子结尾行 df['is_sentence_end'] = df['col2'] != '' # 生成句子分组ID:每个完整句子的所有行将拥有相同ID df['sentence_id'] = df['is_sentence_end'].shift(fill_value=True).cumsum()
2. 提取单个句子
根据sentence_id筛选即可得到对应句子的DataFrame片段:
# 提取第一个句子(对应原数据0-4行) first_sentence = df[df['sentence_id'] == 1] print(first_sentence) # 提取第四个句子(对应原数据14-18行) fourth_sentence = df[df['sentence_id'] == 4] print(fourth_sentence)
3. 批量提取所有句子
如果需要一次性获取所有句子,可以用groupby直接分组:
# 将所有句子存入列表,每个元素是一个句子的DataFrame all_sentences = [group for _, group in df.groupby('sentence_id')] # 访问第二个句子 print(all_sentences[1])
效率说明
这种方法全程使用Pandas内置的矢量化运算和分组机制,避免了低效的逐行遍历,在处理大型数据集时性能优势明显。
内容的提问来源于stack exchange,提问作者striatum
相关产品推荐
相关产品推荐

