如何使用Pandas截断句子左右部分?求DataFrame实现方案
在Pandas DataFrame中实现句子截断功能
没问题,我来帮你把这个单句子的截断逻辑迁移到Pandas DataFrame里,核心思路就是把你写的逻辑封装成自定义函数,再批量应用到DataFrame的列上就行,一步步来:
1. 封装通用的截断函数
先把你原来的逻辑改成更灵活的函数,还能处理根词在句子开头/结尾的边界情况(避免索引越界):
def truncate_sentence(sentence, target_root, before_words=4, after_words=4): try: # 将句子拆分为单词列表 word_list = sentence.split() # 找到目标根词的索引 root_idx = word_list.index(target_root) # 计算截取的起始/结束位置,自动处理边界(比如根词在句首时,起始位置设为0) start_pos = max(0, root_idx - before_words) # 切片是左闭右开,所以结束位置要+1才能包含after_words个后续单词 end_pos = root_idx + after_words + 1 # 拼接成截断后的句子 return ' '.join(word_list[start_pos:end_pos]) except ValueError: # 找不到根词时返回None return None
2. 批量应用到DataFrame列
假设你的DataFrame有一列存储待处理的句子(比如叫sentences),用apply方法就能把函数批量作用到每一行:
import pandas as pd # 创建示例DataFrame sample_data = { 'sentences': [ "lack of association between the promoter polymorphism of the mtnr1a gene and adolescent idiopathic scoliosis", "this is a test sentence without your target word", "mtnr1a sits right at the start of this sentence", "here's a sentence where mtnr1a is near the very end" ] } df = pd.DataFrame(sample_data) # 应用函数生成新的截断列 df['truncated_sentence'] = df['sentences'].apply( lambda x: truncate_sentence(x, target_root='mtnr1a') ) # 查看结果 print(df)
运行结果
你会得到这样的输出:
sentences truncated_sentence 0 lack of association between the promoter polymorphism of the mtnr1a gene and adolescent idiopathic scoliosis promoter polymorphism of the mtnr1a gene and adolescent idiopathic 1 this is a test sentence without your target word None 2 mtnr1a sits right at the start of this sentence mtnr1a sits right at the start of this 3 here's a sentence where mtnr1a is near the very end where mtnr1a is near the very end
额外说明
- 如果根词在句子中可能出现多次,
word_list.index()只会返回第一个出现的位置;若要处理所有出现的情况,可以改成遍历enumerate(word_list)找到所有索引,再分别截断拼接。 - 你可以调整
before_words和after_words参数,自由控制截断时保留根词前后的单词数量。
内容的提问来源于stack exchange,提问作者jurek
相关产品推荐
相关产品推荐

