Python Pandas如何根据索引区间从文本列提取内容生成新列
pandas按指定起止索引截取文本生成新列
两种可直接运行的实现方式:
- apply写法(逻辑直观易读,适配大多数场景)
import pandas as pd # 测试数据 data = {'text': ['They say that all cats land on their feet, but this does not apply to my cat. He not only often falls, but also jumps badly. We have visited the veterinarian more than once with dislocated paws and damaged internal organs.', 'Mom called dad, and when he came home, he took moms car and drove to the store'], 'begin_end':[[128, 139],[20,31]]} df = pd.DataFrame(data) # 核心截取逻辑:按规则取起始到结束索引+1的切片 df['new_col'] = df.apply(lambda row: row['text'][row['begin_end'][0]: row['begin_end'][1]+1], axis=1)
- 列表推导式写法(执行速度更快,适合大数据量场景)
df['new_col'] = [txt[s:e+1] for txt, (s,e) in zip(df['text'], df['begin_end'])]
执行后得到的结果完全匹配预期:
begin_end new_col 0 [128, 139] have visited 1 [20, 31] when he came
内容的提问来源于stack exchange,提问作者Rory
相关产品推荐
相关产品推荐

