Pandas新增列提取其他列字符串指定位置单词的实现方法
Pandas 逐行计算生成新列实现方案
你现有代码存在两个核心问题,修正后即可实现需求:
- 自定义
word_at_pos函数仅通过print输出结果,没有返回值,无法将计算结果赋值给DataFrame列;同时原函数遍历逻辑中break位置错误,会导致拼接结果不符合预期,没必要手动逐字符计数,直接按空格切分字符串即可快速拿到单词列表。 - 逐行调用自定义逻辑可以通过
DataFrame.apply()方法实现,指定axis=1即可让函数按行接收每行的字段值作为参数。
完整可运行代码如下:
import numpy as np import pandas as pd # 构建测试数据集 d = {'ABSTRACT_ID': [14145090,1900667, 8157202,6784974], 'TEXT': [ "velvet antlers vas are commonly used in tradit", "we have taken a basic biologic RPA to elucidat4", "ceftobiprole bpr is an investigational cephalo", "lipoperoxidationderived aldehydes for example",], 'LOCATION': [1, 4, 2, 1]} df = pd.DataFrame(data=d) def word_at_pos(pos, text_str): # 按空格切分文本为单词列表,由于LOCATION从1开始计数,索引需要减1 word_list = text_str.split() return word_list[pos - 1] # 逐行应用函数生成WORD列 df['WORD'] = df.apply(lambda row: word_at_pos(row['LOCATION'], row['TEXT']), axis=1) print(df)
运行后得到的WORD列结果完全匹配需求:
- ABSTRACT_ID=14145090行:
velvet - ABSTRACT_ID=1900667行:
a - ABSTRACT_ID=8157202行:
bpr - ABSTRACT_ID=6784974行:
lipoperoxidationderived
注:
str.split()不传参数时默认按任意空白字符切分,会自动忽略连续空格、首尾空格的影响,比手动指定split(' ')鲁棒性更强。
内容的提问来源于stack exchange,提问作者Shaun Potts
相关产品推荐
相关产品推荐

