Pandas如何按指定索引从DataFrame数组列提取元素生成新列
Pandas按嵌套索引提取列表元素新增列
实现逻辑
逐行对齐tokens列的列表和对应位置的索引列表,通过列表下标直接提取目标元素即可,这种写法比调用apply方法运行效率更高。
完整代码
import pandas as pd # 构造示例数据 data = {'id':[1,2,3], 'tokens': [[ 'in', 'the' , 'morning', 'cat', 'run', 'today', 'very', 'quick'],['dog', 'eat', 'meat', 'chicken', 'from', 'bowl'], ['mouse', 'hides', 'from', 'a', 'cat']]} df = pd.DataFrame(data) lst_index = [[3, 4, 5], [0, 1, 2], [2, 3, 4]] # 新增目标列 df['new'] = [ [token_seq[i] for i in pos] for token_seq, pos in zip(df['tokens'], lst_index) ]
运行结果
执行代码后输出的DataFrame完全符合预期:
id tokens new 0 1 [in, the, morning, cat, run, today, very, quick] [cat, run, today] 1 2 [dog, eat, meat, chicken, from, bowl] [dog, eat, meat] 2 3 [mouse, hides, from, a, cat] [from, a, cat]
注意:如果
lst_index的长度和DataFrame行数不匹配,zip方法会自动按较短的序列长度截断数据,对数据一致性要求高的场景可以提前加一行长度校验,避免数据遗漏。
内容的提问来源于stack exchange,提问作者Rory
相关产品推荐
相关产品推荐

