You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Pandas另一列的元组,从文本列选取可变数量的Token?

问题描述

现有一个包含两列的Pandas DataFrame:sentence列存储文本,selector列存储不同长度的元组数组。示例代码如下:

import pandas as pd
df = pd.DataFrame({'sentence': ['KEEP some of the words from this sentence.',
                                'Keep SOME of THE words from this sentence.',
                                'KEEP some OF the WORDS from this sentence.',
                                'Keep SOME of THE words FROM this SENTENCE.'],
                   'selector': [[(10, 0, 1)],
                                [(10, 1, 2), (10, 3, 4)],
                                [(10, 0, 1), (10, 2, 3), (10, 4, 5)],
                                [(10, 1, 2), (10, 3, 4), (10, 5, 6), (10, 7, 8)]]})

需求:从sentence的分词结果中,选取selector里每个元组的第二个元素(忽略元组首个值)指定位置的词,组成selected_tokens列。最终期望的DataFrame示例如下:

sentence                                    selector                                           selected_tokens
KEEP some of the words from this sentence.  [(10, 0, 1)],                                      ['KEEP']
KEEP some OF the WORDS from this sentence.  [(10, 0, 1), (100, 2, 3), (10, 4, 5)],             ['KEEP', 'OF', 'WORDS']
Keep SOME of THE words from this sentence.  [(10, 1, 2), (10, 3, 4)],                          ['SOME', 'THE']
Keep SOME of THE words FROM this SENTENCE.  [(10, 1, 2), (10, 3, 4), (10, 5, 6), (10, 7, 8)],  ['SOME', 'THE', 'FROM', 'SENTENCE']

此前尝试过提取单个Token的方法,但因selector中元组数量可变(真实数据为0-25个),该方法繁琐且无法适配所有场景,需最优实现方案。

解决方案

方法1:apply + 自定义函数(直观易读)

先对文本分词,再逐行提取指定位置的Token:

# 预分词,避免重复计算
df['tokens'] = df['sentence'].str.split()

def extract_tokens(tokens, selectors):
    return [tokens[tpl[1]] for tpl in selectors]

# 生成目标列
df['selected_tokens'] = df.apply(lambda row: extract_tokens(row['tokens'], row['selector']), axis=1)

# 可选:删除中间列
df = df.drop('tokens', axis=1)

方法2:列表推导式(性能更优)

对于大数据量,列表推导式的执行效率远高于apply:

# 转换为Python列表,避免Pandas的Series操作开销
tokens_list = df['sentence'].str.split().tolist()
selector_list = df['selector'].tolist()

# 批量生成结果
df['selected_tokens'] = [
    [tokens[tpl[1]] for tpl in selectors]
    for tokens, selectors in zip(tokens_list, selector_list)
]

方法3:适配空selector场景

如果存在selector为空列表的情况,可添加判断逻辑(示例返回None,可按需调整):

df['selected_tokens'] = [
    [tokens[tpl[1]] for tpl in selectors] if selectors else None
    for tokens, selectors in zip(tokens_list, selector_list)
]

内容的提问来源于stack exchange,提问作者Ivo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 05:45:36