将tokenize输出元组拆解为DataFrame两列时遇ValueError问题求助
Pandas中progress_apply返回元组时无法拆分为两列的问题解决
问题场景
我编写了一个用于文本编码的tokenize函数,代码如下:
def tokenize(text, max_len=MAX_LEN): encoded = tokenizer.encode_plus( text, add_special_tokens=True, max_length = max_len, padding='max_length', return_attention_mask=True ) return encoded['input_ids'], encoded['attention_mask']
尝试将该函数应用到训练数据的text列,生成input_ids和attention_masks两列:
df_train['input_ids'], df_train['attention_masks'] = df_train['text'].progress_apply(tokenize)
但执行时持续报错:
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Cell In[13], line 1 ----> 1 df_train['input_ids'], df_train['attention_masks'] = df_train['text'].progress_apply(tokenize) 2 df_val['input_ids'], df_val['attention_masks'] = df_val['text'].progress_apply(tokenize) 3 df_test['input_ids'], df_test['attention_masks'] = df_test['text'].progress_apply(tokenize) ValueError: too many values to unpack (expected 2)
我推测是Pandas错误地将元组内的列表拆解,而非将整个列表存入对应列。试过一些方案,但会导致返回结果为Series而非列表,影响后续流程。希望找到无需临时列拆分的优雅解决方法,explode方法也不合适。
附可复现代码:
from transformers import DistilBertTokenizer import pandas as pd from tqdm import tqdm tqdm.pandas() MAX_LEN = 64 tokenizer = DistilBertTokenizer.from_pretrained("distilbert-base-uncased") train_dict = {'text': ['text1', 'text2', 'text3']} df_train = pd.DataFrame.from_dict(train_dict)
解决方法
方法1:利用zip拆分元组序列
直接将progress_apply返回的元组序列用zip拆解,再转为列表赋值给列:
# 拆分元组序列为两个独立的迭代器 input_ids, attention_masks = zip(*df_train['text'].progress_apply(tokenize)) # 转为列表后赋值 df_train['input_ids'] = list(input_ids) df_train['attention_masks'] = list(attention_masks)
方法2:将结果转为DataFrame批量赋值
把progress_apply的结果转为DataFrame,直接一次性赋值给目标列:
# 将元组序列转为DataFrame tokenized_df = df_train['text'].progress_apply(tokenize).apply(pd.Series) # 批量赋值 df_train[['input_ids', 'attention_masks']] = tokenized_df
方法3:修改tokenize函数返回字典
调整tokenize函数返回字典,再用apply(pd.Series)展开为列:
def tokenize(text, max_len=MAX_LEN): encoded = tokenizer.encode_plus( text, add_special_tokens=True, max_length = max_len, padding='max_length', return_attention_mask=True ) # 返回字典而非元组 return {'input_ids': encoded['input_ids'], 'attention_masks': encoded['attention_mask']} # 展开字典为DataFrame并赋值 df_train[['input_ids', 'attention_masks']] = df_train['text'].progress_apply(tokenize).apply(pd.Series)
报错原因说明
progress_apply返回的是一个Series对象,其中每个元素是tokenize函数返回的元组。直接用df_train['a'], df_train['b'] = Series时,Pandas会尝试将Series的每一行拆分为两个值,而非将整个Series拆分为两个部分,因此触发"too many values to unpack"错误。
内容的提问来源于stack exchange,提问作者Will Will
相关产品推荐
相关产品推荐

