Python 2.7 Lambda函数解包报错Too many values to unpack求助
ValueError: too many values to unpack in Lambda Function when Building TFIDF Model
你遇到的问题核心很明确:你的mentions数据是字典类型,而非你Lambda函数里假设的嵌套元组结构,强行按元组规则解包自然会触发too many values to unpack错误。
问题分析
看你提供的mentions示例输出,每个元素都是包含_id、source、span、text键的字典,比如:
{'_id': u'en.wikipedia.org/wiki/William_Cowper', 'source': 'en.wikipedia.org/wiki/Beagle', 'span': (165, 179), 'text': u'References to the dog...'}
但你原来的Lambda函数lambda (target, (span, text)): (target, text)是在假设输入是(target, (span, text))这样的嵌套元组,两者结构完全不匹配,解包失败是必然的。
解决方案
直接从字典中通过键名提取你需要的字段即可,不需要尝试元组解包:
1. 修复你的可复现示例代码
修改后的代码如下:
import math import numpy Data = [ {'_id': '333981', 'source': 'Apple', 'span': (100, 119), 'text': ' It is native to the northern Pacific.'}, {'_id': '27262', 'source': 'Apple', 'span': (4, 20), 'text': ' Apples are yummy.'} ] # 直接从字典中提取_id和text字段 m = map(lambda x: (x['_id'], x['text']), Data) print(list(m))
运行后会得到正确输出:
[('333981', ' It is native to the northern Pacific.'), ('27262', ' Apples are yummy.')]
2. 应用到你的TFIDF类build方法
把第一个map的Lambda函数改成从字典提取字段的形式,修改后的代码片段:
class tfidf(ModelBuilder, Model): def __init__(self, max_ngram=1, normalize = True): self.max_ngram = max_ngram self.normalize = normalize def build(self, mentions, idfs): m = mentions\ .map(lambda x: (x['_id'], x['text']))\ # 从字典提取_id作为target,text作为文本内容 .mapValues(lambda v: ngrams(v, self.max_ngram))\ .flatMap(lambda (target, tokens): (((target, t), 1) for t in tokens))\ .reduceByKey(add)\ .map(lambda ((target, token), count): (token, (target, count)))\ .leftOuterJoin(idfs)\
补充说明
在Python 2.7的Spark环境中,当RDD的元素是字典时,Lambda函数的参数是整个字典对象,你需要通过键名来访问对应的值,而不是试图按元组的方式解包。之前的错误本质是混淆了输入数据的结构——把字典当成了嵌套元组来处理。
内容的提问来源于stack exchange,提问作者user3446905
相关产品推荐
相关产品推荐

