Python函数求助:生成含字典与标签的元组列表不符合预期
修正featureExtraction函数代码
现有代码的问题在于两层循环遍历到每个单词,每次生成单个单词的字典,还错误使用了未定义的tup变量,正确应该操作函数内的features列表。
修正后的代码:
def featureExtraction(clean_tokenized, label): features = [] for line in clean_tokenized: # 为当前评论的所有单词生成一个字典 feature_dict = {word: True for word in line} features.append((feature_dict, label)) return features
修改说明
- 移除了遍历单个单词的内层循环,改为针对每条评论(子列表)生成一个完整的特征字典
- 使用字典推导式快速构建包含当前评论所有单词的字典,每个单词对应值为
True - 直接将(特征字典, 标签)元组添加到
features列表中,确保每条评论对应一个元组
测试验证
调用示例:
input_data = [['hate', 'movie'], ['acting', 'terrible']] print(featureExtraction(input_data, 'neg'))
输出结果:
[({'hate': True, 'movie': True}, 'neg'), ({'acting': True, 'terrible': True}, 'neg')]
完全符合预期输出。
内容的提问来源于stack exchange,提问作者Kevin Veeder
相关产品推荐
相关产品推荐

