如何将Python循环输出保存至Excel并解决分数重复问题
问题根源
生成Excel时所有行分数值重复,核心原因是构建DataFrame时误用了函数对象(positive_score、negative_score),而非已经计算完成的分数列表(positive_scores、negative_scores)。你的代码已经正确算出每个URL对应的正负分数,但最后组装输出数据时用错了变量。
修复方案
1. 修正Excel生成代码
直接使用计算好的分数列表替换函数名:
# 先计算极性分数(如果需要添加到输出) polarity_scores = [polarity_score(pos, neg) for pos, neg in zip(positive_scores, negative_scores)] # 构建正确的数据字典 data_collection = { 'URL_ID': url_ids, 'URL': urls, 'POSITIVE SCORE': positive_scores, 'NEGATIVE SCORE': negative_scores, 'POLARITY SCORE': polarity_scores # 可选,不需要可以删除 } # 生成Excel excel_data_df = pd.DataFrame(data_collection) excel_data_df.to_excel("Output.xlsx", index=False)
2. 确保分数与URL_ID顺序一致
注意os.walk遍历文件的顺序不一定和原始url_ids的顺序匹配,会导致分数和URL对应错误。建议按url_id的顺序读取文件:
# 重新生成positive_scores(替换原循环) positive_scores = [] for url_id in url_ids: file_path = os.path.join(new_texts_folder, f"{url_id}.txt") if os.path.exists(file_path): with codecs.open(file_path, encoding='utf-8', errors='ignore') as info: new_content = eval(info.read()) positive_scores.append(positive_score(new_content)) else: positive_scores.append(0) # 文件不存在时默认0分 # 重新生成negative_scores(替换原循环) negative_scores = [] for url_id in url_ids: file_path = os.path.join(new_texts_folder, f"{url_id}.txt") if os.path.exists(file_path): with codecs.open(file_path, encoding='utf-8', errors='ignore') as info: new_content = eval(info.read()) negative_scores.append(negative_score(new_content)) else: negative_scores.append(0)
3. 优化:一次遍历计算正负分数
避免重复遍历文件夹,提升效率:
positive_scores = [] negative_scores = [] for url_id in url_ids: file_path = os.path.join(new_texts_folder, f"{url_id}.txt") if os.path.exists(file_path): with codecs.open(file_path, encoding='utf-8', errors='ignore') as info: new_content = eval(info.read()) # 一次计算两个分数 pos_score = positive_score(new_content) neg_score = negative_score(new_content) positive_scores.append(pos_score) negative_scores.append(neg_score) else: positive_scores.append(0) negative_scores.append(0)
内容的提问来源于stack exchange,提问作者Abuchi
相关产品推荐
相关产品推荐

