如何不使用索引向pandas列末尾追加值及解决append相关报错
问题原因
- 作用域错误:你在
ReviewSearchEngine的add方法内对全局变量DB做赋值操作时,Python默认会将DB识别为该方法的局部变量,在赋值前读取该变量就会触发UnboundLocalError。 - 方法弃用问题:pandas 1.4.0及以上版本已经正式弃用了
DataFrame.append()方法,即使修复作用域问题,后续运行也会触发弃用警告,官方推荐使用pd.concat实现追加逻辑。
解决方案
方案1:快速修复现有代码
在add方法开头添加global声明,明确使用全局变量DB,同时替换append为concat:
def add(self, review:Review): global DB x = pd.DataFrame({'reviews': [review]}) DB = pd.concat([DB, x], ignore_index=True)
方案2:更合理的工程实现(推荐)
将存储数据的DB作为ReviewSearchEngine的实例属性,避免全局变量污染,更符合“在类内部完成操作”的需求:
import pandas as pd import json class Review: def __init__(self, json_string): self.json_string=json_string def get_text(self): json_dict=json.loads(self.json_string) return json_dict['body'] class ReviewSearchEngine: def __init__(self): # 把存储结构放在类内部管理,不需要全局变量 self.DB = pd.DataFrame(columns=['reviews']) def add(self, review:Review): x = pd.DataFrame({'reviews': [review]}) self.DB = pd.concat([self.DB, x], ignore_index=True) if __name__ == '__main__': search_engine = ReviewSearchEngine() file_path = "./review_data.txt" # 用with上下文管理打开文件,避免资源泄漏 with open(file_path, 'r', encoding='utf-8') as f: lines = f.readlines() for line in lines: review = Review(line) search_engine.add(review) # 输出即可得到你预期的格式 print(search_engine.DB)
运行后search_engine.DB的输出格式和你预期完全一致。
内容的提问来源于stack exchange,提问作者Asael Bar Ilan
相关产品推荐
相关产品推荐

