Pandas技术问题:如何不使用.append()向DataFrame重复追加单行?
替代.append()的高效实现方案
核心思路
抛弃效率低下的双重行遍历,改用Pandas的向量化筛选+批量合并的方式,既避免弃用警告,又大幅提升处理速度。
实现代码
import pandas as pd # 初始化空列表用于收集匹配结果 matches_list = [] # 遍历每个地名记录 for _, place_row in placenames.iterrows(): # 筛选出文本中包含当前地名的所有铭文记录 matched_inscriptions = inscriptions[ inscriptions['text'].str.contains(place_row['name'], na=False) ] # 如果有匹配结果,添加地名相关字段并加入列表 if not matched_inscriptions.empty: # 复制数据避免修改原DataFrame matched_copy = matched_inscriptions.copy() matched_copy['place'] = place_row['name'] matched_copy['latitude'] = place_row['latitude'] matches_list.append(matched_copy) # 一次性合并所有匹配结果 matches = pd.concat(matches_list, ignore_index=True)
关键优化点
- 向量化筛选:用
str.contains替代逐行文本匹配,Pandas底层会用优化后的逻辑处理,比Python循环快数倍。 - 批量合并:先把所有匹配结果存入列表,最后用
pd.concat一次性合并,避免循环中反复修改DataFrame的性能损耗。 - 避免原数据修改:用
.copy()复制匹配的铭文行,防止意外修改原inscriptions数据。
可选:精确匹配(避免子串误匹配)
如果需要避免类似"Paris"匹配"Parisian"的情况,可以用正则边界匹配,同时转义地名中的特殊字符:
import re matches_list = [] for _, place_row in placenames.iterrows(): # 转义地名中的正则特殊字符,生成精确匹配的正则表达式 pattern = rf'\b{re.escape(place_row["name"])}\b' matched_inscriptions = inscriptions[ inscriptions['text'].str.contains(pattern, regex=True, na=False) ] if not matched_inscriptions.empty: matched_copy = matched_inscriptions.copy() matched_copy['place'] = place_row['name'] matched_copy['latitude'] = place_row['latitude'] matches_list.append(matched_copy) matches = pd.concat(matches_list, ignore_index=True)
内容的提问来源于stack exchange,提问作者Thomas Leibundgut
相关产品推荐
相关产品推荐

