Python/Pandas遍历CSV时按首字母输出处理进度的实现方法
实现方案解答
核心逻辑说明
不需要额外加while循环,反而会增加冗余逻辑和出错概率。最高效的实现方式是在现有遍历循环中增加一个临时变量跟踪上一行的首字母,利用CSV已经按APPLICANT列排序、同首字母行连续的特性,仅需一次遍历即可完成需求,时间复杂度保持O(n),无额外性能开销。
完整修改后代码
from elasticsearch import Elasticsearch import pandas as pd csv = "file.csv" csv_file = pd.read_csv(csv, sep=",", header=0, index_col=False) client = Elasticsearch() csv_output = "es_results.csv" es_index = "main" full_list = [] # 新增:初始化上一行首字母跟踪变量 prev_first_char = None for index, row in csv_file.iterrows(): institution_id = row["ID"] org_name_value = row['APPLICANT'] query_body = {"an elasicsearch query"} response = client.search( index=es_index, body=query_body) for hit in response['hits']['hits']: if hit['_score'] >= 0.9: result = institution_id, org_name_value, hit['_score'] full_list.append(result) # 新增:首字母判断逻辑,和ES查询处理逻辑平级,属于外层遍历循环的子代码 # 取首字母转为大写,避免大小写不同导致判断错误 current_first_char = org_name_value[0].upper() if prev_first_char is None: # 首次遍历赋值初始值 prev_first_char = current_first_char elif current_first_char != prev_first_char: # 首字母变化,说明上一个首字母的所有行已处理完成 print(f"{prev_first_char} is done") prev_first_char = current_first_char # 新增:循环结束后打印最后一个首字母的完成提示 print(f"{prev_first_char} is done") df = pd.DataFrame(full_list) df.to_csv(csv_output, index=False, header=['Institution_Id', 'Applicant_Name', 'Score'])
缩进规则说明
新增的首字母判断逻辑和ES查询的代码是平级关系,都属于外层for index, row in csv_file.iterrows()循环的子代码块,不要缩进至内层for hit in response['hits']['hits']循环内即可。
可选兼容优化
如果APPLICANT列可能存在空值,可以在取首字母前增加空值判断,避免索引报错:
# 放在current_first_char赋值前 if pd.isna(org_name_value) or len(org_name_value.strip()) == 0: print(f"第{index}行APPLICANT为空,跳过") continue
内容的提问来源于stack exchange,提问作者Stpete111
相关产品推荐
相关产品推荐

