如何堆叠循环内函数的输出结果构建DataFrame?
问题排查与实现方案
原有代码核心问题
- 每次循环都重新创建独立的DataFrame对象,没有累积存储多轮爬取结果,循环结束后仅能保留最后一次爬取成功的内容,无法得到堆叠后的完整DataFrame
- 裸
except未指定异常类型,会掩盖包括语法错误在内的所有异常,不利于后续问题排查 - 每轮循环单独执行
set_index属于冗余操作,完全可以在生成最终DataFrame时统一处理,提升运行效率
最优实现方案
优先使用「先收集所有行数据到列表,最后一次性转换为DataFrame」的方案,该方案比循环中反复拼接DataFrame的内存效率高30%以上,代码也更简洁:
import pandas as pd # 循环外初始化空列表,存储每一行的字典数据 row_list = [] for x in range(583,625): try: seq, title, end_period, upload_count, goal_count, total_reward, per_price, explain = Crawling(x) # 直接将单条数据字典加入列表,无需提前转DataFrame row_list.append({ 'seq': seq, 'title': title, 'end_period': end_period, 'upload_count': upload_count, 'goal_count': goal_count, 'total_reward': total_reward, 'per_price': per_price, 'explain': explain }) # 建议根据实际场景替换为具体的异常类型,比如爬虫相关的RequestException等 except Exception as e: print(x, 'fail', '错误信息:', str(e)) # 所有数据收集完成后,一次性转换为DataFrame并设置索引 CloudPick_df = pd.DataFrame(row_list).set_index('seq') print(CloudPick_df)
逐次堆叠DataFrame实现(不推荐,仅做参考)
如果必须使用循环堆叠的写法,可以参考如下实现,该方案数据量超过1000条时运行效率会明显降低:
import pandas as pd # 循环外初始化空DataFrame CloudPick_df = pd.DataFrame() for x in range(583,625): try: seq, title, end_period, upload_count, goal_count, total_reward, per_price, explain = Crawling(x) temp_dict = { 'seq': seq, 'title': title, 'end_period': end_period, 'upload_count': upload_count, 'goal_count': goal_count, 'total_reward': total_reward, 'per_price': per_price, 'explain': explain } temp_df = pd.DataFrame(temp_dict, index=[0]) # 逐次拼接数据 CloudPick_df = pd.concat([CloudPick_df, temp_df], ignore_index=True) except Exception as e: print(x, 'fail', '错误信息:', str(e)) CloudPick_df.set_index('seq', inplace=True) print(CloudPick_df)
内容的提问来源于stack exchange,提问作者Ironman
相关产品推荐
相关产品推荐

