如何优化遍历万级词表的Python网页爬取脚本
Python代码优化方案
以下是针对遍历词表发起搜索查询代码的优化建议及优化后版本:
核心优化点
- 精准异常捕获:替换宽泛的
except:,捕获特定异常(如网络请求异常、JSON解析异常、键不存在异常),便于定位问题 - 文件操作改进:用追加模式写入结果,并用
with语句管理文件句柄,避免资源泄漏与数据覆盖 - 安全的URL参数传递:使用
requests.get的params参数自动处理URL编码,避免拼接参数导致的错误 - 规范化日期判断:借助
datetime模块解析时间戳字符串,替代手动字符串截取,提升可靠性 - 会话复用提升性能:使用
requests.Session()复用TCP连接,大幅减少10000次请求的耗时 - 冗余代码清理:移除未使用的变量、多余的类型转换,清理注释掉的无效代码
- 添加进度提示:实时显示当前处理进度,便于跟踪任务状态
- 请求超时设置:为每个请求设置超时时间,避免单个请求卡住整个程序
优化后代码
import requests from datetime import datetime # 初始化会话,复用TCP连接提升效率 session = requests.Session() # 定义目标日期:2018年10月14日 target_date = datetime(2018, 10, 14) # 读取词表 with open("wordlist.txt", "r", encoding="utf-8") as file: lines = [line.rstrip() for line in file] total_words = len(lines) print(f"开始处理,共{total_words}个单词...") for idx, word in enumerate(lines, 1): try: # 安全传递URL参数,自动处理编码 response = session.get( "https://us-central1-sandtable-8d0f7.cloudfunctions.net/api/creations", params={"title": word}, timeout=10 # 设置10秒超时,避免请求卡住 ) response.raise_for_status() # 主动抛出HTTP状态码异常 responsedict = response.json() # 直接用内置方法解析JSON if responsedict: # 取最后一条结果 last_item = responsedict[-1] item_data = last_item["data"] timestamp_str = item_data["timestamp"] # 解析时间戳字符串为datetime对象 item_datetime = datetime.strptime(timestamp_str, "%Y-%m-%dT%H:%M:%S.%fZ") # 判断是否早于目标日期 if item_datetime <= target_date: item_title = item_data["title"] item_id = item_data["id"] item_url = f"https://sandspiel.club/#{item_id}" # 格式化输出 print(f"\n[{idx}/{total_words}] 找到符合条件的结果:") print(f" 标题: {item_title}") print(f" 链接: {item_url}") print(f" 发布日期: {timestamp_str[:10]}") print(f" 完整时间戳: {timestamp_str}") print(f" 搜索词: {word}") # 追加写入结果到文件,自动管理文件句柄 with open("posts.txt", "a", encoding="utf-8") as out_file: out_file.write(f"{item_title}\n{item_url}\n{timestamp_str}\n\n") except requests.exceptions.RequestException as e: print(f"\n[{idx}/{total_words}] 搜索词 '{word}' 请求失败: {str(e)}") except (KeyError, ValueError) as e: print(f"\n[{idx}/{total_words}] 搜索词 '{word}' 数据解析错误: {str(e)}") except Exception as e: print(f"\n[{idx}/{total_words}] 搜索词 '{word}' 出现未知错误: {str(e)}") print("\n\n处理完成!")
优化说明
- 会话复用:
requests.Session()会保持TCP连接,相比每次新建请求,处理10000个单词时能节省大量连接建立/关闭的耗时 - 日期判断:使用
datetime.strptime解析标准格式的时间戳字符串,直接和目标日期比较,避免手动截取字符串的出错风险 - 文件写入:用
'a'模式追加写入,保留所有符合条件的结果;with语句确保文件自动关闭,不会出现资源泄漏 - 异常处理:分类型捕获异常,清晰区分请求错误、解析错误和未知错误,便于排查问题
- 进度显示:通过
enumerate显示当前处理进度,让你清楚任务执行状态 - 超时设置:10秒超时避免因网络问题导致单个请求长时间阻塞,提升整体任务的稳定性
内容的提问来源于stack exchange,提问作者Vixey
相关产品推荐
相关产品推荐

