You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化遍历万级词表的Python网页爬取脚本

Python代码优化方案

以下是针对遍历词表发起搜索查询代码的优化建议及优化后版本:

核心优化点

  • 精准异常捕获:替换宽泛的except:,捕获特定异常(如网络请求异常、JSON解析异常、键不存在异常),便于定位问题
  • 文件操作改进:用追加模式写入结果,并用with语句管理文件句柄,避免资源泄漏与数据覆盖
  • 安全的URL参数传递:使用requests.get的params参数自动处理URL编码,避免拼接参数导致的错误
  • 规范化日期判断:借助datetime模块解析时间戳字符串,替代手动字符串截取,提升可靠性
  • 会话复用提升性能:使用requests.Session()复用TCP连接,大幅减少10000次请求的耗时
  • 冗余代码清理:移除未使用的变量、多余的类型转换,清理注释掉的无效代码
  • 添加进度提示:实时显示当前处理进度,便于跟踪任务状态
  • 请求超时设置:为每个请求设置超时时间,避免单个请求卡住整个程序

优化后代码

import requests
from datetime import datetime

# 初始化会话,复用TCP连接提升效率
session = requests.Session()
# 定义目标日期:2018年10月14日
target_date = datetime(2018, 10, 14)

# 读取词表
with open("wordlist.txt", "r", encoding="utf-8") as file:
    lines = [line.rstrip() for line in file]
total_words = len(lines)

print(f"开始处理,共{total_words}个单词...")

for idx, word in enumerate(lines, 1):
    try:
        # 安全传递URL参数,自动处理编码
        response = session.get(
            "https://us-central1-sandtable-8d0f7.cloudfunctions.net/api/creations",
            params={"title": word},
            timeout=10  # 设置10秒超时,避免请求卡住
        )
        response.raise_for_status()  # 主动抛出HTTP状态码异常
        responsedict = response.json()  # 直接用内置方法解析JSON

        if responsedict:
            # 取最后一条结果
            last_item = responsedict[-1]
            item_data = last_item["data"]
            timestamp_str = item_data["timestamp"]
            # 解析时间戳字符串为datetime对象
            item_datetime = datetime.strptime(timestamp_str, "%Y-%m-%dT%H:%M:%S.%fZ")
            
            # 判断是否早于目标日期
            if item_datetime <= target_date:
                item_title = item_data["title"]
                item_id = item_data["id"]
                item_url = f"https://sandspiel.club/#{item_id}"
                
                # 格式化输出
                print(f"\n[{idx}/{total_words}] 找到符合条件的结果:")
                print(f"  标题: {item_title}")
                print(f"  链接: {item_url}")
                print(f"  发布日期: {timestamp_str[:10]}")
                print(f"  完整时间戳: {timestamp_str}")
                print(f"  搜索词: {word}")
                
                # 追加写入结果到文件,自动管理文件句柄
                with open("posts.txt", "a", encoding="utf-8") as out_file:
                    out_file.write(f"{item_title}\n{item_url}\n{timestamp_str}\n\n")
    
    except requests.exceptions.RequestException as e:
        print(f"\n[{idx}/{total_words}] 搜索词 '{word}' 请求失败: {str(e)}")
    except (KeyError, ValueError) as e:
        print(f"\n[{idx}/{total_words}] 搜索词 '{word}' 数据解析错误: {str(e)}")
    except Exception as e:
        print(f"\n[{idx}/{total_words}] 搜索词 '{word}' 出现未知错误: {str(e)}")

print("\n\n处理完成!")

优化说明

  1. 会话复用:requests.Session()会保持TCP连接,相比每次新建请求,处理10000个单词时能节省大量连接建立/关闭的耗时
  2. 日期判断:使用datetime.strptime解析标准格式的时间戳字符串,直接和目标日期比较,避免手动截取字符串的出错风险
  3. 文件写入:用'a'模式追加写入,保留所有符合条件的结果;with语句确保文件自动关闭,不会出现资源泄漏
  4. 异常处理:分类型捕获异常,清晰区分请求错误、解析错误和未知错误,便于排查问题
  5. 进度显示:通过enumerate显示当前处理进度,让你清楚任务执行状态
  6. 超时设置:10秒超时避免因网络问题导致单个请求长时间阻塞,提升整体任务的稳定性

内容的提问来源于stack exchange,提问作者Vixey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 22:42:39