You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效从70万条URL批量抓取文本并写入DataFrame?

提速方案:并行处理URL请求

你当前的代码是单线程串行请求每个URL,大部分时间都在等待网络响应,完全浪费了服务器的硬件资源。下面是几个能大幅提速的实用方案:

1. 多线程并行(最易实现,适合IO密集型任务)

用concurrent.futures.ThreadPoolExecutor开启多线程,让多个请求同时发起,在等待一个请求响应的间隙处理其他请求,把服务器的带宽和CPU资源充分利用起来。

代码示例:

import pandas as pd
from urllib.request import urlopen
from concurrent.futures import ThreadPoolExecutor

def get_text(url):
    try:
        return urlopen(url, timeout=10).read().decode()  # 加超时避免卡壳
    except Exception as e:
        print(f"请求{url}失败: {str(e)}")
        return ""

# 线程数建议设50-200(IO密集型任务可以开多些)
with ThreadPoolExecutor(max_workers=150) as executor:
    df['txt'] = list(executor.map(get_text, df['Text url']))

2. 异步请求(效率更高)

用aiohttp做异步HTTP请求,比多线程的资源占用更低、并发量更高,更适合处理几十万级的URL请求场景。

代码示例:

import pandas as pd
import aiohttp
import asyncio

async def fetch_text(session, url):
    try:
        async with session.get(url, timeout=10) as resp:
            return await resp.text()
    except Exception as e:
        print(f"请求{url}失败: {str(e)}")
        return ""

async def batch_fetch(urls):
    async with aiohttp.ClientSession() as session:
        tasks = [fetch_text(session, url) for url in urls]
        return await asyncio.gather(*tasks)

# 执行异步任务
df['txt'] = asyncio.run(batch_fetch(df['Text url'].tolist()))

额外优化点

  • 超时设置:必须加超时限制,防止个别慢请求卡住整个任务流程
  • 错误捕获:处理请求失败的异常情况,避免单个请求出错导致程序崩溃
  • 分批处理:如果一次性处理70万条数据内存吃紧,可以把DataFrame分成若干批次,每批处理完再合并结果
  • 连接复用:异步的ClientSession会自动复用HTTP连接,比每次新建连接节省大量开销

内容的提问来源于stack exchange,提问作者Remrem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 06:05:07