You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取网页Twitter超链接运行速度过慢如何优化

爬取Twitter链接效率优化方案

你当前代码速度慢的核心原因是串行执行网络IO请求,单条请求哪怕仅耗时0.5秒,1万条串行跑下来光等响应就要近1.5小时,再加上冗余逻辑、无超时卡死、重复请求等问题,速度会更慢,可按以下方向优化:

1. 核心优化:替换串行请求为异步并发

  • swifter对网络请求这类IO密集型任务没有加速效果,它的优化场景是CPU密集的pandas数值计算,不用在这个场景下使用。
  • 直接用aiohttp+asyncio实现异步并发请求,设置30-80的并发数(根据自身带宽调整,别开太高触发站点反爬),请求效率能比串行提升几十倍,1万条数据通常几分钟就能跑完。
  • 一定要给请求设置超时阈值,遇到无响应的站点等5-10秒直接跳过,避免单个请求卡死拖慢整体进度。

2. 砍掉函数内冗余逻辑

  • 不要每次调用爬取函数都新建HTTP客户端实例,全局创建一次复用连接池,能省掉大量TCP握手、连接建立的开销。
  • 你要匹配的域名twitter.com是固定值,不用每次进函数都调用urlparse做解析,直接定义为全局常量即可。
  • 不要写裸except捕获所有异常,也不要捕获到异常后只写个ConnectionRefusedError不做任何返回,要覆盖超时、DNS解析失败、SSL错误、4xx/5xx状态码等常见异常,异常场景直接返回空列表即可。
  • 现有匹配逻辑会误抓包含twitter.com字符串的第三方跳转、统计链接,可补充域名校验逻辑,只保留host为twitter.com的链接,减少无效数据。

3. 优化解析与前置处理

  • 把BeautifulSoup的解析器从默认的html.parser换成lxml,解析速度能提升2-3倍,安装完lxml库后直接修改解析器参数即可。
  • 先对DataFrame里的网站地址列做去重,重复站点不用重复爬取;再加个内存/本地缓存,已经爬过的站点直接读缓存结果,任务中途中断后重跑也不用从头开始。
  • 你当前用SoupStrainer只提取a标签的逻辑是对的,不用解析完整DOM树,这部分可以保留。

参考实现代码

import asyncio
import aiohttp
from bs4 import BeautifulSoup, SoupStrainer
import pandas as pd

# 提前定义固定常量
SEARCH_DOMAIN = "twitter.com"
LINK_STRAINER = SoupStrainer("a")
REQUEST_TIMEOUT = aiohttp.ClientTimeout(total=8)
# 并发数,根据网络情况调整,建议30-80之间
CONCURRENT_LIMIT = 50

# 爬取结果缓存
cache = {}

async def get_twitter_links(session, website):
    if website in cache:
        return website, cache[website]
    try:
        target_url = f"https://{website}"
        async with session.get(target_url, timeout=REQUEST_TIMEOUT) as resp:
            if resp.status != 200:
                cache[website] = []
                return website, []
            html_content = await resp.text()
            twitter_links = set()
            for link in BeautifulSoup(html_content, "lxml", parseOnlyThese=LINK_STRAINER):
                href = link.get("href")
                if href and SEARCH_DOMAIN in href:
                    twitter_links.add(href)
            cache[website] = list(twitter_links)
            return website, list(twitter_links)
    except Exception:
        cache[website] = []
        return website, []

async def batch_crawl(site_list):
    sem = asyncio.Semaphore(CONCURRENT_LIMIT)
    async with aiohttp.ClientSession() as session:
        async def limited_crawl(site):
            async with sem:
                return await get_twitter_links(session, site)
        tasks = [asyncio.create_task(limited_crawl(site)) for site in site_list]
        results = await asyncio.gather(*tasks)
        return dict(results)

if __name__ == "__main__":
    # 读取你的原始DataFrame
    df = pd.read_csv("your_website_list.csv")
    # 先去重减少无效请求
    unique_sites = df["Website address"].dropna().unique().tolist()
    # 批量爬取
    site_twitter_map = asyncio.run(batch_crawl(unique_sites))
    # 结果映射回原DataFrame
    df["twitter_id"] = df["Website address"].map(site_twitter_map)

内容的提问来源于stack exchange,提问作者Sasa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 10:12:22