You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

aiohttp+asyncio网页解析初始正常,运行后无数据返回求助

异步批量爬取域名主页无响应的修复方案

你的脚本初期运行正常,后期突然无数据返回,核心原因是一次性创建所有请求任务导致并发量过高:aiohttp的ClientSession默认连接池资源有限,瞬间大量请求会耗尽连接池,导致后续请求无法发起;同时高并发也容易触发目标网站反爬机制或网络限流,造成请求“静默失败”。

此外还有几个潜在问题:

  • 全局字典data在高并发场景下虽无线程安全问题,但异步任务同时写入可能引发数据覆盖或异常;
  • 仅使用http://前缀,部分域名会强制跳转HTTPS,可能引发超时或错误;
  • 全局120秒超时设置过于宽泛,单个慢请求会阻塞整体流程。

解决方案

  • 限制并发数:用asyncio.Semaphore控制同时运行的任务数量,避免请求过载;
  • 避免全局共享状态:让任务返回结果,统一收集数据,替代全局字典;
  • 支持HTTPS并优化请求逻辑:优先尝试HTTPS,适配主流网站配置;
  • 细粒度超时控制:给单个请求设置独立超时,避免慢请求拖垮整体;
  • 完善错误日志:方便快速定位请求失败原因。

修改后的代码

import asyncio
import time
import aiohttp
import pandas as pd
import json
from bs4 import BeautifulSoup

async def get_data(session, url, semaphore):
    async with semaphore:
        try:
            # 单个请求设置30秒超时,避免慢请求阻塞
            async with session.get(url, timeout=aiohttp.ClientTimeout(total=30)) as resp:
                html = await resp.text()
                rawhtml = BeautifulSoup(html, 'lxml')
                title = rawhtml.title.string if rawhtml.title else 'None'
                return {
                    'url': url,
                    'resp': resp.status,
                    'title': title,
                    'html': str(rawhtml)
                }
        except Exception as e:
            print(f"请求失败: {url} | 错误: {str(e)}")
            return {
                'url': url,
                'resp': str(e),
                'title': 'None',
                'html': 'None'
            }

async def main():
    df = pd.read_excel('work_file.xlsx')
    domains_count = df.shape[0]
    print(f"总域名数: {domains_count}")
    
    # 限制并发数,根据网络情况调整,建议10-20
    semaphore = asyncio.Semaphore(15)
    tasks = []
    
    async with aiohttp.ClientSession() as session:
        for domain in df.domain:
            # 优先尝试HTTPS,适配主流网站配置
            url = f'https://{domain.lower()}'
            task = asyncio.create_task(get_data(session, url, semaphore))
            tasks.append(task)
        
        # 等待所有任务完成并收集结果
        results = await asyncio.gather(*tasks)
        
        # 转为原脚本的字典格式(可选)
        data = {item['url']: item for item in results}
        
        # 写入JSON文件,指定编码和缩进提升可读性
        with open('parsed_data.json', 'w', encoding='utf-8') as file:
            json.dump(data, file, ensure_ascii=False, indent=2)

if __name__ == '__main__':
    start_time = time.time()
    asyncio.run(main())
    end_time = time.time() - start_time
    print(f"总耗时: {end_time:.2f}秒")

关键修改说明

  1. 用Semaphore限制并发数,避免瞬间请求过载导致连接池耗尽;
  2. get_data函数返回结果,在main中统一收集,避免全局共享状态的潜在问题;
  3. 单个请求设置独立超时,不会因某个慢请求阻塞所有任务;
  4. 默认使用HTTPS,适配当前绝大多数网站的配置;
  5. 优化title提取逻辑,直接获取标题文本而非标签字符串;
  6. 写入JSON时指定UTF-8编码和缩进,提升文件可读性。

内容的提问来源于stack exchange,提问作者Artem Biryukov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 12:05:18