You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用aiohttp优化Python HTTP请求未获性能提升的问题排查求助

aiohttp优化Python HTTP请求未获性能提升的问题排查求助

嘿,我仔细看了你的代码,发现几个关键问题导致异步编程的优势没发挥出来,难怪和同步版本性能差不多!咱们一个个说:

1. 致命问题:同步请求彻底阻塞了异步流程

你的individual_listing_scraper函数里用了**同步的requests.get**来获取单个列表项的详情页。这就相当于在异步代码里插了个“阻塞钉子”——每次调用这个函数,整个事件循环都会停下来等待同步请求完成,完全抵消了aiohttp异步的优势。

改法:

把这个函数改成异步的,用aiohttp的session来发请求,比如:

async def individual_listing_scraper(session, individual_metadata, raw_listing_html, link_metadata, site_headers, timeout_sec):
    try:
        # 提取列表URL部分逻辑不变
        listing_url_tag = raw_listing_html.find(link_metadata['tag'], class_=link_metadata['classname'])
        if not listing_url_tag:
            print("Error: Listing URL not found in the raw HTML.")
            return {}

        url = listing_url_tag.get(link_metadata['attrib'])
        if not url:
            print("Error: URL attribute not found in the tag.")
            return {}

        url = link_metadata['baseurlLink'] + url

        # 替换成异步请求
        async with session.get(url, headers=site_headers, timeout=timeout_sec, verify=False) as response:
            response.raise_for_status()
            html_text = await response.text()
            individual_page_html = BeautifulSoup(html_text, 'html.parser')
        
        # 后续解析逻辑不变...
        tag = individual_metadata['cassingIdentity']['tag']
        classname = individual_metadata['cassingIdentity']['classname']
        _id = individual_metadata['cassingIdentity']['id']
        index = individual_metadata['cassingIdentity']['index']

        listing_raw_html = individual_page_html.find_all(tag, class_=classname, id=_id)
        if not listing_raw_html or len(listing_raw_html) <= index:
            print(f"Error: Listing data not found or index out of range.")
            return {}

        listing_data = listing_raw_html[index]
        listing = bulklisting_scraper(individual_metadata, listing_data)
        listing['link'] = url
        return listing

    except aiohttp.ClientError as e:
        print(f"Request error: {e}")
    except Exception as e:
        print(f"Unexpected error: {e}")
        traceback.print_tb(e.__traceback__)
    return {}

注意要把session传进来复用,不要每次创建新session。

2. Session管理混乱,重复创建且未正确复用

  • 你在autoScraper里调用process_page时传了一个aiohttp.ClientSession(),但process_page内部又自己创建了一个新的async with aiohttp.ClientSession() as session:,这不仅浪费资源,还破坏了session的复用(比如cookie共享)。
  • autoScraper里获取初始cookie的方式也有问题:await aiohttp.ClientSession().get("https://www.google.com/")创建了session但没关闭,应该用async with来管理生命周期。

改法:

在autoScraper里统一创建session,传递给后续函数复用:

async def autoScraper(SITE_INDEX, START_PAGE_NUM=1, END_PAGE_NUM=-1):
    new_listings = []
    if SITE_INDEX >= len(sites):
        return "site number not found"

    site = sites[SITE_INDEX]
    try:
        # 用async with管理session生命周期
        async with aiohttp.ClientSession() as session:
            # 获取初始cookie
            async with session.get("https://www.google.com/") as initial_response:
                cookie_jar = initial_response.cookies

            # 把session和cookie传给process_page,不要用全局变量
            await process_page(session, SITE_INDEX, START_PAGE_NUM, END_PAGE_NUM, new_listings, cookie_jar)

        return new_listings
    except Exception as e:
        print(f"Error in autoScraper: {str(e)}")
        return new_listings

然后修改process_page的参数,去掉内部创建session的代码,用传进来的session。

3. 页面处理是串行的,没真正并发请求多个页面

当前代码里process_page是一个页面请求、处理完,再去请求下一个页面,本质上还是串行执行。要提升性能,应该把多个页面的请求任务并发起来,比如用asyncio.gather。

改法思路:

先生成所有要请求的页面URL,然后创建多个异步任务一起执行,还可以控制并发数避免被目标网站封禁:

async def process_page(session, site_index, start_page_num, end_page_num, new_listings, cookie_jar):
    site = sites[site_index]
    cur_url = site['url']
    site_headers = site['headers']
    cassing = site['casings'][0]['cassingIdentity']
    link_MD = site['casings'][0]['link']
    sitename = site['pageTitle']
    individual_Scraper_Metadata = site['individualCasings'][0]
    COUNTER = site['URL_Counter']
    
    current_page_num = start_page_num
    end_page_num *= COUNTER if end_page_num != -1 else float('inf')
    tasks = []
    max_concurrent = 10  # 控制并发数,根据目标网站调整

    while current_page_num <= end_page_num:
        # 生成页面URL
        if cur_url.count('{}') > 1:
            cur_url2 = cur_url.format(current_page_num-1, current_page_num-1)
        else:
            cur_url2 = cur_url.format(current_page_num)
        # 把每个页面的处理包装成任务
        tasks.append(process_single_page(session, cur_url2, site, current_page_num, new_listings, cookie_jar))
        
        # 达到并发上限就执行一批任务
        if len(tasks) >= max_concurrent:
            await asyncio.gather(*tasks)
            tasks = []
        
        if end_page_num != float('inf') and current_page_num >= end_page_num:
            break
        current_page_num += COUNTER
    
    # 处理剩余任务
    if tasks:
        await asyncio.gather(*tasks)
    return finalReturn(sitename, new_listings)

# 新增单个页面处理的异步函数
async def process_single_page(session, url, site, page_num, new_listings, cookie_jar):
    print(f"Requesting URL: {url}")
    try:
        soup = await get_soup(session, url, site['headers'], cookie_jar)
        raw_listings = soup.find_all(site['casings'][0]['cassingIdentity']['tag'], class_=site['casings'][0]['cassingIdentity']['classname'])
        print(f"Found {len(raw_listings)} raw listings on page {page_num}")
        if not raw_listings:
            print(f"No listings found at {url}")
            return
        
        # 并发获取所有列表项详情
        listing_tasks = []
        for raw_listing in raw_listings:
            listing_tasks.append(individual_listing_scraper(session, site['individualCasings'][0], raw_listing, site['casings'][0]['link'], site['headers'], 30))
        
        page_listings = await asyncio.gather(*listing_tasks)
        for listing in page_listings:
            if listing:
                listing['pageNo'] = page_num
                listing['siteName'] = site['pageTitle']
                new_listings.append(listing)
    except Exception as e:
        print(f"Error processing page {page_num}: {str(e)}")
        traceback.print_exc()

这样页面之间、列表项之间都能并发处理,性能才会上去。

4. 全局变量COOKIEJAR带来的隐患

用全局变量传递cookie很容易在并发场景下出问题(比如多个任务同时修改),最好通过函数参数传递,就像上面改法里那样。

把这些问题都修复后,你的异步脚本应该能明显比同步版本快很多,毕竟真正实现了请求的并发处理。

备注:内容来源于stack exchange,提问作者user119264

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 15:48:10