You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

异步Web Scraping获取数据报错问题排查及优化需求

问题分析与解决方案

核心错误原因

  1. 解析函数参数错误:所有解析函数(title_bs4、url_bs4等)期望接收HTML文本,但你传入的是aiohttp的Response对象,导致BeautifulSoup无法解析,直接触发异常。
  2. 数据收集逻辑混乱:
    • url_bs4中错误提取了标签文本而非链接的href属性;
    • 把URL、价格结果错误添加到list_titles列表,而非对应的list_urls、list_prices,导致最终数据为空。
  3. 分页请求缺失:循环中仅发起了第一页的请求,后续拿到下一页URL后未调用session.get获取新页面,一直复用初始的Response。
  4. 异常处理过于宽泛:直接用except:捕获所有异常,无法定位具体错误(比如选择器失效、网络错误等)。
  5. 不必要的异步封装:BeautifulSoup和lxml解析是CPU密集型操作,不需要封装为异步函数,反而增加异步调度开销。

修正后的代码

import asyncio
import aiohttp
from bs4 import BeautifulSoup
import lxml.etree as etree
import pandas as pd
import re

# 解析函数改为普通函数(CPU密集操作无需异步)
def soup(html):
    return BeautifulSoup(html, 'html.parser')

def title_bs4(html, tag, classes):
    soup_obj = soup(html)
    titles = soup_obj.findAll(tag, attrs={"class": classes})
    return [i.text.strip() for i in titles]

def url_bs4(html, tag, classes):
    soup_obj = soup(html)
    urls = soup_obj.findAll(tag, attrs={"class": classes})
    return [i.get('href') for i in urls]  # 提取href属性而非文本

def price_xpath(html):
    soup_obj = soup(html)
    dom = etree.HTML(str(soup_obj))
    # 简化xpath,提升鲁棒性
    price_nodes = dom.xpath('//span[@class="price-tag-fraction"]')
    return [i.text.replace('.', '') for i in price_nodes if i.text]

def page_number_bs4(html):
    soup_obj = soup(html)
    # 选择当前激活的页码(带--active类)
    current_page = soup_obj.find('span', attrs={"class": "andes-pagination__link andes-pagination__link--active"})
    if not current_page:
        return 1
    return int(current_page.text.strip())

def number_of_pages_bs4(html):
    soup_obj = soup(html)
    page_count = soup_obj.find('li', attrs={"class": "andes-pagination__page-count"})
    if not page_count:
        return 1
    return int(page_count.text.strip().split(" ")[1])

def next_xpath(html):
    soup_obj = soup(html)
    dom = etree.HTML(str(soup_obj))
    next_links = dom.xpath('//li[contains(@class,"--next")]/a/@href')
    return next_links[0] if next_links else None

async def fetch_page(session, url):
    # 添加请求头,避免被反爬识别为爬虫
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    async with session.get(url, headers=headers) as response:
        if response.status != 200:
            raise Exception(f"请求失败,状态码:{response.status}")
        return await response.text()

async def main(product):
    web = "Mercado libre"
    list_titles = []
    list_urls = []
    list_prices = []
    current_url = f'https://listado.mercadolibre.com.co/{product}'
    
    async with aiohttp.ClientSession() as session:
        while current_url:
            try:
                # 先获取当前页面的HTML文本
                html = await fetch_page(session, current_url)
                
                # 解析数据
                titles = title_bs4(html, 'h2', 'ui-search-item__title shops__item-title')
                list_titles.extend(titles)
                
                urls = url_bs4(html, 'a', 'ui-search-item__group__element shops__items-group-details ui-search-link')
                list_urls.extend(urls)
                
                prices = price_xpath(html)
                list_prices.extend(prices)
                
                # 获取分页信息
                current_page = page_number_bs4(html)
                total_pages = number_of_pages_bs4(html)
                
                print(f"已抓取第 {current_page}/{total_pages} 页")
                
                # 判断是否到最后一页
                if current_page >= total_pages:
                    break
                
                # 获取下一页URL
                current_url = next_xpath(html)
                
            except Exception as e:
                # 打印具体异常信息,方便排查
                print(f"错误详情:{str(e)}")
                break
    
    # 处理数据对齐(避免因解析错误导致列表长度不一致)
    min_length = min(len(list_titles), len(list_urls), len(list_prices))
    list_titles = list_titles[:min_length]
    list_urls = list_urls[:min_length]
    list_prices = list_prices[:min_length]
    
    # 生成DataFrame并处理价格
    df = pd.DataFrame({
        "shop": web,
        "titles": list_titles,
        "links": list_urls,
        "prices": list_prices
    })
    
    # 安全处理价格转换
    df['prices'] = df['prices'].apply(lambda x: float(re.search(r"\d+", str(x)).group(0)) if re.search(r"\d+", str(x)) else None)
    
    df.to_json("templates/product.json", orient='records', ensure_ascii=False)
    return df

if __name__ == "__main__":
    try:
        asyncio.run(main('samsung'))
    except KeyboardInterrupt:
        print("程序被用户中断")

额外优化建议

  • 并发请求分页:如果想要进一步提升速度,可以一次性获取所有分页URL,然后用asyncio.gather并发请求所有页面,避免串行等待。
  • 添加重试机制:针对网络请求失败的情况,添加重试逻辑,提升稳定性。
  • 使用代理IP:如果频繁请求被网站限制,可以添加代理IP池。
  • 选择器优化:尽量使用更简洁、鲁棒的CSS选择器替代冗长的XPath,降低页面结构变化导致的解析失败概率。

内容的提问来源于stack exchange,提问作者YTorres

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 13:55:38