异步Web Scraping获取数据报错问题排查及优化需求
问题分析与解决方案
核心错误原因
- 解析函数参数错误:所有解析函数(
title_bs4、url_bs4等)期望接收HTML文本,但你传入的是aiohttp的Response对象,导致BeautifulSoup无法解析,直接触发异常。 - 数据收集逻辑混乱:
url_bs4中错误提取了标签文本而非链接的href属性;- 把URL、价格结果错误添加到
list_titles列表,而非对应的list_urls、list_prices,导致最终数据为空。
- 分页请求缺失:循环中仅发起了第一页的请求,后续拿到下一页URL后未调用
session.get获取新页面,一直复用初始的Response。 - 异常处理过于宽泛:直接用
except:捕获所有异常,无法定位具体错误(比如选择器失效、网络错误等)。 - 不必要的异步封装:BeautifulSoup和lxml解析是CPU密集型操作,不需要封装为异步函数,反而增加异步调度开销。
修正后的代码
import asyncio import aiohttp from bs4 import BeautifulSoup import lxml.etree as etree import pandas as pd import re # 解析函数改为普通函数(CPU密集操作无需异步) def soup(html): return BeautifulSoup(html, 'html.parser') def title_bs4(html, tag, classes): soup_obj = soup(html) titles = soup_obj.findAll(tag, attrs={"class": classes}) return [i.text.strip() for i in titles] def url_bs4(html, tag, classes): soup_obj = soup(html) urls = soup_obj.findAll(tag, attrs={"class": classes}) return [i.get('href') for i in urls] # 提取href属性而非文本 def price_xpath(html): soup_obj = soup(html) dom = etree.HTML(str(soup_obj)) # 简化xpath,提升鲁棒性 price_nodes = dom.xpath('//span[@class="price-tag-fraction"]') return [i.text.replace('.', '') for i in price_nodes if i.text] def page_number_bs4(html): soup_obj = soup(html) # 选择当前激活的页码(带--active类) current_page = soup_obj.find('span', attrs={"class": "andes-pagination__link andes-pagination__link--active"}) if not current_page: return 1 return int(current_page.text.strip()) def number_of_pages_bs4(html): soup_obj = soup(html) page_count = soup_obj.find('li', attrs={"class": "andes-pagination__page-count"}) if not page_count: return 1 return int(page_count.text.strip().split(" ")[1]) def next_xpath(html): soup_obj = soup(html) dom = etree.HTML(str(soup_obj)) next_links = dom.xpath('//li[contains(@class,"--next")]/a/@href') return next_links[0] if next_links else None async def fetch_page(session, url): # 添加请求头,避免被反爬识别为爬虫 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } async with session.get(url, headers=headers) as response: if response.status != 200: raise Exception(f"请求失败,状态码:{response.status}") return await response.text() async def main(product): web = "Mercado libre" list_titles = [] list_urls = [] list_prices = [] current_url = f'https://listado.mercadolibre.com.co/{product}' async with aiohttp.ClientSession() as session: while current_url: try: # 先获取当前页面的HTML文本 html = await fetch_page(session, current_url) # 解析数据 titles = title_bs4(html, 'h2', 'ui-search-item__title shops__item-title') list_titles.extend(titles) urls = url_bs4(html, 'a', 'ui-search-item__group__element shops__items-group-details ui-search-link') list_urls.extend(urls) prices = price_xpath(html) list_prices.extend(prices) # 获取分页信息 current_page = page_number_bs4(html) total_pages = number_of_pages_bs4(html) print(f"已抓取第 {current_page}/{total_pages} 页") # 判断是否到最后一页 if current_page >= total_pages: break # 获取下一页URL current_url = next_xpath(html) except Exception as e: # 打印具体异常信息,方便排查 print(f"错误详情:{str(e)}") break # 处理数据对齐(避免因解析错误导致列表长度不一致) min_length = min(len(list_titles), len(list_urls), len(list_prices)) list_titles = list_titles[:min_length] list_urls = list_urls[:min_length] list_prices = list_prices[:min_length] # 生成DataFrame并处理价格 df = pd.DataFrame({ "shop": web, "titles": list_titles, "links": list_urls, "prices": list_prices }) # 安全处理价格转换 df['prices'] = df['prices'].apply(lambda x: float(re.search(r"\d+", str(x)).group(0)) if re.search(r"\d+", str(x)) else None) df.to_json("templates/product.json", orient='records', ensure_ascii=False) return df if __name__ == "__main__": try: asyncio.run(main('samsung')) except KeyboardInterrupt: print("程序被用户中断")
额外优化建议
- 并发请求分页:如果想要进一步提升速度,可以一次性获取所有分页URL,然后用
asyncio.gather并发请求所有页面,避免串行等待。 - 添加重试机制:针对网络请求失败的情况,添加重试逻辑,提升稳定性。
- 使用代理IP:如果频繁请求被网站限制,可以添加代理IP池。
- 选择器优化:尽量使用更简洁、鲁棒的CSS选择器替代冗长的XPath,降低页面结构变化导致的解析失败概率。
内容的提问来源于stack exchange,提问作者YTorres
相关产品推荐
相关产品推荐

