如何加速基于BeautifulSoup4的手表市场网络爬虫脚本?
问题背景
我正在做一个小型项目用于丰富GitHub内容并学习网络爬虫,目标是从手表市场网站爬取部分热门品牌的手表数据。目前已能以数组形式获取部分数据,但爬取更多数据时耗时过长,请问有哪些优化该脚本的建议?
当前脚本代码
from bs4 import BeautifulSoup import requests headers = {"User-Agent":"Mozilla/5.0"} chrono24 = 'https://www.chrono24.com' single_watch_url = 'https://www.chrono24.com/rolex/rolex-rolex-gmt-master-ii-batman-oyster-bracelet-126710blnr-2023--id20505792.htm' watch_list_url = 'https://www.chrono24.com/rolex/index.htm' def parse_watch_page(url): """ function to parse a single watch page :param url: url to a singular watch page :return: a list of descriptor/info tuples """ # get page as bs4 html parser html = requests.get(url, headers=headers) page = BeautifulSoup(html.content, 'lxml') # searches for table that contains "Basic Info" text basic_info_table = page.find_all(lambda tag: tag.name =='tbody' and "Basic Info" in tag.text)[0] # gets rows of basic info table table_rows = basic_info_table.find_all('tr') # parses each row for td tags and puts pairings into lists datalist = [] for row in table_rows: elements = row.find_all('td') test = tuple([element.text.strip() for element in elements]) datalist.append(test) datalist.pop(0) return datalist def parse_watch_list(url): """ function to parse a watch list page :param url: url to a page to a list of watches and next button :return: link to next? """ # get page as bs4 html parser html = requests.get(url, headers=headers) page = BeautifulSoup(html.content, 'lxml') # finds all links to watches on the list watches = page.find_all('a', 'article-item block-item rcard') # parses each watch page for watch in watches: parse_watch_page(chrono24 + watch['href']) parse_watch_list(watch_list_url)
优化建议
改用异步请求提升并发效率
原脚本用requests同步请求,每次只能等一个请求完成再发起下一个,大量数据爬取时等待时间叠加导致效率极低。建议用aiohttp结合asyncio实现异步并发,同时发起多个手表页面请求,大幅压缩整体耗时。核心是把parse_watch_page改成异步函数,用任务队列批量调度请求,避免串行等待。添加请求延迟与UA随机化
无间隔连续请求容易触发网站反爬限流,也会降低请求成功率。建议在请求之间加入1-3秒的随机延迟,同时维护一个User-Agent列表,每次请求随机选取一个,模拟真实用户访问行为。同步场景用time.sleep(random.uniform(1,3)),异步场景用asyncio.sleep(random.uniform(1,3))。优化页面解析逻辑
- 原代码遍历所有
tbody查找目标表格效率偏低,可先定位包含"Basic Info"的标题元素,再通过父节点/兄弟节点精准定位表格,或用更高效的CSS选择器(如tbody:has(td:contains("Basic Info")),需确认lxml支持)。 - 原代码中
datalist.pop(0)属于冗余操作,可在遍历表格行时直接跳过第一行,避免先添加再删除的无效步骤。
- 原代码遍历所有
使用会话保持复用连接
每次requests.get()都会新建TCP连接,握手过程额外耗时。改用requests.Session()可复用连接,减少重复握手开销,提升请求效率。只需初始化一次session,绑定headers后用session.get()发起请求即可。批量处理与内存优化
- 先批量收集所有手表链接,再统一发起解析请求,避免边爬边处理的零散操作。
- 解析数据时用生成器替代列表,减少内存占用,爬取大量数据时效果更明显。
启用本地缓存避免重复请求
用requests-cache库对已爬取页面做本地缓存,后续请求相同页面直接读取缓存内容,既节省时间带宽,也降低被网站封禁的风险。
内容的提问来源于stack exchange,提问作者tybrucker

