使用Python+BeautifulSoup爬取汽配网站无法获取分类链接如何解决
问题原因
- 重复发起未携带请求头的请求:代码中第二次调用
requests.get(url)未传入headers参数,触发网站反爬策略,返回的页面内容无有效DOM结构,无法定位元素。 - BeautifulSoup API使用错误:
findAll('a', 'href')的写法不符合API规范,该语法是筛选a标签的属性值为href的元素,而非提取a标签的href属性,也无法匹配到任何有效元素。 - 类名筛选语法错误:
find('div', {'class',"category-cell large novehicle"})中第二个参数是集合结构,不是字典结构,无法正确匹配元素的class属性。 - 解析器指定不规范:
BeautifulSoup(response.content, "html")未明确指定解析器,部分环境会抛出警告或解析失败。
修复后的代码
import requests from bs4 import BeautifulSoup url = "https://www.ipdusa.com" headers = { "Accept": "*/*", "User-Agent": "Mozilla/5.0 (iPad; CPU OS 11_0 like Mac OS X) AppleWebKit/604.1.34 (KHTML, like Gecko) Version/11.0 Mobile/15A5341f Safari/604.1" } # 只发起一次带请求头的请求 response = requests.get(url, headers=headers) # 显式指定html.parser解析器 soup = BeautifulSoup(response.content, "html.parser") results = [] # 先匹配分类容器,添加异常捕获避免找不到元素时报错 try: category_container = soup.find('div', class_="category-cell large novehicle") # 查找容器下所有带href属性的a标签 a_tags = category_container.find_all('a', href=True) for a in a_tags: href = a['href'] # 补全相对路径为完整URL full_url = url + href if href.startswith('/') else href results.append(full_url) except AttributeError: print("未找到指定分类容器") print(results)
代码说明
- 所有请求统一携带自定义请求头,规避基础反爬拦截
- 先定位分类容器,再筛选所有带有效href属性的a标签,逐一提取链接并补全为可直接访问的完整路径
- 增加异常捕获逻辑,避免页面结构变动导致的脚本崩溃问题
内容的提问来源于stack exchange,提问作者Kodesenin
相关产品推荐
相关产品推荐

