使用Requests+BeautifulSoup爬取G2网站产品标题失败求助
解决G2营销自动化分类页产品标题爬取失败问题
问题原因分析
- 请求头缺失被反爬拦截:直接使用
requests.get发送请求时,默认请求头会被识别为爬虫,返回的HTML不含目标产品卡片内容。 - 动态内容未渲染:页面产品列表由JavaScript动态加载,
requests获取的是初始静态HTML,无法获取渲染后的元素。 - CSS类名变更:网站可能更新了元素类名,原
product-card__product-name可能已失效。
解决方案
1. 完善请求头(Requests+BeautifulSoup)
添加模拟浏览器的请求头,绕过基础反爬:
import requests from bs4 import BeautifulSoup url = 'https://www.g2.com/categories/marketing-automation' # 模拟Chrome浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9' } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 先检查页面是否存在目标类名,或用其他定位方式(如标签+属性) product_names = soup.find_all(class_='product-card__product-name') # 输出标题文本 print([name.get_text(strip=True) for name in product_names])
2. 使用Selenium等待动态加载
若内容完全由JS渲染,用Selenium等待元素加载完成:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = 'https://www.g2.com/categories/marketing-automation' # 初始化Chrome浏览器 driver = webdriver.Chrome() driver.get(url) # 等待产品标题元素加载,超时时间10秒 wait = WebDriverWait(driver, 10) product_elements = wait.until( EC.presence_of_all_elements_located((By.CLASS_NAME, 'product-card__product-name')) ) # 提取并打印标题 for elem in product_elements: print(elem.text.strip()) driver.quit()
3. 额外注意事项
- 若仍无结果,打开浏览器开发者工具(F12),检查当前页面产品标题的实际CSS类名,可能已更新。
- G2部分内容需登录才能查看完整列表,若爬取全部350+产品,可能需要模拟登录或处理Cookie。
内容的提问来源于stack exchange,提问作者Confused_Programmer
相关产品推荐
相关产品推荐

