You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Requests+BeautifulSoup爬取G2网站产品标题失败求助

解决G2营销自动化分类页产品标题爬取失败问题

问题原因分析

  • 请求头缺失被反爬拦截:直接使用requests.get发送请求时,默认请求头会被识别为爬虫,返回的HTML不含目标产品卡片内容。
  • 动态内容未渲染:页面产品列表由JavaScript动态加载,requests获取的是初始静态HTML,无法获取渲染后的元素。
  • CSS类名变更:网站可能更新了元素类名,原product-card__product-name可能已失效。

解决方案

1. 完善请求头(Requests+BeautifulSoup)

添加模拟浏览器的请求头,绕过基础反爬:

import requests
from bs4 import BeautifulSoup

url = 'https://www.g2.com/categories/marketing-automation'
# 模拟Chrome浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept-Language': 'en-US,en;q=0.9'
}

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

# 先检查页面是否存在目标类名,或用其他定位方式(如标签+属性)
product_names = soup.find_all(class_='product-card__product-name')
# 输出标题文本
print([name.get_text(strip=True) for name in product_names])

2. 使用Selenium等待动态加载

若内容完全由JS渲染,用Selenium等待元素加载完成:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = 'https://www.g2.com/categories/marketing-automation'
# 初始化Chrome浏览器
driver = webdriver.Chrome()
driver.get(url)

# 等待产品标题元素加载,超时时间10秒
wait = WebDriverWait(driver, 10)
product_elements = wait.until(
    EC.presence_of_all_elements_located((By.CLASS_NAME, 'product-card__product-name'))
)

# 提取并打印标题
for elem in product_elements:
    print(elem.text.strip())

driver.quit()

3. 额外注意事项

  • 若仍无结果,打开浏览器开发者工具(F12),检查当前页面产品标题的实际CSS类名,可能已更新。
  • G2部分内容需登录才能查看完整列表,若爬取全部350+产品,可能需要模拟登录或处理Cookie。

内容的提问来源于stack exchange,提问作者Confused_Programmer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 00:04:01