使用Beautiful Soup 4爬取Lochinvar热水器规格表格失败求助
问题核心
目标页面的规格表格是通过JavaScript动态渲染生成的。requests.get()仅能获取页面初始静态HTML,此时表格内容尚未被浏览器渲染,因此BeautifulSoup无法定位到那些动态生成的类名(如Table__Cell-sc-1e0v68l-0 kdksLO)。
解决方案1:直接抓取后端API(推荐)
通过浏览器开发者工具抓包,可找到页面加载表格数据的后端API接口,直接请求接口获取结构化数据,效率远高于模拟浏览器。
操作步骤:
- 打开目标页面,按F12唤起开发者工具,切换至
Network标签 - 刷新页面,筛选
XHR/Fetch类型请求,定位返回规格数据的接口(通常路径包含spec、product等关键词) - 复制接口URL,用
requests直接请求并解析JSON数据
示例代码(需替换为实际抓包得到的API地址):
import requests API_URL = "实际抓包获取的API接口地址" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(API_URL, headers=headers) spec_data = response.json() # 解析JSON提取规格信息 for spec_item in spec_data.get("specifications", []): print(f"{spec_item['label']}: {spec_item['value']}")
解决方案2:使用Selenium模拟浏览器渲染
若无法找到API接口,可通过Selenium模拟真实浏览器加载页面,等待JS渲染完成后提取数据。
示例代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup URL = "https://www.lochinvar.com/products/commercial-water-heaters/armor-condensing-water-heater" # 初始化Chrome浏览器(需提前安装ChromeDriver并配置环境变量) driver = webdriver.Chrome() driver.get(URL) # 等待表格元素加载完成 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CLASS_NAME, "Table__Wrapper-sc-1e0v68l-3"))) # 获取渲染后的页面源码 page_source = driver.page_source driver.quit() # 解析提取表格数据 soup = BeautifulSoup(page_source, "html.parser") table_rows = soup.find_all("div", class_="Table__Row-sc-1e0v68l-2") for row in table_rows: cells = row.find_all("div", class_="Table__Cell-sc-1e0v68l-0") if len(cells) == 2: print(f"{cells[0].text.strip()}: {cells[1].text.strip()}")
内容的提问来源于stack exchange,提问作者John Provost
相关产品推荐
相关产品推荐

