网页爬取问题:指定class的div无法被BeautifulSoup捕获
问题:无法定位指定class的div元素
我尝试爬取带有如下属性的div元素:
<div data-v-28872a74="" class="col-lg-10 col-md-10 col-sm-12 col-12 offset-lg-1 offset-md-1 offset-sm-0 offset-0">
使用BeautifulSoup的find_all('div', class_ = 'col-lg-10 col-md-10 col-sm-12 col-12 offset-lg-1 offset-md-1 offset-sm-0 offset-0')调用后,返回空列表[]。
第一份测试代码(直接用requests请求)
import requests from bs4 import BeautifulSoup as bs url = 'https://remart.az/yasayis-kompleksi?cities=1&districts=' result = requests.get(url) soup = bs(result.text, 'html.parser') code= soup.find_all('div', class_ = 'col-lg-10 col-md-10 col-sm-12 col-12 offset-lg-1 offset-md-1 offset-sm-0 offset-0') print(code)
第二份测试代码(Selenium获取列表URL,但详情页爬取仍失败)
先获取列表页的详情URL:
driver = webdriver.Chrome(r'C:\Program Files (x86)\chromedriver_win32\chromedriver.exe') driver.get('https://remart.az/yasayis-kompleksi?cities=1&districts=') time.sleep(3) aze = driver.find_element(By.XPATH, '//*[@id="app"]/div[2]/div[1]/div[2]/div[6]/button') for a in range(1,2): aze.click() time.sleep(1) soup = bs(driver.page_source, "html.parser") aezexx = soup.find_all('div', class_ = 'bitem') for parent in aezexx: a_tag = parent.find("a") URRL = a_tag.attrs['href'] print(URRL)
后续爬取详情页时,同样出现元素定位失败:
soup = bs(driver.page_source, "html.parser") aezexx = soup.find_all('div', class_ = 'bitem') for parent in aezexx: a_tag = parent.find("a") URRL = a_tag.attrs['href'] result = requests.get(URRL) soup = bs(result.text, 'html.parser') are = soup.find_all("div", class_ = 'bottom-panel-descripton cut-text') for aes in are: azzzz = aes.find_all('p') print(azzzz)
解决方案
1. 问题本质
- 直接用
requests请求只能拿到服务端返回的静态HTML骨架,目标元素是前端通过JavaScript动态渲染生成的,静态页面里根本没有这些元素,所以find_all找不到内容。 - 详情页用
requests请求也会踩同样的坑,因为详情页的核心内容同样是JS动态加载的。
2. 具体解决办法
办法一:全程用Selenium渲染页面
不管列表页还是详情页,都用Selenium加载页面(确保JS执行完毕)后再解析,这样能拿到完整的渲染后页面:
from selenium import webdriver from selenium.webdriver.common.by import By from bs4 import BeautifulSoup as bs import time # 初始化Chrome浏览器 driver = webdriver.Chrome(r'C:\Program Files (x86)\chromedriver_win32\chromedriver.exe') # 打开列表页 driver.get('https://remart.az/yasayis-kompleksi?cities=1&districts=') time.sleep(3) # 等待页面加载完成 # 点击加载更多按钮(按需执行) aze = driver.find_element(By.XPATH, '//*[@id="app"]/div[2]/div[1]/div[2]/div[6]/button') for _ in range(1): # 这里控制点击次数,原代码是1次 aze.click() time.sleep(1) # 解析列表页,提取所有详情页URL soup = bs(driver.page_source, "html.parser") detail_urls = [item.find("a")["href"] for item in soup.find_all('div', class_='bitem')] # 遍历每个详情页,用Selenium加载后解析 for url in detail_urls: driver.get(url) time.sleep(2) # 等待详情页加载 soup = bs(driver.page_source, "html.parser") # 定位目标元素 descriptions = soup.find_all("div", class_='bottom-panel-descripton cut-text') for desc in descriptions: paragraphs = desc.find_all('p') print(paragraphs) # 关闭浏览器 driver.quit()
办法二:简化class匹配规则
不需要完全匹配所有class属性值,挑一个唯一且稳定的class片段来定位,比如:
# 只匹配关键class,避免空格或顺序问题 code = soup.find_all('div', class_='col-lg-10 offset-lg-1') # 或者结合属性过滤,更精准 code = soup.find_all('div', {'data-v-28872a74': '', 'class': lambda cls: 'col-lg-10' in cls})
办法三:直接请求后端API(最优解)
打开浏览器开发者工具的「网络」面板,筛选XHR/fetch请求,找到页面加载数据的接口(比如类似/api/projects的请求),直接请求这些API获取JSON格式的数据,比解析HTML效率高得多,也避免了JS渲染的问题。
内容的提问来源于stack exchange,提问作者Rasim Dilbani
相关产品推荐
相关产品推荐

