使用BeautifulSoup无法获取class为listing_LinkedListingCard__5SRvZ的div元素求助
解决方法
1. 添加请求头模拟浏览器访问
网站可能会拦截非浏览器发起的请求,导致返回的页面结构缺失目标元素。给requests.get添加headers参数模拟浏览器请求:
import requests from bs4 import BeautifulSoup headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } web_page = requests.get("https://sa.aqar.fm/%D9%81%D9%84%D9%84-%D9%84%D9%84%D8%A8%D9%8A%D8%B9/%D8%A8%D8%B1%D9%8A%D8%AF%D8%A9", headers=headers) def main(page): src = page.content soup = BeautifulSoup(src, 'lxml') houses = soup.find_all("div", class_='listing_LinkedListingCard__5SRvZ') print(f"找到{len(houses)}个房屋卡片") main(web_page)
2. 处理动态渲染内容
如果添加请求头后仍无法获取元素,说明页面内容是通过JavaScript动态加载的,BeautifulSoup仅能解析静态HTML。此时需要使用支持执行JS的工具(如Selenium):
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import time options = Options() options.add_argument('--headless=new') # 无头模式运行 options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=options) driver.get("https://sa.aqar.fm/%D9%81%D9%84%D9%84-%D9%84%D9%84%D8%A8%D9%8A%D8%B9/%D8%A8%D8%B1%D9%8A%D8%AF%D8%A9") time.sleep(3) # 等待页面加载完成 src = driver.page_source soup = BeautifulSoup(src, 'lxml') houses = soup.find_all("div", class_='listing_LinkedListingCard__5SRvZ') print(f"找到{len(houses)}个房屋卡片") driver.quit()
3. 验证并调整元素选择器
部分网站的class名称会动态生成(带随机后缀),可通过以下方式匹配:
- 模糊匹配class:
soup.find_all("div", class_=lambda x: x and 'listing_LinkedListingCard' in x) - 使用元素的其他属性定位(如
data-testid等,需查看页面实际元素属性):soup.find_all("div", attrs={"data-testid": "listing-card"})
内容的提问来源于stack exchange,提问作者محمد
相关产品推荐
相关产品推荐

