You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫返回空列表问题排查(疑似HTML类选择错误)

Python爬虫返回空列表问题排查(疑似HTML类选择错误)

兄弟,我看你这爬虫跑出来空列表,核心问题确实出在元素选择器写错了,还有页面等待的方式不够靠谱!我帮你拆解下问题,再给你调整后的可运行代码。

问题出在哪?

  1. 多类名的写法完全错了
    你原代码里用class_='th-h4 text-scarlet'定位价格元素,但BeautifulSoup不认这种空格分隔的写法——元素的多个class是独立属性,你得用列表形式传参:class_=['th-h4', 'text-scarlet'],不然它会把整个字符串当成一个完整的类名去匹配,自然找不到任何元素。

  2. 页面等待太死板
    用time.sleep(30)完全看运气,要是页面加载慢31秒,你还是拿不到数据;加载快的话又浪费时间。换成Selenium的显式等待才是正确姿势——直到目标元素出现再继续执行,既高效又可靠。

  3. 元素定位范围太广
    你直接找所有带aria-label的div,会拿到一堆和公寓无关的元素(比如页面导航、广告的元素),就算找到价格元素,zip的时候也容易因为数量不匹配丢数据。应该先定位每个公寓的父容器,再在容器里找价格和标题,这样更精准。

调整后的代码

我把你的代码修正了这些问题,还加了调试小技巧:

# Importing selenium, CSV, and time
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import TimeoutException
from bs4 import BeautifulSoup
import csv
import time
from webdriver_manager.chrome import ChromeDriverManager

# Running the browser in the background without GPU and Sandbox
chrome_options = Options()
chrome_options.add_argument('--headless')
chrome_options.add_argument('--disable-gpu')
chrome_options.add_argument('--no-sandbox')

# Using Service and CDM to specify the driver path
service = Service(ChromeDriverManager().install())

# Initializing the driver
driver = webdriver.Chrome(service=service, options=chrome_options)

# Opening the developer's URL
print("Opening the page...")
driver.get('https://etalongroup.ru/msk/object/voxhall/')
print("The page is opened.")

# 替换死板的sleep为显式等待:直到公寓列表容器加载完成
wait = WebDriverWait(driver, 30)
try:
    # 这里的定位器可以用F12查看实际页面的公寓列表容器类名,我假设是"flats-list"
    wait.until(EC.presence_of_element_located((By.CLASS_NAME, "flats-list")))
    print("公寓列表已加载完成")
except TimeoutException:
    print("等待公寓列表加载超时,请检查网络或页面结构")
    driver.quit()
    exit()

# 【调试技巧】把爬取到的页面保存成HTML,你可以打开看看实际结构
page_source = driver.page_source
with open('page.html', 'w', encoding='utf-8') as f:
    f.write(page_source)

# Closing the driver
driver.quit()

# Parsing HTML with bs4
soup = BeautifulSoup(page_source, 'html.parser')

# List with apartment data
apartments = []

# 先找到每个公寓的父容器(这里的类名需要你用F12确认,比如实际是"flats-list__item")
apartment_cards = soup.find_all('div', class_='flats-list__item')

# 遍历每个公寓卡片,精准提取数据
for card in apartment_cards:
    # 用列表形式传多类名,匹配价格元素
    price_element = card.find('span', class_=['th-h4', 'text-scarlet'])
    # 在当前卡片内找带aria-label的标题元素
    title_element = card.find('div', {'aria-label': True})
    
    # 确保两个元素都存在再添加数据
    if price_element and title_element:
        price = price_element.text.strip()
        title = title_element['aria-label'].strip()
        apartments.append({'Title': title, 'Price': price})

print(apartments)

# Script completion message
print("The script has finished executing.")

额外调试建议

如果还是拿不到数据,先把chrome_options.add_argument('--headless')注释掉,让浏览器可视化打开,看看页面实际加载后,公寓列表是不是需要滚动才加载?或者有没有反爬验证?另外,打开保存的page.html文件,直接在里面搜索你要找的类名(比如text-scarlet),确认这些元素有没有出现在爬取到的源码里——如果没有,说明数据是通过AJAX动态加载的,可能需要用Selenium模拟滚动,或者直接抓接口。

备注:内容来源于stack exchange,提问作者Danny Mxxre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 14:23:03