Python爬虫返回空列表问题排查(疑似HTML类选择错误)
兄弟,我看你这爬虫跑出来空列表,核心问题确实出在元素选择器写错了,还有页面等待的方式不够靠谱!我帮你拆解下问题,再给你调整后的可运行代码。
问题出在哪?
多类名的写法完全错了
你原代码里用class_='th-h4 text-scarlet'定位价格元素,但BeautifulSoup不认这种空格分隔的写法——元素的多个class是独立属性,你得用列表形式传参:class_=['th-h4', 'text-scarlet'],不然它会把整个字符串当成一个完整的类名去匹配,自然找不到任何元素。页面等待太死板
用time.sleep(30)完全看运气,要是页面加载慢31秒,你还是拿不到数据;加载快的话又浪费时间。换成Selenium的显式等待才是正确姿势——直到目标元素出现再继续执行,既高效又可靠。元素定位范围太广
你直接找所有带aria-label的div,会拿到一堆和公寓无关的元素(比如页面导航、广告的元素),就算找到价格元素,zip的时候也容易因为数量不匹配丢数据。应该先定位每个公寓的父容器,再在容器里找价格和标题,这样更精准。
调整后的代码
我把你的代码修正了这些问题,还加了调试小技巧:
# Importing selenium, CSV, and time from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait from selenium.common.exceptions import TimeoutException from bs4 import BeautifulSoup import csv import time from webdriver_manager.chrome import ChromeDriverManager # Running the browser in the background without GPU and Sandbox chrome_options = Options() chrome_options.add_argument('--headless') chrome_options.add_argument('--disable-gpu') chrome_options.add_argument('--no-sandbox') # Using Service and CDM to specify the driver path service = Service(ChromeDriverManager().install()) # Initializing the driver driver = webdriver.Chrome(service=service, options=chrome_options) # Opening the developer's URL print("Opening the page...") driver.get('https://etalongroup.ru/msk/object/voxhall/') print("The page is opened.") # 替换死板的sleep为显式等待:直到公寓列表容器加载完成 wait = WebDriverWait(driver, 30) try: # 这里的定位器可以用F12查看实际页面的公寓列表容器类名,我假设是"flats-list" wait.until(EC.presence_of_element_located((By.CLASS_NAME, "flats-list"))) print("公寓列表已加载完成") except TimeoutException: print("等待公寓列表加载超时,请检查网络或页面结构") driver.quit() exit() # 【调试技巧】把爬取到的页面保存成HTML,你可以打开看看实际结构 page_source = driver.page_source with open('page.html', 'w', encoding='utf-8') as f: f.write(page_source) # Closing the driver driver.quit() # Parsing HTML with bs4 soup = BeautifulSoup(page_source, 'html.parser') # List with apartment data apartments = [] # 先找到每个公寓的父容器(这里的类名需要你用F12确认,比如实际是"flats-list__item") apartment_cards = soup.find_all('div', class_='flats-list__item') # 遍历每个公寓卡片,精准提取数据 for card in apartment_cards: # 用列表形式传多类名,匹配价格元素 price_element = card.find('span', class_=['th-h4', 'text-scarlet']) # 在当前卡片内找带aria-label的标题元素 title_element = card.find('div', {'aria-label': True}) # 确保两个元素都存在再添加数据 if price_element and title_element: price = price_element.text.strip() title = title_element['aria-label'].strip() apartments.append({'Title': title, 'Price': price}) print(apartments) # Script completion message print("The script has finished executing.")
额外调试建议
如果还是拿不到数据,先把chrome_options.add_argument('--headless')注释掉,让浏览器可视化打开,看看页面实际加载后,公寓列表是不是需要滚动才加载?或者有没有反爬验证?另外,打开保存的page.html文件,直接在里面搜索你要找的类名(比如text-scarlet),确认这些元素有没有出现在爬取到的源码里——如果没有,说明数据是通过AJAX动态加载的,可能需要用Selenium模拟滚动,或者直接抓接口。
备注:内容来源于stack exchange,提问作者Danny Mxxre

