Selenium WebDriver元素定位失效,Python爬虫代码求助
解决Selenium爬取JLL页面PDF链接返回空列表的问题
问题说明
原本正常运行的Python/Selenium爬取代码突然失效,执行后返回空列表。需求为访问指定JLL物业页面,提取其中的PDF下载链接,但使用class="pt-res-link"定位元素无法可靠获取目标内容。
原代码
driver_service = Service(executable_path="C:\\WPy64-39100\\chromedriver.exe") chrome_options = Options() chrome_options.add_experimental_option("detach", True) chrome_options.add_argument("--headless") driver = webdriver.Chrome(service = driver_service, options=chrome_options) site = 'https://powersearch.jll.com/ca-en/property/52770/centurion-plaza-10335-172-street' driver.get(site) time.sleep(10) elements = driver.find_elements(By.CLASS_NAME, 'pt-res-link') links = [e.get_attribute("href") for e in elements] print(links)
技术建议
1. 替换固定等待为显式等待
硬编码的time.sleep()无法适配页面动态加载的波动,改用WebDriverWait等待目标元素出现,确保元素加载完成后再进行定位:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 替换time.sleep(10)为: wait = WebDriverWait(driver, 20) elements = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'pt-res-link')))
2. 检查DOM结构变更,更换定位方式
网站可能更新了页面结构,pt-res-link类名可能被修改或元素层级变化。可以通过浏览器开发者工具重新确认元素属性:
- 若类名失效,改用CSS选择器定位包含PDF的链接:
elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'a[href$=".pdf"]'))) - 或用XPath定位带有特定关键词的链接:
elements = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//a[contains(@href, "brochure") and contains(@href, ".pdf")]')))
3. 优化无头浏览器配置
部分网站会检测无头模式,添加以下参数模拟正常浏览器:
chrome_options.add_argument("--window-size=1920,1080") chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") chrome_options.add_argument("--disable-blink-features=AutomationControlled")
4. 验证页面加载状态
若仍无法定位,可在代码中添加页面源码打印或截图,确认目标元素是否存在:
# 打印页面源码 print(driver.page_source) # 截图保存 driver.save_screenshot("page.png")
修改后的示例代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver_service = Service(executable_path="C:\\WPy64-39100\\chromedriver.exe") chrome_options = Options() chrome_options.add_experimental_option("detach", True) chrome_options.add_argument("--headless") # 添加模拟正常浏览器的参数 chrome_options.add_argument("--window-size=1920,1080") chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") chrome_options.add_argument("--disable-blink-features=AutomationControlled") driver = webdriver.Chrome(service=driver_service, options=chrome_options) site = 'https://powersearch.jll.com/ca-en/property/52770/centurion-plaza-10335-172-street' driver.get(site) try: # 显式等待PDF链接元素加载 wait = WebDriverWait(driver, 20) # 用CSS选择器定位所有PDF链接 elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'a[href$=".pdf"]'))) links = [e.get_attribute("href") for e in elements] print(links) finally: driver.quit()
内容的提问来源于stack exchange,提问作者SMS
相关产品推荐
相关产品推荐

