使用Headless Selenium抓取动态网站:为何无法获取脚本生成内容?
问题分析与解决方案
你的问题核心在于无头Chrome被网站检测或页面未完全加载就获取源码,导致无法拿到动态渲染的有效HTML。以下是针对性解决方法:
1. 等待页面元素加载完成
直接调用driver.get(url)后立即获取page_source,大概率页面还在动态渲染数据,此时拿到的是未加载完成的无效HTML。改用显式等待,直到目标元素出现再继续执行。
2. 伪装无头Chrome规避检测
很多网站会通过浏览器特征识别无头模式,需要添加配置参数让无头Chrome更接近正常浏览器。
修改后的完整代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url='https://www.mesitis.com.cy/Search.aspx?&isAsc=0&isRent=0&districts=Lefkosia&status=1&apartFloor=&types=0dfe6c47-04c5-e511-ae61-a4badb3ceace&refno=&priceFrom=170000&priceTo=260000&priceRentFrom=0&priceRentTo=10000&areaFrom=0&areaTo=1000&intAreaFrom=130&intAreaTo=500&densityFrom=0&densityTo=200¤tPage=1' options = Options() options.headless = True # 添加伪装参数,模拟正常浏览器 options.add_argument("--window-size=1920,1080") options.add_argument("--disable-blink-features=AutomationControlled") options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(options=options, executable_path='C:\Downloads\chromedriver_win32\chromedriver.exe') # 移除自动化标识 driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") driver.get(url) # 显式等待目标元素加载,最长等待10秒 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "properties-list")) ) finally: page_source = driver.page_source driver.quit() soup = BeautifulSoup(page_source,"html.parser") properties = soup.find('div', attrs={'class':'row properties-list'}) if properties: allinks = properties.find_all('h3') for d in allinks: if d.a: print('https://www.mesitis.com.cy/'+d.a.get('href')) else: print("未找到目标元素,请检查页面加载情况或选择器是否正确")
额外注意事项
- 确保
chromedriver版本与你的Chrome浏览器版本完全匹配,版本不兼容会导致渲染异常。 - 若网站反爬机制较强,可额外添加随机延迟、代理等策略,但当前场景下上述修改足以解决问题。
内容的提问来源于stack exchange,提问作者Giganoob
相关产品推荐
相关产品推荐

