You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Headless Selenium抓取动态网站:为何无法获取脚本生成内容?

问题分析与解决方案

你的问题核心在于无头Chrome被网站检测或页面未完全加载就获取源码,导致无法拿到动态渲染的有效HTML。以下是针对性解决方法:

1. 等待页面元素加载完成

直接调用driver.get(url)后立即获取page_source,大概率页面还在动态渲染数据,此时拿到的是未加载完成的无效HTML。改用显式等待,直到目标元素出现再继续执行。

2. 伪装无头Chrome规避检测

很多网站会通过浏览器特征识别无头模式,需要添加配置参数让无头Chrome更接近正常浏览器。

修改后的完整代码

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url='https://www.mesitis.com.cy/Search.aspx?&isAsc=0&isRent=0&districts=Lefkosia&status=1&apartFloor=&types=0dfe6c47-04c5-e511-ae61-a4badb3ceace&refno=&priceFrom=170000&priceTo=260000&priceRentFrom=0&priceRentTo=10000&areaFrom=0&areaTo=1000&intAreaFrom=130&intAreaTo=500&densityFrom=0&densityTo=200&currentPage=1'

options = Options()
options.headless = True
# 添加伪装参数,模拟正常浏览器
options.add_argument("--window-size=1920,1080")
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)

driver = webdriver.Chrome(options=options, executable_path='C:\Downloads\chromedriver_win32\chromedriver.exe')
# 移除自动化标识
driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")

driver.get(url)

# 显式等待目标元素加载,最长等待10秒
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "properties-list"))
    )
finally:
    page_source = driver.page_source
    driver.quit()

soup = BeautifulSoup(page_source,"html.parser")
properties = soup.find('div', attrs={'class':'row properties-list'})

if properties:
    allinks = properties.find_all('h3')
    for d in allinks:
        if d.a:
            print('https://www.mesitis.com.cy/'+d.a.get('href'))
else:
    print("未找到目标元素,请检查页面加载情况或选择器是否正确")

额外注意事项

  • 确保chromedriver版本与你的Chrome浏览器版本完全匹配,版本不兼容会导致渲染异常。
  • 若网站反爬机制较强,可额外添加随机延迟、代理等策略,但当前场景下上述修改足以解决问题。

内容的提问来源于stack exchange,提问作者Giganoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 22:15:36