You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium获取网页源码不完整问题求助

解决Selenium爬取sharky.fi/lend仅获部分HTML的问题

问题描述

爬取https://sharky.fi/lend时,使用Selenium的page_source只能获取部分HTML:直接执行代码仅拿到页面顶部内容;滚动到底部后,又仅能获取底部内容,顶部元素从DOM中丢失,无法拿到完整页面源码。

原问题代码

from selenium import webdriver
from webdriver_manager.firefox import GeckoDriverManager
from fake_useragent import UserAgent
import time

url = 'https://sharky.fi/lend'
user_agent = UserAgent().random
options = webdriver.FirefoxOptions()
options.add_argument(f'user-agent={user_agent}')
driver = webdriver.Firefox(executable_path=GeckoDriverManager().install(), options=options)
driver.get(url)
time.sleep(15)
browser_source = driver.page_source
driver.close()
print(browser_source)

问题原因

该网站采用了**虚拟滚动(Virtual Scrolling)**技术,仅渲染当前视口内的内容,当元素滚动出视口后会被从DOM中移除,因此直接获取page_source只能拿到当前浏览器渲染的部分内容。

解决方案

方法1:逐步滚动加载所有内容

通过循环滚动页面,等待新内容加载,直到页面高度不再变化,确保所有元素都被渲染到DOM中:

from selenium import webdriver
from webdriver_manager.firefox import GeckoDriverManager
from fake_useragent import UserAgent
import time

url = 'https://sharky.fi/lend'
user_agent = UserAgent().random
options = webdriver.FirefoxOptions()
options.add_argument(f'user-agent={user_agent}')

driver = webdriver.Firefox(executable_path=GeckoDriverManager().install(), options=options)
driver.get(url)
time.sleep(5)  # 等待初始页面框架加载

# 循环滚动直到内容加载完成
last_scroll_height = driver.execute_script("return document.body.scrollHeight")
while True:
    # 滚动到当前页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待新内容渲染
    time.sleep(3)
    # 获取最新页面高度
    current_scroll_height = driver.execute_script("return document.body.scrollHeight")
    # 高度不变说明没有新内容加载
    if current_scroll_height == last_scroll_height:
        break
    last_scroll_height = current_scroll_height

# 获取完整页面源码
full_page_source = driver.page_source
driver.quit()
print(full_page_source)

方法2:强制禁用虚拟滚动(针对顽固场景)

如果方法1无效,可通过JavaScript修改虚拟滚动容器的样式,强制渲染所有内容:

from selenium import webdriver
from webdriver_manager.firefox import GeckoDriverManager
from fake_useragent import UserAgent
import time

url = 'https://sharky.fi/lend'
user_agent = UserAgent().random
options = webdriver.FirefoxOptions()
options.add_argument(f'user-agent={user_agent}')

driver = webdriver.Firefox(executable_path=GeckoDriverManager().install(), options=options)
driver.get(url)
time.sleep(5)

# 执行JS禁用虚拟滚动容器的高度限制(需根据网站实际DOM结构调整选择器)
driver.execute_script("""
// 查找虚拟滚动容器(示例选择器,需根据实际页面修改)
const scrollContainer = document.querySelector('[data-testid="lend-market-list"]');
if (scrollContainer) {
    scrollContainer.style.height = 'auto';
    scrollContainer.style.overflow = 'visible';
}
""")
time.sleep(3)

# 获取完整源码
full_page_source = driver.page_source
driver.quit()
print(full_page_source)

注意事项

  • 调整time.sleep()的等待时长,根据网站加载速度灵活修改,避免因等待不足导致内容未加载完成。
  • 方法2中的DOM选择器需要根据网站实际结构调整,可通过浏览器开发者工具定位虚拟滚动的容器元素。

内容的提问来源于stack exchange,提问作者Jordan Iliev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 07:25:34