You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Selenium平滑精准滚动到DOM指定区域并提取文本

问题描述

尝试用Selenium滚动到网页指定区域并提取文本,目标区域为DOM中的“Production / Artist”板块。网页通过user-select: none和-webkit-user-select: none禁用文本选中,已可通过JS解除该限制,但核心问题是滚动行为不稳定,无法精准定位到目标区域。现有代码如下:

from selenium import webdriver
from selenium.webdriver.common.by import By

# Initialize WebDriver
driver = webdriver.Chrome()

# Open the URL
url = "https://www.art-mate.net/doc/78492?name=%E6%A8%82%E3%83%BB%E8%AA%BC%E7%8D%A8%E5%A5%8F%E5%AE%B6%E6%A8%82%E5%9C%98%E2%94%80%E2%94%80%E5%A4%A7%E6%8F%90%E7%90%B4%E8%88%87%E9%A6%AC%E7%89%B9%E8%AB%BE%E7%90%B4%E3"
driver.get(url)

# Scroll to the "Production / Artist" section
element = driver.find_element(By.XPATH, "//h2[text()='Production / Artist']")
driver.execute_script("arguments[0].scrollIntoView();", element)

# Now attempt to copy the text from the section
production_artist_section = driver.find_element(By.XPATH, "//div[contains(text(), 'Production / Artist')]")
print(production_artist_section.text)

# Close the driver
driver.quit()

遇到的问题:能识别目标元素,但滚动不流畅、不可靠,偶尔无法精准滚动到指定区域。

解决方案

1. 优化滚动参数,提升精准度

scrollIntoView()默认行为可能导致元素贴边,添加平滑滚动和居中对齐参数,让滚动更稳定且元素定位更准确:

driver.execute_script("arguments[0].scrollIntoView({behavior: 'smooth', block: 'center'});", element)

2. 等待元素完全可见后再滚动

页面加载存在延迟,用显式等待确保目标元素渲染完成后再执行滚动,避免滚动操作提前触发:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 最长等待15秒,直到目标h2元素可见
element = WebDriverWait(driver, 15).until(
    EC.visibility_of_element_located((By.XPATH, "//h2[text()='Production / Artist']"))
)

3. 修复文本提取的定位逻辑

原代码中提取文本的XPath不准确,应基于目标h2元素定位其相邻的内容容器:

# 找到h2元素后,获取其后续的兄弟div内容块
production_artist_section = element.find_element(By.XPATH, "./following-sibling::div")
print(production_artist_section.text)

4. 滚动后增加短暂等待

滚动完成后给页面预留渲染时间,避免因页面未稳定导致文本提取失败:

driver.implicitly_wait(2)  # 等待2秒让页面稳定

5. 全局解除文本选中限制(可选)

如果提取文本时仍受user-select属性影响,执行JS批量解除限制:

driver.execute_script("""
document.body.style.userSelect = 'auto';
document.body.style.webkitUserSelect = 'auto';
document.querySelectorAll('*').forEach(el => {
    el.style.userSelect = 'auto';
    el.style.webkitUserSelect = 'auto';
});
""")
完整优化代码
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Initialize WebDriver
driver = webdriver.Chrome()
driver.maximize_window()  # 最大化窗口减少滚动偏差

# Open the URL
url = "https://www.art-mate.net/doc/78492?name=%E6%A8%82%E3%83%BB%E8%AA%BC%E7%8D%A8%E5%A5%8F%E5%AE%B6%E6%A8%82%E5%9C%98%E2%94%80%E2%94%80%E5%A4%A7%E6%8F%90%E7%90%B4%E8%88%87%E9%A6%AC%E7%89%B9%E8%AB%BE%E7%90%B4%E3"
driver.get(url)

try:
    # 等待目标元素可见
    element = WebDriverWait(driver, 15).until(
        EC.visibility_of_element_located((By.XPATH, "//h2[text()='Production / Artist']"))
    )
    
    # 平滑滚动到元素居中位置
    driver.execute_script("arguments[0].scrollIntoView({behavior: 'smooth', block: 'center'});", element)
    
    # 等待页面稳定
    driver.implicitly_wait(2)
    
    # 解除文本选中限制
    driver.execute_script("""
    document.body.style.userSelect = 'auto';
    document.body.style.webkitUserSelect = 'auto';
    document.querySelectorAll('*').forEach(el => {
        el.style.userSelect = 'auto';
        el.style.webkitUserSelect = 'auto';
    });
    """)
    
    # 提取目标区域文本
    production_artist_section = element.find_element(By.XPATH, "./following-sibling::div")
    print("提取的文本内容:")
    print(production_artist_section.text)

finally:
    # Close the driver
    driver.quit()

内容的提问来源于stack exchange,提问作者poe trenton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 02:20:10