You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取问题:无法完整获取文本且无法排除删除线文本

解决Selenium爬取的两个问题

问题1:排除<strike>标签的删除线文本

你当前代码存在两个明显问题:

  • 用find_element仅能获取第一个<strike>元素,无法覆盖页面中所有删除线内容
  • 变量名错误(undesired_text未定义,实际应为strikedout_text)

推荐两种可靠解决方案:

方案1:通过JavaScript移除所有<strike>元素

直接修改页面DOM,彻底排除删除线内容后再提取文本:

# 执行JS移除所有strike标签
driver.execute_script("""
    const strikes = document.querySelectorAll('strike');
    strikes.forEach(strike => strike.remove());
""")

# 获取处理后的body文本
all_text = driver.find_element(By.TAG_NAME, 'body').text

方案2:遍历替换所有删除线文本

不修改页面DOM,逐个替换所有<strike>元素的文本内容:

strikedout_elems = driver.find_elements(By.TAG_NAME, 'strike')
all_text = driver.find_element(By.TAG_NAME, 'body').text

for elem in strikedout_elems:
    all_text = all_text.replace(elem.text, '')

# 清理多余换行
wanted_text = all_text.replace('\n\n', '')

问题2:文本截断、无法获取全部内容

文本截断通常源于页面未完全加载(懒加载、动态渲染)或text属性的局限性,可从以下三点优化:

1. 等待页面完全加载

使用WebDriverWait等待关键元素加载,确保内容渲染完毕:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver.get("https://custom.statenet.com/public/resources.cgi?id=ID:bill:KY2022000H740&ciq=ncsl32&client_md=4cf283759f8caf88663f3af9e2707c25&mode=current_text")

# 等待body元素加载,最长等待10秒
WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, 'body')))

2. 滚动页面触发懒加载

若页面存在滚动加载内容,执行JS滚动到页面底部:

# 滚动到页面底部
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
# 等待内容加载(时长可根据页面调整)
time.sleep(2)

3. 用innerHTML提取原始内容

相比text属性,innerHTML能获取更完整的原始内容,结合BeautifulSoup处理:

from bs4 import BeautifulSoup

body_html = driver.find_element(By.TAG_NAME, 'body').get_attribute('innerHTML')
soup = BeautifulSoup(body_html, 'html.parser')

# 移除所有strike标签
for strike in soup.find_all('strike'):
    strike.decompose()

# 提取纯文本并清理
wanted_text = soup.get_text().replace('\n\n', '')

整合后的完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

driver = webdriver.Chrome()
try:
    url = "https://custom.statenet.com/public/resources.cgi?id=ID:bill:KY2022000H740&ciq=ncsl32&client_md=4cf283759f8caf88663f3af9e2707c25&mode=current_text"
    driver.get(url)

    # 等待页面核心元素加载
    WebDriverWait(driver, 15).until(EC.presence_of_element_located((By.TAG_NAME, 'body')))
    # 滚动触发懒加载内容
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(3)

    # 提取并处理页面内容
    body_html = driver.find_element(By.TAG_NAME, 'body').get_attribute('innerHTML')
    soup = BeautifulSoup(body_html, 'html.parser')
    for strike in soup.find_all('strike'):
        strike.decompose()

    wanted_text = soup.get_text().replace('\n\n', '')
    print(wanted_text)
finally:
    driver.quit()

内容的提问来源于stack exchange,提问作者ZZ Top

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 23:55:24