Selenium爬取问题:无法完整获取文本且无法排除删除线文本
解决Selenium爬取的两个问题
问题1:排除<strike>标签的删除线文本
你当前代码存在两个明显问题:
- 用
find_element仅能获取第一个<strike>元素,无法覆盖页面中所有删除线内容 - 变量名错误(
undesired_text未定义,实际应为strikedout_text)
推荐两种可靠解决方案:
方案1:通过JavaScript移除所有<strike>元素
直接修改页面DOM,彻底排除删除线内容后再提取文本:
# 执行JS移除所有strike标签 driver.execute_script(""" const strikes = document.querySelectorAll('strike'); strikes.forEach(strike => strike.remove()); """) # 获取处理后的body文本 all_text = driver.find_element(By.TAG_NAME, 'body').text
方案2:遍历替换所有删除线文本
不修改页面DOM,逐个替换所有<strike>元素的文本内容:
strikedout_elems = driver.find_elements(By.TAG_NAME, 'strike') all_text = driver.find_element(By.TAG_NAME, 'body').text for elem in strikedout_elems: all_text = all_text.replace(elem.text, '') # 清理多余换行 wanted_text = all_text.replace('\n\n', '')
问题2:文本截断、无法获取全部内容
文本截断通常源于页面未完全加载(懒加载、动态渲染)或text属性的局限性,可从以下三点优化:
1. 等待页面完全加载
使用WebDriverWait等待关键元素加载,确保内容渲染完毕:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver.get("https://custom.statenet.com/public/resources.cgi?id=ID:bill:KY2022000H740&ciq=ncsl32&client_md=4cf283759f8caf88663f3af9e2707c25&mode=current_text") # 等待body元素加载,最长等待10秒 WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, 'body')))
2. 滚动页面触发懒加载
若页面存在滚动加载内容,执行JS滚动到页面底部:
# 滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待内容加载(时长可根据页面调整) time.sleep(2)
3. 用innerHTML提取原始内容
相比text属性,innerHTML能获取更完整的原始内容,结合BeautifulSoup处理:
from bs4 import BeautifulSoup body_html = driver.find_element(By.TAG_NAME, 'body').get_attribute('innerHTML') soup = BeautifulSoup(body_html, 'html.parser') # 移除所有strike标签 for strike in soup.find_all('strike'): strike.decompose() # 提取纯文本并清理 wanted_text = soup.get_text().replace('\n\n', '')
整合后的完整代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time driver = webdriver.Chrome() try: url = "https://custom.statenet.com/public/resources.cgi?id=ID:bill:KY2022000H740&ciq=ncsl32&client_md=4cf283759f8caf88663f3af9e2707c25&mode=current_text" driver.get(url) # 等待页面核心元素加载 WebDriverWait(driver, 15).until(EC.presence_of_element_located((By.TAG_NAME, 'body'))) # 滚动触发懒加载内容 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(3) # 提取并处理页面内容 body_html = driver.find_element(By.TAG_NAME, 'body').get_attribute('innerHTML') soup = BeautifulSoup(body_html, 'html.parser') for strike in soup.find_all('strike'): strike.decompose() wanted_text = soup.get_text().replace('\n\n', '') print(wanted_text) finally: driver.quit()
内容的提问来源于stack exchange,提问作者ZZ Top
相关产品推荐
相关产品推荐

