You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python动态HTML标签取值异常:Selenium爬取未获预期数字结果

解决动态HTML标签值获取问题(从短横线到目标数字)

你遇到的核心问题是:目标标签初始渲染为短横线-,实际数值由JavaScript动态加载,直接用BeautifulSoup解析页面源码拿到的是未更新的初始值,而非网页上显示的真实数字。

方案1:用Selenium直接定位元素获取文本(推荐)

跳过BeautifulSoup中转,直接用Selenium的元素定位方法获取动态渲染后的文本,同时用显式等待替代固定睡眠,确保元素加载完成后再操作,可靠性更高。

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def scrape_content_from_dynamic_websites():
    url = "https://statusinvest.com.br/acoes/petr4/"
    driver = webdriver.Chrome()
    driver.get(url)
    
    # 最多等待10秒,直到目标元素出现
    target_element = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "strong[data-item='avg_F']"))
    )
    # 获取元素渲染后的文本
    result = target_element.text
    driver.close()
    return result

def main09():
    print(scrape_content_from_dynamic_websites())

if __name__ == "__main__":
    main09()

方案2:保留BeautifulSoup,确保获取渲染后的源码

如果坚持用BeautifulSoup解析,需要确保页面完全渲染,可通过滚动触发JS加载,再配合显式等待:

from selenium import webdriver
import time
import bs4
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def scrape_content_from_dynamic_websites():
    url = "https://statusinvest.com.br/acoes/petr4/"
    driver = webdriver.Chrome()
    driver.get(url)
    
    # 等待目标元素加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "strong[data-item='avg_F']"))
    )
    # 滚动到元素位置,触发JS渲染
    target_element = driver.find_element(By.CSS_SELECTOR, "strong[data-item='avg_F']")
    driver.execute_script("arguments[0].scrollIntoView();", target_element)
    time.sleep(1)  # 给渲染留一点缓冲时间
    
    html = driver.page_source
    soup = bs4.BeautifulSoup(html, "html.parser")
    # 直接获取标签文本,而非标签对象
    result = soup.find("strong", {"data-item":"avg_F"}).text
    driver.close()

    return result

def main09():
    print(scrape_content_from_dynamic_websites())

if __name__ == "__main__":
    main09()

关键注意点

  • 固定time.sleep(5)不可靠,网络或设备性能差异会导致页面未完全加载,显式等待能精准等待目标元素就绪。
  • 动态渲染的内容,直接用Selenium获取元素文本比解析HTML源码更直接,避免初始DOM和渲染后DOM不一致的问题。

内容的提问来源于stack exchange,提问作者Renato Sousa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 17:40:19