You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium或BeautifulSoup爬取实时金融数据为何返回空列表?

问题根因与修复方案

导致返回空列表的三个主要原因

  • 目标<td>元素全部归属页面内嵌的iframe框架,Selenium默认读取顶层页面DOM,未主动切入iframe的情况下,顶层DOM不存在对应节点,自然匹配不到内容。
  • BeautifulSoup使用html.parser解析HTML时会自动将所有标签名转为小写,代码中使用大写TD作为查找参数,无法匹配到解析后生成的小写<td>节点。
  • 固定3秒硬等待鲁棒性极差,受网络波动、服务端响应速度影响,很容易出现等待时间结束后iframe及内部节点尚未完成渲染的情况,提前读取页面源码也会导致匹配为空。

可直接运行的修复代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import os

driver = webdriver.Chrome(os.getcwd() + '/chromedriver')
driver.get("https://www.amarkets.com/research-education/sentiment-widget/")

# 显式等待iframe加载完成后切入iframe上下文
wait = WebDriverWait(driver, 10)
target_iframe = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "iframe[src*='sentiment']")))
driver.switch_to.frame(target_iframe)

# 等待iframe内部td标签完成渲染
wait.until(EC.presence_of_element_located((By.TAG_NAME, "td")))

real_soup = BeautifulSoup(driver.page_source, 'html.parser')
# 使用小写标签名查找
sentiments = real_soup.find_all("td")
print(sentiments)

# 操作完成后切回顶层DOM
driver.switch_to.default_content()
driver.quit()

内容的提问来源于stack exchange,提问作者Bayrem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 16:24:08