You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+BeautifulSoup4爬取债券表格仅获标签无对应值求助

解决BeautifulSoup爬取法兰克福证券交易所表格数据缺失的问题

问题原因

目标页面的表格数值是通过JavaScript动态渲染的,直接使用requests.get()获取的静态HTML中并不包含这些数据,因此BeautifulSoup无法解析到对应的数值内容。

解决方案

方案1:使用Selenium模拟浏览器加载动态内容

通过模拟浏览器行为,等待页面动态内容加载完成后再解析HTML,确保能获取到完整的表格数据。

步骤:

  1. 安装Selenium库和对应浏览器的驱动(如ChromeDriver,需与浏览器版本匹配)
  2. 使用以下代码实现:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

url = "https://www.boerse-frankfurt.de/bond/xs0216072230"

# 初始化Chrome驱动(需确保ChromeDriver路径正确)
driver = webdriver.Chrome()
driver.get(url)

# 等待目标表格加载完成,超时时间10秒
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "table.table.widget-table:nth-of-type(3)")))

# 获取渲染后的页面源码
page_source = driver.page_source
driver.quit()

# 解析页面内容
soup = BeautifulSoup(page_source, 'html.parser')
target_table = soup.select_one("table.table.widget-table:nth-of-type(3)")
table_body = target_table.find('tbody')
rows = table_body.find_all('tr')

# 提取每行的标签和对应数值
for row in rows:
    cols = row.find_all('td')
    label = cols[0].get_text(strip=True)
    value = cols[1].get_text(strip=True) if len(cols) > 1 else ""
    print(f"{label}: {value}")

方案2:直接调用网站API接口(推荐)

分析页面的网络请求,找到加载债券数据的API接口,直接请求接口获取JSON格式的数据,这种方式更高效且无需模拟浏览器。

示例代码:

import requests

# 目标债券的API接口
api_url = "https://api.boerse-frankfurt.de/api/v1/instruments/XS0216072230/bond"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

response = requests.get(api_url, headers=headers)
bond_data = response.json()

# 提取所需字段(可根据JSON结构调整)
print(f"发行人: {bond_data['issuer']['name']}")
print(f"行业: {bond_data['issuer']['industry']}")
print(f"ISIN编码: {bond_data['isin']}")
print(f"票面利率: {bond_data['interest']['coupon']}")
print(f"到期日: {bond_data['maturityDate']}")

两种方案对比

  • Selenium方法:直观易实现,适合复杂动态页面,但运行速度较慢,需要维护浏览器驱动。
  • API接口方法:速度快、效率高,数据结构清晰,但需要自行分析网络请求找到正确的接口。

内容的提问来源于stack exchange,提问作者Yogita Negi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 23:33:24