You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas提取HTML表格中td标签的data-custom-date属性值

解决Pandas抓取网页表格时提取标签属性的问题

问题分析

pd.read_html() 仅能提取HTML表格单元格内的文本内容,无法保留标签的属性(比如data-custom-date),因此需要结合BeautifulSoup解析原始HTML,手动提取所需的文本和属性值。

解决方案代码

import pandas as pd
import requests
import pprint
from bs4 import BeautifulSoup

url = "https://www.myfxbook.com/forex-economic-calendar/interest-rates"

r = requests.get(url)
soup = BeautifulSoup(r.text, "html.parser")

# 定位目标表格(若页面存在多个表格,可通过class/id精准定位)
target_table = soup.find("table")

data = []
# 遍历表格行(跳过表头行)
for row in target_table.find_all("tr")[1:]:
    cells = row.find_all("td")
    row_dict = {
        "Country": cells[0].get_text(strip=True),
        "Unnamed: 1": cells[1].get_text(strip=True),
        "Central Bank": cells[2].get_text(strip=True),
        "Current Rate": cells[3].get_text(strip=True),
        "Previous Rate": cells[4].get_text(strip=True),
        "Change": cells[5].get_text(strip=True),
        "Last Meeting": cells[6].get_text(strip=True),
        # 提取data-custom-date属性值替换原文本
        "Next Meeting": cells[7].get("data-custom-date")
    }
    data.append(row_dict)

pp = pprint.PrettyPrinter(depth=4)
pp.pprint(data)

代码说明

  • HTML解析:用BeautifulSoup定位目标表格,若页面有多个表格,可通过class或id缩小范围(例如soup.find("table", class_="target-table-class"))。
  • 属性提取:通过cells[7].get("data-custom-date")直接获取对应<td>标签的属性值,替换原本的"1 day"文本。
  • 文本清理:get_text(strip=True)用于去除文本前后的空格、换行符,让数据更整洁。

优化建议

  • 若表格结构可能变动,可先提取表头文本,通过表头匹配列的位置,避免硬编码索引。
  • 添加异常处理逻辑,防止因页面结构变更导致的索引越界错误。

内容的提问来源于stack exchange,提问作者30ThreeDegrees

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 08:25:32