如何用Pandas提取HTML表格中td标签的data-custom-date属性值
解决Pandas抓取网页表格时提取标签属性的问题
问题分析
pd.read_html() 仅能提取HTML表格单元格内的文本内容,无法保留标签的属性(比如data-custom-date),因此需要结合BeautifulSoup解析原始HTML,手动提取所需的文本和属性值。
解决方案代码
import pandas as pd import requests import pprint from bs4 import BeautifulSoup url = "https://www.myfxbook.com/forex-economic-calendar/interest-rates" r = requests.get(url) soup = BeautifulSoup(r.text, "html.parser") # 定位目标表格(若页面存在多个表格,可通过class/id精准定位) target_table = soup.find("table") data = [] # 遍历表格行(跳过表头行) for row in target_table.find_all("tr")[1:]: cells = row.find_all("td") row_dict = { "Country": cells[0].get_text(strip=True), "Unnamed: 1": cells[1].get_text(strip=True), "Central Bank": cells[2].get_text(strip=True), "Current Rate": cells[3].get_text(strip=True), "Previous Rate": cells[4].get_text(strip=True), "Change": cells[5].get_text(strip=True), "Last Meeting": cells[6].get_text(strip=True), # 提取data-custom-date属性值替换原文本 "Next Meeting": cells[7].get("data-custom-date") } data.append(row_dict) pp = pprint.PrettyPrinter(depth=4) pp.pprint(data)
代码说明
- HTML解析:用BeautifulSoup定位目标表格,若页面有多个表格,可通过
class或id缩小范围(例如soup.find("table", class_="target-table-class"))。 - 属性提取:通过
cells[7].get("data-custom-date")直接获取对应<td>标签的属性值,替换原本的"1 day"文本。 - 文本清理:
get_text(strip=True)用于去除文本前后的空格、换行符,让数据更整洁。
优化建议
- 若表格结构可能变动,可先提取表头文本,通过表头匹配列的位置,避免硬编码索引。
- 添加异常处理逻辑,防止因页面结构变更导致的索引越界错误。
内容的提问来源于stack exchange,提问作者30ThreeDegrees
相关产品推荐
相关产品推荐

