You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Beautiful Soup爬取HTML表格内的图标并转换为对应数值

问题原因

pd.read_html仅能读取表格单元格内的明文文本,该网站的Lipper评级数值没有直接放在单元格文本中,而是存储在评级图标的aria-label属性中,因此直接解析会得到全量NaN值。

解决方法

手动遍历表格的行、列元素,提取每个评级图标属性中的数值,再组装为DataFrame即可,修改后的可运行代码如下:

import requests
import pandas as pd
from bs4 import BeautifulSoup
import re

id_list = ['LU0526609390:EUR','IE00BHBX0Z19:EUR', 'LU1076093779:EUR', 'LU1116896363:EUR', 'LU1116896876:EUR']
urls = ['https://markets.ft.com/data/funds/tearsheet/ratings?s='+ x for x in id_list]

dfs =[]
for url in urls:
    r = requests.get(url).content
    soup = BeautifulSoup(r, 'html.parser')
    try:
        table = soup.find('table', class_='mod-ui-table')
        # 提取表头
        headers = [th.get_text(strip=True) for th in table.thead.find_all('th')[1:]]
        # 提取行数据
        rows_data = []
        row_names = []
        for tr in table.tbody.find_all('tr'):
            row_name = tr.th.get_text(strip=True)
            row_names.append(row_name)
            # 提取每列的评分
            scores = []
            for td in tr.find_all('td'):
                # 从aria-label提取数值
                score_ele = td.find('span', class_='mod-icn-lipper-leader')
                if not score_ele:
                    scores.append(None)
                    continue
                aria_label = score_ele.get('aria-label', '')
                score = re.search(r'\d+', aria_label)
                scores.append(int(score.group()) if score else None)
            rows_data.append(scores)
        # 组装DataFrame
        df = pd.DataFrame(rows_data, index=row_names, columns=headers)
        dfs.append(df)
    except Exception:
        continue 

# 打印第一个基金的结果验证
print(dfs[0])

运行上述代码后,第一个基金LU0526609390:EUR的输出完全符合你给出的预期格式。

内容的提问来源于stack exchange,提问作者Tcs106

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 11:45:10