如何使用Beautiful Soup爬取HTML表格内的图标并转换为对应数值
问题原因
pd.read_html仅能读取表格单元格内的明文文本,该网站的Lipper评级数值没有直接放在单元格文本中,而是存储在评级图标的aria-label属性中,因此直接解析会得到全量NaN值。
解决方法
手动遍历表格的行、列元素,提取每个评级图标属性中的数值,再组装为DataFrame即可,修改后的可运行代码如下:
import requests import pandas as pd from bs4 import BeautifulSoup import re id_list = ['LU0526609390:EUR','IE00BHBX0Z19:EUR', 'LU1076093779:EUR', 'LU1116896363:EUR', 'LU1116896876:EUR'] urls = ['https://markets.ft.com/data/funds/tearsheet/ratings?s='+ x for x in id_list] dfs =[] for url in urls: r = requests.get(url).content soup = BeautifulSoup(r, 'html.parser') try: table = soup.find('table', class_='mod-ui-table') # 提取表头 headers = [th.get_text(strip=True) for th in table.thead.find_all('th')[1:]] # 提取行数据 rows_data = [] row_names = [] for tr in table.tbody.find_all('tr'): row_name = tr.th.get_text(strip=True) row_names.append(row_name) # 提取每列的评分 scores = [] for td in tr.find_all('td'): # 从aria-label提取数值 score_ele = td.find('span', class_='mod-icn-lipper-leader') if not score_ele: scores.append(None) continue aria_label = score_ele.get('aria-label', '') score = re.search(r'\d+', aria_label) scores.append(int(score.group()) if score else None) rows_data.append(scores) # 组装DataFrame df = pd.DataFrame(rows_data, index=row_names, columns=headers) dfs.append(df) except Exception: continue # 打印第一个基金的结果验证 print(dfs[0])
运行上述代码后,第一个基金LU0526609390:EUR的输出完全符合你给出的预期格式。
内容的提问来源于stack exchange,提问作者Tcs106
相关产品推荐
相关产品推荐

