You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas read_html获取Azure区域产品表的层级化数据?

获取Azure区域产品可用性的层级化(多级索引)数据

你提到的需求就是要构建Pandas的多级索引(MultiIndex) DataFrame,对应页面category-row→service-row→capability-row的三层结构。直接用pd.read_html只能得到扁平表格,你可以通过BeautifulSoup手动遍历层级元素收集数据,再构建多级索引,具体实现如下:

步骤1:遍历页面层级,收集原始数据

页面的表格结构中,每个一级分类(category-row)下包含多个二级服务(service-row),每个服务下又有多个三级功能(capability-row)。我们逐个遍历这些元素,提取名称和对应区域的可用性状态:

from selenium import webdriver
from selenium.webdriver.firefox.options import Options
from bs4 import BeautifulSoup
import pandas as pd

options = Options()
options.add_argument('--headless')
driver = webdriver.Firefox(options=options)
driver.implicitly_wait(30)

url = 'https://azure.microsoft.com/en-us/explore/global-infrastructure/products-by-region/?regions=us-east-2,canada-central,canada-east&products=all'
driver.get(url)

# 提取表格HTML
table_html = driver.find_element_by_id("primary-table").get_attribute('outerHTML')
tree = BeautifulSoup(table_html, "html5lib")
table = tree.find('table', class_='primary-table')

# 获取区域表头
region_headers = [th.get_text(strip=True) for th in table.find('tr', {'class': 'region-headers-row'}).find_all('th')]

# 遍历层级收集数据
data = []
for category_row in table.find_all('tr', class_='category-row'):
    category = category_row.find('th').get_text(strip=True)
    # 遍历当前分类下的所有服务
    service_rows = category_row.find_next_siblings('tr', class_='service-row')
    for service_row in service_rows:
        service = service_row.find('th').get_text(strip=True)
        # 遍历当前服务下的所有功能
        cap_rows = service_row.find_next_siblings('tr', class_='capability-row')
        for cap_row in cap_rows:
            # 遇到更高层级的行则终止当前循环
            if 'category-row' in cap_row.get('class', []) or 'service-row' in cap_row.get('class', []):
                break
            capability = cap_row.find('th').get_text(strip=True)
            # 提取对应区域的可用性状态
            statuses = [td.get_text(strip=True) for td in cap_row.find_all('td')[:len(region_headers)]]
            data.append([category, service, capability] + statuses)

driver.quit()

步骤2:构建多级索引DataFrame

用收集到的数据创建DataFrame,将前三级设置为多级索引:

# 定义列名
columns = ['分类', '服务', '功能'] + region_headers
df = pd.DataFrame(data, columns=columns)

# 设置三级多级索引
df_multi = df.set_index(['分类', '服务', '功能'])

# 查看结果示例
print(df_multi.head())

最终得到的df_multi会完美匹配页面的三层层级结构,你可以通过多级索引快速筛选某分类、某服务下的功能可用性数据。

内容的提问来源于stack exchange,提问作者amsteel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 17:46:07