You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何从无指定表格类的网站提取CSV格式表格数据?

解决方法

首先,你的代码里缺少requests库的导入,这是第一个需要修正的点。另外,目标网站里的有效表格是带有class="table"的那个,另一个无class的表格属于无关内容,可以忽略。下面提供两种提取表格数据的方案:

方案一:用Pandas直接读取表格(最简单)

Pandas的read_html方法可以直接抓取网页中的表格,自动处理解析逻辑,非常适合新手:

import pandas as pd
import requests

url = "https://www.canada.ca/en/immigration-refugees-citizenship/corporate/mandate/policies-operational-instructions-agreements/ministerial-instructions/express-entry-rounds.html"
# 发送请求获取页面内容
response = requests.get(url)
# 读取页面中的所有表格,返回DataFrame列表
tables = pd.read_html(response.text)
# 第一个表格就是目标数据(对应你打印的['table']的那个)
target_table = tables[0]
# 查看数据前几行
print(target_table.head())
# 保存为CSV文件(可选)
target_table.to_csv('express_entry_rounds.csv', index=False)

方案二:用BeautifulSoup手动解析表格

如果你想手动处理解析逻辑,可以这样做:

from bs4 import BeautifulSoup
import pandas as pd
import requests

url = "https://www.canada.ca/en/immigration-refugees-citizenship/corporate/mandate/policies-operational-instructions-agreements/ministerial-instructions/express-entry-rounds.html"
html_text = requests.get(url).text
soup = BeautifulSoup(html_text, 'html.parser')

# 定位到目标表格(class为table的那个)
target_table = soup.find('table', class_='table')
if not target_table:
    print("未找到目标表格")
else:
    data = []
    # 提取表头
    headers = [th.get_text(strip=True) for th in target_table.find_all('th')]
    data.append(headers)
    # 提取每一行的数据
    for row in target_table.find_all('tr')[1:]:  # 跳过表头行
        row_data = [td.get_text(strip=True) for td in row.find_all('td')]
        data.append(row_data)
    # 转换为DataFrame
    df = pd.DataFrame(data[1:], columns=data[0])
    print(df.head())

关键说明

  • 你之前打印的None是页面中第二个无class的表格,不是我们需要的目标数据,直接忽略即可。
  • 确保已经安装所需依赖:pip install requests beautifulsoup4 pandas

内容的提问来源于stack exchange,提问作者Liam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 20:15:44