You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup抓取HTML表格行数据失败,寻求解决方案

问题:BeautifulSoup抓取HTML表格td数据失败

我尝试用bs4抓取HTML表格,但代码无法正常运行,需要获取表格中td的行数据来写入CSV文件。目标HTML结构如下:

<table class="sc-jAaTju bVEWLO">
    <thead>
        <tr>
            <td width="10%">Rank</td>
            <td>Trending Topic</td>
            <td width="30%">Tweet Volume</td>
        </tr>
        </thead>
        <tbody>
        <tr>
            <td>1</td>
            <td><a href="http:///example.com/search?q=%23One" target="_blank" without="true" rel="noopener noreferrer">#One</a></td>
            <td>1006.4K tweets</td>
        </tr>
        <tr>
            <td>2</td>
            <td><a href="http:///example.com/search?q=%23Two" target="_blank" without="true" rel="noopener noreferrer">#Two</a></td>
            <td>1028.7K tweets</td>
        </tr>
        <tr>
            <td>3</td>
            <td><a href="http:///example.com/search?q=%23Three" target="_blank" without="true" rel="noopener noreferrer">#Three</a></td>
            <td>Less than 10K tweets</td>
        </tr>
    </tbody>
</table>

我的第一次尝试代码:

url = requests.get(f"https://www.exportdata.io/trends/italy/2020-01-01/0")
soup = BeautifulSoup(url.text, "html.parser")

table = soup.find_all("table", attrs={"class":"sc-jAaTju bVEWLO"})

我的第二次尝试代码:

tables = soup.find_all('table') 

for table in tables:
    td = tables.td.text.strip()

两段代码均无法正常工作,请问我忽略了什么?


问题分析与解决

两次尝试的问题点

  1. 第一次代码:仅完成了表格定位,没有后续提取行和单元格数据的逻辑,相当于只找到了表格,但没做数据提取。另外find_all返回的是列表,若目标表格唯一,用find更合适。
  2. 第二次代码:循环中错误调用tables.td(tables是表格列表,应使用循环变量table),且仅提取了第一个单元格,没有遍历所有行和每行内的所有td。

正确实现代码

import requests
from bs4 import BeautifulSoup
import csv

# 发送请求并验证状态
response = requests.get("https://www.exportdata.io/trends/italy/2020-01-01/0")
if response.status_code != 200:
    print(f"请求失败,状态码:{response.status_code}")
else:
    soup = BeautifulSoup(response.text, "html.parser")
    # 定位目标表格(用class_简化属性传参)
    target_table = soup.find("table", class_="sc-jAaTju bVEWLO")
    
    if not target_table:
        print("未找到目标表格,请检查class是否正确")
    else:
        # 提取表头
        headers = [header.text.strip() for header in target_table.thead.find_all("td")]
        # 提取tbody内的所有行
        rows = target_table.tbody.find_all("tr")
        
        # 写入CSV文件
        with open("trending_topics.csv", "w", newline="", encoding="utf-8") as csv_file:
            writer = csv.writer(csv_file)
            writer.writerow(headers)
            
            for row in rows:
                # 提取每行的所有td文本
                row_content = [cell.text.strip() for cell in row.find_all("td")]
                writer.writerow(row_content)
        
        print("数据已成功写入CSV文件")

额外注意事项

  • 如果页面数据是JS动态渲染,requests无法获取到表格内容,需改用Selenium、Playwright等工具模拟浏览器加载页面。
  • 若网站的class是动态生成(会随机变化),可换用其他定位方式,比如通过表头文本匹配表格:
    # 示例:通过表头包含"Rank"和"Tweet Volume"定位表格
    for table in soup.find_all("table"):
        thead_text = table.thead.get_text(strip=True) if table.thead else ""
        if "Rank" in thead_text and "Tweet Volume" in thead_text:
            target_table = table
            break
    

内容的提问来源于stack exchange,提问作者Kome Gognome

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 09:24:25