You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Beautiful Soup 4中将<td>和<tr>转换为可用行列表格?

解决Cars.com规格表格抓取问题

首先,你可以遍历每个抓取到的specs-table表格,然后逐行解析<tr>和<td>元素,将数据整理成结构化格式(比如列表或字典)。以下是修改后的代码:

from bs4 import BeautifulSoup
import requests

headers = {'User Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36"}
url = 'https://www.cars.com/research/ford-fusion-2020/specs/407870/'

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html5lib')

# 获取所有规格表格
tables = soup.find_all(class_="specs-table")

# 遍历每个表格解析数据
for table in tables:
    # 获取表格内的所有行
    rows = table.find_all('tr')
    for row in rows:
        # 获取每行的单元格
        cells = row.find_all('td')
        # 提取单元格文本并清理空格
        cell_texts = [cell.get_text(strip=True) for cell in cells]
        # 打印结构化后的行数据
        print(cell_texts)

关键说明:

  • 使用class_参数替代attrs={'class': ...},是BeautifulSoup更简洁的写法
  • 嵌套遍历table→tr→td,逐层提取数据
  • get_text(strip=True)可以自动清理文本中的多余空格和换行符

如果需要将数据整理成更易读的表格形式(比如CSV),可以进一步修改代码:

import csv

# 打开CSV文件准备写入
with open('ford_fusion_specs.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    for table in tables:
        rows = table.find_all('tr')
        for row in rows:
            cells = row.find_all('td')
            cell_texts = [cell.get_text(strip=True) for cell in cells]
            writer.writerow(cell_texts)

这样就能把抓取到的规格数据保存为本地CSV文件,直接用表格软件打开查看。

内容的提问来源于stack exchange,提问作者Odells_other_acl_406

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 11:45:37