如何在Beautiful Soup 4中将<td>和<tr>转换为可用行列表格?
解决Cars.com规格表格抓取问题
首先,你可以遍历每个抓取到的specs-table表格,然后逐行解析<tr>和<td>元素,将数据整理成结构化格式(比如列表或字典)。以下是修改后的代码:
from bs4 import BeautifulSoup import requests headers = {'User Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36"} url = 'https://www.cars.com/research/ford-fusion-2020/specs/407870/' response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html5lib') # 获取所有规格表格 tables = soup.find_all(class_="specs-table") # 遍历每个表格解析数据 for table in tables: # 获取表格内的所有行 rows = table.find_all('tr') for row in rows: # 获取每行的单元格 cells = row.find_all('td') # 提取单元格文本并清理空格 cell_texts = [cell.get_text(strip=True) for cell in cells] # 打印结构化后的行数据 print(cell_texts)
关键说明:
- 使用
class_参数替代attrs={'class': ...},是BeautifulSoup更简洁的写法 - 嵌套遍历
table→tr→td,逐层提取数据 get_text(strip=True)可以自动清理文本中的多余空格和换行符
如果需要将数据整理成更易读的表格形式(比如CSV),可以进一步修改代码:
import csv # 打开CSV文件准备写入 with open('ford_fusion_specs.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) for table in tables: rows = table.find_all('tr') for row in rows: cells = row.find_all('td') cell_texts = [cell.get_text(strip=True) for cell in cells] writer.writerow(cell_texts)
这样就能把抓取到的规格数据保存为本地CSV文件,直接用表格软件打开查看。
内容的提问来源于stack exchange,提问作者Odells_other_acl_406
相关产品推荐
相关产品推荐

