如何使用Python抓取HTML文件中的多结构表格并导出为CSV?
解决HTML多表格抓取转CSV的问题
问题核心
你遇到的难点是第三个表格的行同时包含<th>和<td>标签,常规按标签类型区分表头/数据的逻辑失效,而需求明确要求非加粗文本作为列标题、加粗文本作为行数据,所以需要跳出标签限制,以文本样式作为判断依据。
解决方案思路
- 锚定判断标准:完全以文本是否被
<b>/<strong>包裹作为区分标题(非加粗)和数据(加粗)的核心依据,忽略<th>/<td>标签的差异; - 分步骤解析:
- 遍历每个表格,先提取所有非加粗的有效文本作为CSV列标题;
- 再提取所有加粗文本,按列标题的数量分组,生成对应行数据;
- 过滤空文本和冗余空格,保证CSV内容整洁。
代码示例(Python + BeautifulSoup)
以下是针对混合标签表格的适配性解析代码:
from bs4 import BeautifulSoup import csv def html_to_csv(html_file, output_csv): # 读取HTML内容 with open(html_file, 'r', encoding='utf-8') as f: soup = BeautifulSoup(f.read(), 'html.parser') all_headers = [] all_data_rows = [] # 遍历每个表格 for table in soup.find_all('table'): # 提取非加粗文本作为列标题 headers = [] non_bold_texts = table.find_all(text=lambda t: t.parent and t.parent.name not in ['b', 'strong']) for text in non_bold_texts: cleaned_text = text.strip() if cleaned_text and cleaned_text not in headers: headers.append(cleaned_text) # 提取加粗文本作为行数据 bold_texts = table.find_all(['b', 'strong']) data_cells = [t.get_text(strip=True) for t in bold_texts if t.get_text(strip=True)] # 按列标题数量分组生成行 if headers: chunk_size = len(headers) # 处理数据长度不足的情况,补空值对齐 while len(data_cells) % chunk_size != 0: data_cells.append('') # 分割成多行 table_rows = [data_cells[i:i+chunk_size] for i in range(0, len(data_cells), chunk_size)] all_headers = headers # 若多表格列标题不同,需改为追加或单独处理 all_data_rows.extend(table_rows) # 写入CSV文件 with open(output_csv, 'w', newline='', encoding='utf-8') as csv_f: writer = csv.writer(csv_f) writer.writerow(all_headers) writer.writerows(all_data_rows) # 调用示例 html_to_csv('your_full_page.html', 'output.csv')
关键适配点
- 脱离标签限制:不再依赖
<th>/<td>判断表头,完全以文本加粗属性为标准,适配混合标签的表格; - 数据对齐处理:针对数据长度与列标题数量不匹配的情况,自动补空值保证CSV结构正确;
- 去重与清洗:过滤重复标题和空文本,避免CSV出现无效内容。
如果多表格的列标题不一致,可以修改逻辑为每个表格单独生成CSV,或者合并时补充缺失列的标题与空值。
内容的提问来源于stack exchange,提问作者RageAgainstheMachine
相关产品推荐
相关产品推荐

