You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup解析表格不识别换行,如何保留原格式存入dataframe?

解决方案

核心原因

pandas的read_html方法默认会丢弃HTML中的<br>标签,直接将标签前后的文本拼接,所以即使BeautifulSoup保留了<br/>标签,转成DataFrame时换行还是会丢失,导致内容混乱。

修复步骤

  1. 在定位到目标表格CompTab后,先遍历替换表格内所有<br>标签为文本换行符\n
  2. 再将处理后的表格对象传给read_html生成DataFrame

完整修改后代码

import requests
from bs4 import BeautifulSoup as bs
import pandas as pd

resp = requests.get(url_to_use, headers=headers)
soup = bs(resp.text, "html.parser")
tables = soup.find_all("table")
# 选择目标表格
CompTab = None
for table in tables:
    table_str = str(table)
    if 'Name' in table_str and 'Year' in table_str and 'Salary' in table_str:
        CompTab = table
        break

if CompTab:
    # 替换所有br标签为换行符
    for br in CompTab.find_all("br"):
        br.replace_with("\n")
    # 转换为DataFrame
    df = pd.read_html(str(CompTab))[0]

补充说明

如果后续需要将DataFrame导出到Excel等工具中正常展示换行,只要开启单元格的「自动换行」格式即可正常显示多行效果。

内容的提问来源于stack exchange,提问作者xxgaryxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 05:24:01