You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python爬取行中嵌套行的网页表格并生成结构化DataFrame

解决方案

1. 安装依赖

需要先安装用到的第三方库:

pip install beautifulsoup4 pandas lxml

2. 核心实现代码

from bs4 import BeautifulSoup
import pandas as pd

# 此处替换为你实际获取的网页源码,selenium可通过driver.page_source获取,requests可通过response.text获取
html_str = """
<td>
    <strong>    Important Information1  </strong>
    <br>    Some information
    <br>    Some information
    <br>    Some information
    <br>    Some information
    <br>    Some information
    <br>    Some information
</td>
<td>
    <strong>    Important Information 2 </strong>
    <br>    Some information 2
    <br>    Some information 2
    <br>    Some information 2
    <br>    Some information 2
    <br>    Some information 2  
    <br>    Some information 2
    <br>    Some information 2  
    <br>    Some information 2
    <br>    Some information 2  
    <br>    Some information 2  
</td>
<td>
    <strong>    Important Information 3 </strong>
    <br>    Some information 3
    <br>    Some information 3
    <br>    Some information 3
    <br>    Some information 3  
</td>
"""

soup = BeautifulSoup(html_str, 'lxml')
td_list = soup.find_all('td')
result = []

for td in td_list:
    # 提取strong标签内的重要信息
    important_info = td.find('strong').get_text(strip=True)
    # 拆分td内所有文本,过滤掉首行(即strong内容)后遍历普通信息
    all_lines = td.get_text('\n', strip=True).split('\n')
    for info_line in all_lines[1:]:
        result.append({
            "Important Header": important_info,
            "Some Information Header": info_line.strip()
        })

# 转换为DataFrame
df = pd.DataFrame(result)
# 可直接打印验证结果
print(df)

3. 可选鲁棒性优化

如果部分td结构不符合预期,可以增加异常捕获逻辑,避免程序中断:

for td in td_list:
    try:
        important_info = td.find('strong').get_text(strip=True)
        all_lines = td.get_text('\n', strip=True).split('\n')
        for info_line in all_lines[1:]:
            result.append({
                "Important Header": important_info,
                "Some Information Header": info_line.strip()
            })
    except:
        # 可自行增加日志打印排查异常td
        continue

运行代码后得到的df就是你需要的结构,可直接用于后续导出、分析等操作。

内容的提问来源于stack exchange,提问作者pkpto39

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 14:24:04