如何使用Python爬取行中嵌套行的网页表格并生成结构化DataFrame
解决方案
1. 安装依赖
需要先安装用到的第三方库:
pip install beautifulsoup4 pandas lxml
2. 核心实现代码
from bs4 import BeautifulSoup import pandas as pd # 此处替换为你实际获取的网页源码,selenium可通过driver.page_source获取,requests可通过response.text获取 html_str = """ <td> <strong> Important Information1 </strong> <br> Some information <br> Some information <br> Some information <br> Some information <br> Some information <br> Some information </td> <td> <strong> Important Information 2 </strong> <br> Some information 2 <br> Some information 2 <br> Some information 2 <br> Some information 2 <br> Some information 2 <br> Some information 2 <br> Some information 2 <br> Some information 2 <br> Some information 2 <br> Some information 2 </td> <td> <strong> Important Information 3 </strong> <br> Some information 3 <br> Some information 3 <br> Some information 3 <br> Some information 3 </td> """ soup = BeautifulSoup(html_str, 'lxml') td_list = soup.find_all('td') result = [] for td in td_list: # 提取strong标签内的重要信息 important_info = td.find('strong').get_text(strip=True) # 拆分td内所有文本,过滤掉首行(即strong内容)后遍历普通信息 all_lines = td.get_text('\n', strip=True).split('\n') for info_line in all_lines[1:]: result.append({ "Important Header": important_info, "Some Information Header": info_line.strip() }) # 转换为DataFrame df = pd.DataFrame(result) # 可直接打印验证结果 print(df)
3. 可选鲁棒性优化
如果部分td结构不符合预期,可以增加异常捕获逻辑,避免程序中断:
for td in td_list: try: important_info = td.find('strong').get_text(strip=True) all_lines = td.get_text('\n', strip=True).split('\n') for info_line in all_lines[1:]: result.append({ "Important Header": important_info, "Some Information Header": info_line.strip() }) except: # 可自行增加日志打印排查异常td continue
运行代码后得到的df就是你需要的结构,可直接用于后续导出、分析等操作。
内容的提问来源于stack exchange,提问作者pkpto39
相关产品推荐
相关产品推荐

