使用BeautifulSoup解析HTML出现多余输出的原因及解决方法
问题描述
在Jupyter Notebook中使用BeautifulSoup解析包含重复HTML结构的文档时,出现多余的重复输出:先逐单元格分行打印内容,再按整行格式打印,但仅需要整行打印的结果。
HTML示例
<table class="tableBorder" width="100%" cellspacing="0" cellpadding="0" border="0"> <tbody> <tr> <td colspan="2" width="100%" valign="top" bgcolor="#f0f0f0"> <h3 class="formtitle"> Title <a href="somelink">Title</a> <span class="subText"> Date: 21/Dec/22 </span> </h3> </td> </tr> <tr> <td width="20%"><b>Status</b></td> <td width="80%">shipping</td> </tr> </tbody> </table> <table class="grid" width="100%" cellspacing="0" cellpadding="0" border="0"> <tbody> <tr> <td width="20%" valign="top" bgcolor="#f0f0f0"> <b>some data</b></td> <td width="30%" valign="top" bgcolor="#ffffff"> some data </td> <td bgcolor="#f0f0f0"> <b>some data:</b>some data</td> <td valign="top" nowrap="" bgcolor="#ffffff">vsome data </td> </tr> <tr> <td width="20%" valign="top" bgcolor="#f0f0f0"> <b>some data:</b> </td> </tr> </tbody> </table> <table class="grid" width="100%" cellspacing="0" cellpadding="0" border="0"> <tbody> <tr> <td width="20%" valign="top" bgcolor="#f0f0f0"> <b>Sections</b> </td> <td class="noPadding" valign="top" bgcolor="#ffffff"> <table class="blank" width="100%" cellspacing="0" cellpadding="0" border="0"> <tbody> <tr> <td colspan="4" bgcolor="#f0f0f0"> <b>Section 1</b> </td> </tr> <tr> <td> Test 1 </td> <td> <a href="somelink"> Test 1 Code </a> </td> <td> Test 1 Description </td> <td> Test 1 Extended Description </td> </tr> <tr> <td colspan="4" bgcolor="#f0f0f0"> <b>Section 2</b> </td> </tr> <tr> <td> Test 2 </td> <td> <a href="somelink"> Test 2 Code </a> </td> <td> Test 2 Description </td> <td> Test 2 Extended Description </td> </tr> <tr> <td> Test 3 </td> <td> <a href="somelink"> Test 3 Code </a> </td> <td> Test 3 Description </td> <td> Test 3 Extended Description </td> </tr> </tbody> </table> </td> </tr> </tbody> </table>
解析代码
mainHtml = soup.find_all('table', class_='tableBorder') for main in mainHtml: print () print ("URL : ", main.tbody.tr.td.h3.a["href"]) print ("Title : ", main.tbody.tr.td.h3.a.text) print ("Status : ", main.tbody.select('tr')[1].select('td')[1].text) linked = main.find_next_sibling('table', class_='grid') if linked: linked = linked.find_next_sibling('table', class_='grid') if linked: rows = linked.find_all('tr') # Iterate through the rows and extract the information for row in rows: cells = row.find_all('td') if len(cells) >= 4: # Extract the information from the cells a= cells[0].text.strip() b = cells[1].text.strip() c = cells[2].text.strip() d = cells[3].text.strip() print(a, b, c, d)
当前输出
Test 1 Test 1 Code Test 1 Description Test 1 Extended Description Test 2 Test 2 Code Test 2 Description Test 2 Extended Description Test 3 Test 3 Code Test 3 Description Test 3 Extended Description Test 1 Test 1 Code Test 1 Description Test 1 Extended Description Test 2 Test 2 Code Test 2 Description Test 2 Extended Description Test 3 Test 3 Code Test 3 Description Test 3 Extended Description
期望输出
Test 1 Test 1 Code Test 1 Description Test 1 Extended Description Test 2 Test 2 Code Test 2 Description Test 2 Extended Description Test 3 Test 3 Code Test 3 Description Test 3 Extended Description
原因分析
- 代码逻辑冗余:原代码用
linked.find_all('tr')递归查找所有层级的<tr>,包括外层表格和嵌套表格的行,若调试时曾单独执行过单元格文本输出语句(如print(a)),会触发逐行输出。 - Jupyter特性影响:Jupyter会自动输出表达式结果,若代码中存在单独的变量名(如直接写
a而非a=...),会触发单元格内容的逐行打印。
解决办法
1. 精准定位目标表格,避免递归获取多余行
修改代码中获取行的逻辑,直接定位到嵌套的table.blank,再从中提取目标行:
mainHtml = soup.find_all('table', class_='tableBorder') for main in mainHtml: print () print ("URL : ", main.tbody.tr.td.h3.a["href"]) print ("Title : ", main.tbody.tr.td.h3.a.text) print ("Status : ", main.tbody.select('tr')[1].select('td')[1].text) linked = main.find_next_sibling('table', class_='grid') if linked: linked = linked.find_next_sibling('table', class_='grid') if linked: # 精准找到嵌套的blank表格 blank_table = linked.find('table', class_='blank') if blank_table: rows = blank_table.find_all('tr') # 遍历行并提取信息 for row in rows: cells = row.find_all('td') if len(cells) >= 4: a = cells[0].text.strip() b = cells[1].text.strip() c = cells[2].text.strip() d = cells[3].text.strip() print(a, b, c, d)
2. 清理代码中的多余输出语句
检查当前代码块及Jupyter历史代码,删除所有单独打印单元格内容的语句(如print(a)、print(b)等),确保只有print(a, b, c, d)这一行输出整行内容。
3. 重启Jupyter内核(可选)
若仍存在残留输出,可重启Jupyter内核后重新运行代码,清除历史输出缓存。
内容的提问来源于stack exchange,提问作者Vanzy M
相关产品推荐
相关产品推荐

