You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析HTML出现多余输出的原因及解决方法

问题描述

在Jupyter Notebook中使用BeautifulSoup解析包含重复HTML结构的文档时,出现多余的重复输出:先逐单元格分行打印内容,再按整行格式打印,但仅需要整行打印的结果。

HTML示例

<table class="tableBorder" width="100%" cellspacing="0" cellpadding="0" border="0">
   <tbody>
      <tr>
         <td colspan="2" width="100%" valign="top" bgcolor="#f0f0f0">
            <h3 class="formtitle"> Title <a href="somelink">Title</a>
               <span class="subText"> Date: 21/Dec/22 </span>
            </h3>
         </td>
      </tr>
      <tr>
         <td width="20%"><b>Status</b></td>
         <td width="80%">shipping</td>
      </tr>
   </tbody>
</table>
      
<table class="grid" width="100%" cellspacing="0" cellpadding="0" border="0">
   <tbody>
      <tr>
         <td width="20%" valign="top" bgcolor="#f0f0f0"> <b>some data</b></td>
         <td width="30%" valign="top" bgcolor="#ffffff"> some data </td>
         <td bgcolor="#f0f0f0"> <b>some data:</b>some data</td>
         <td valign="top" nowrap="" bgcolor="#ffffff">vsome data </td>
      </tr>
      <tr>
         <td width="20%" valign="top" bgcolor="#f0f0f0"> <b>some data:</b> </td>
      </tr>
   </tbody>
</table>

<table class="grid" width="100%" cellspacing="0" cellpadding="0" border="0">
   <tbody>
      <tr>
         <td width="20%" valign="top" bgcolor="#f0f0f0">
            <b>Sections</b>
         </td>
         <td class="noPadding" valign="top" bgcolor="#ffffff">
            <table class="blank" width="100%" cellspacing="0" cellpadding="0" border="0">
               <tbody>
                  <tr>
                     <td colspan="4" bgcolor="#f0f0f0"> <b>Section 1</b> </td>
                  </tr>
                  <tr>
                     <td> Test 1 </td>
                     <td> <a href="somelink"> Test 1 Code </a> </td>
                     <td> Test 1 Description </td>
                     <td> Test 1 Extended Description </td>
                  </tr>
                  <tr>
                     <td colspan="4" bgcolor="#f0f0f0"> <b>Section 2</b> </td>
                  </tr>
                  <tr>
                     <td> Test 2 </td>
                     <td> <a href="somelink"> Test 2 Code </a> </td>
                     <td> Test 2 Description </td>
                     <td> Test 2 Extended Description </td>
                  </tr>
                  <tr>
                     <td> Test 3 </td>
                     <td> <a href="somelink"> Test 3 Code </a> </td>
                     <td> Test 3 Description </td>
                     <td> Test 3 Extended Description </td>
                  </tr>
               </tbody>
            </table>
         </td>
      </tr>
   </tbody>
</table>

解析代码

mainHtml = soup.find_all('table', class_='tableBorder')

for main in mainHtml:
    
    print ()
    print ("URL : ", main.tbody.tr.td.h3.a["href"])
    print ("Title : ", main.tbody.tr.td.h3.a.text)
    print ("Status : ", main.tbody.select('tr')[1].select('td')[1].text)

    linked = main.find_next_sibling('table', class_='grid')
    if linked:
        linked = linked.find_next_sibling('table', class_='grid')
    
    if linked:
        rows = linked.find_all('tr')

#       Iterate through the rows and extract the information
        for row in rows:
        
            cells = row.find_all('td')
            
            if len(cells) >= 4:
                
#               Extract the information from the cells
                a= cells[0].text.strip()
                b = cells[1].text.strip()
                c = cells[2].text.strip()
                d = cells[3].text.strip()
                
                print(a, b, c, d)

当前输出

Test 1 
Test 1 Code 
Test 1 Description 
Test 1 Extended Description

Test 2 
Test 2 Code 
Test 2 Description 
Test 2 Extended Description

Test 3 
Test 3 Code 
Test 3 Description 
Test 3 Extended Description

Test 1 Test 1 Code Test 1 Description Test 1 Extended Description
Test 2 Test 2 Code Test 2 Description Test 2 Extended Description
Test 3 Test 3 Code Test 3 Description Test 3 Extended Description

期望输出

Test 1 Test 1 Code Test 1 Description Test 1 Extended Description
Test 2 Test 2 Code Test 2 Description Test 2 Extended Description
Test 3 Test 3 Code Test 3 Description Test 3 Extended Description
原因分析
  1. 代码逻辑冗余:原代码用linked.find_all('tr')递归查找所有层级的<tr>,包括外层表格和嵌套表格的行,若调试时曾单独执行过单元格文本输出语句(如print(a)),会触发逐行输出。
  2. Jupyter特性影响:Jupyter会自动输出表达式结果,若代码中存在单独的变量名(如直接写a而非a=...),会触发单元格内容的逐行打印。
解决办法

1. 精准定位目标表格,避免递归获取多余行

修改代码中获取行的逻辑,直接定位到嵌套的table.blank,再从中提取目标行:

mainHtml = soup.find_all('table', class_='tableBorder')

for main in mainHtml:
    
    print ()
    print ("URL : ", main.tbody.tr.td.h3.a["href"])
    print ("Title : ", main.tbody.tr.td.h3.a.text)
    print ("Status : ", main.tbody.select('tr')[1].select('td')[1].text)

    linked = main.find_next_sibling('table', class_='grid')
    if linked:
        linked = linked.find_next_sibling('table', class_='grid')
    
    if linked:
        # 精准找到嵌套的blank表格
        blank_table = linked.find('table', class_='blank')
        if blank_table:
            rows = blank_table.find_all('tr')

            # 遍历行并提取信息
            for row in rows:
                cells = row.find_all('td')
                if len(cells) >= 4:
                    a = cells[0].text.strip()
                    b = cells[1].text.strip()
                    c = cells[2].text.strip()
                    d = cells[3].text.strip()
                    print(a, b, c, d)

2. 清理代码中的多余输出语句

检查当前代码块及Jupyter历史代码,删除所有单独打印单元格内容的语句(如print(a)、print(b)等),确保只有print(a, b, c, d)这一行输出整行内容。

3. 重启Jupyter内核(可选)

若仍存在残留输出,可重启Jupyter内核后重新运行代码,清除历史输出缓存。


内容的提问来源于stack exchange,提问作者Vanzy M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 14:40:15