You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何BeautifulSoup会忽略<body>内的第一个<p>标签?

BeautifulSoup解析时第一个

标签被忽略的原因分析

问题现象

在使用BeautifulSoup解析HTML时,出现第一个<p>标签被忽略的异常情况:

  1. 第一段代码运行后,p1标签被完全忽略:
from bs4 import BeautifulSoup

html = '''
    <html>
        <head><meta charset="UTF-8><title>something1</title><head>
        <body>             
            <p id= "p1">something2</p>
            <p id= "p2" class='red'>somthing3</p>            
        </body>
    </html>
'''

sp = BeautifulSoup(html, 'html.parser')
print(sp.find('p'))
print(sp.find_all('p'))
print(sp.find('p',{'id':'p2', 'class':'red'}))
print(sp.find("p", id='p0'))
  1. 在p1上方添加p0标签后,运行代码发现<body>内的第一个<p>标签仍被忽略:
from bs4 import BeautifulSoup

html = '''
    <html>
        <head><meta charset="UTF-8><title>something1</title><head>
        <body>            
            <p id= "p0">something0</p>
            <p id= "p1">something2</p>
            <p id= "p2" class='red'>somthing3</p>            
        </body>
    </html>
'''

sp = BeautifulSoup(html, 'html.parser')
print(sp.find('p'))
print(sp.find_all('p'))
print(sp.find('p',{'id':'p2', 'class':'red'}))
print(sp.find("p", id='p0'))

原因分析

问题根源在于HTML代码的两处语法错误:

  • <meta charset="UTF-8>的双引号未闭合,Python内置的html.parser容错能力有限,会错误地将后续内容(包括第一个<p>标签)识别为meta标签的属性值,直到遇到下一个双引号才结束错误解析状态,导致第一个<p>被“吞掉”。
  • <head>标签未正确闭合,原代码结尾处用了<head>而非</head>,进一步干扰了解析器的结构识别。

只要修正这两处错误:将meta标签改为<meta charset="UTF-8">,将结尾的<head>替换为</head>,BeautifulSoup就能正常识别所有<p>标签。

内容的提问来源于stack exchange,提问作者Chiu Sung Yu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 14:07:24