You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取<li>标签时无法获取完整整行内容的问题求助

Fix: BeautifulSoup Outputs on a Separate Line

Hey there, I see the issue you're facing—when printing <li> elements extracted with BeautifulSoup, the closing </li> tag is getting split onto its own line instead of staying with the rest of the tag content. This is actually just a quirk of the default html.parser that comes with Python, and there are a few easy fixes:

Option 1: Convert the Tag Object to a String Directly

If all you need is the full <li> tag as a single string (without extra formatting), just cast the BeautifulSoup tag object to a string using str():

from urllib.request import urlopen
from bs4 import BeautifulSoup
html = urlopen('https://www.champlain.edu/current-students')
bs = BeautifulSoup(html.read(), 'html.parser')
soup = bs.find(class_='secondary-nav secondary-nav-sm has-callouts')
for li in soup.find_all('li'):
    print(str(li))

This will output each <li> element as a single, complete line with the opening and closing tags together.

Option 2: Switch to a More Robust Parser (lxml or html5lib)

The built-in html.parser has pretty basic formatting logic. Using a third-party parser like lxml will give you more compact, clean output that keeps your <li> tags intact. First, install the parser:

pip install lxml

Then update your code to use it:

from urllib.request import urlopen
from bs4 import BeautifulSoup
html = urlopen('https://www.champlain.edu/current-students')
# Replace html.parser with lxml
bs = BeautifulSoup(html.read(), 'lxml')
soup = bs.find(class_='secondary-nav secondary-nav-sm has-callouts')
for li in soup.find_all('li'):
    print(li)

The lxml parser will render the <li> elements with the opening and closing tags on the same line (or in the compact format you're expecting).

Option 3: Tweak the prettify() Method (For html.parser Users)

If you want to stick with html.parser, you can use the prettify() method with custom formatting parameters to control line breaks. This is more for fine-tuning, but it works:

from urllib.request import urlopen
from bs4 import BeautifulSoup
html = urlopen('https://www.champlain.edu/current-students')
bs = BeautifulSoup(html.read(), 'html.parser')
soup = bs.find(class_='secondary-nav secondary-nav-sm has-callouts')
for li in soup.find_all('li'):
    # Use formatter=None to avoid extra formatting
    print(li.prettify(formatter=None))

Note that this might still have some minor spacing, but it will keep the </li> tag attached to the rest of the element content.

At the root of this issue is how different parsers handle output formatting. The default html.parser prioritizes readability with line breaks, while third-party parsers or direct string conversion give you the compact structure you want.

内容的提问来源于stack exchange,提问作者zeusbella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 06:37:29