使用BeautifulSoup提取<li>标签时无法获取完整整行内容的问题求助
Hey there, I see the issue you're facing—when printing <li> elements extracted with BeautifulSoup, the closing </li> tag is getting split onto its own line instead of staying with the rest of the tag content. This is actually just a quirk of the default html.parser that comes with Python, and there are a few easy fixes:
Option 1: Convert the Tag Object to a String Directly
If all you need is the full <li> tag as a single string (without extra formatting), just cast the BeautifulSoup tag object to a string using str():
from urllib.request import urlopen from bs4 import BeautifulSoup html = urlopen('https://www.champlain.edu/current-students') bs = BeautifulSoup(html.read(), 'html.parser') soup = bs.find(class_='secondary-nav secondary-nav-sm has-callouts') for li in soup.find_all('li'): print(str(li))
This will output each <li> element as a single, complete line with the opening and closing tags together.
Option 2: Switch to a More Robust Parser (lxml or html5lib)
The built-in html.parser has pretty basic formatting logic. Using a third-party parser like lxml will give you more compact, clean output that keeps your <li> tags intact. First, install the parser:
pip install lxml
Then update your code to use it:
from urllib.request import urlopen from bs4 import BeautifulSoup html = urlopen('https://www.champlain.edu/current-students') # Replace html.parser with lxml bs = BeautifulSoup(html.read(), 'lxml') soup = bs.find(class_='secondary-nav secondary-nav-sm has-callouts') for li in soup.find_all('li'): print(li)
The lxml parser will render the <li> elements with the opening and closing tags on the same line (or in the compact format you're expecting).
Option 3: Tweak the prettify() Method (For html.parser Users)
If you want to stick with html.parser, you can use the prettify() method with custom formatting parameters to control line breaks. This is more for fine-tuning, but it works:
from urllib.request import urlopen from bs4 import BeautifulSoup html = urlopen('https://www.champlain.edu/current-students') bs = BeautifulSoup(html.read(), 'html.parser') soup = bs.find(class_='secondary-nav secondary-nav-sm has-callouts') for li in soup.find_all('li'): # Use formatter=None to avoid extra formatting print(li.prettify(formatter=None))
Note that this might still have some minor spacing, but it will keep the </li> tag attached to the rest of the element content.
At the root of this issue is how different parsers handle output formatting. The default html.parser prioritizes readability with line breaks, while third-party parsers or direct string conversion give you the compact structure you want.
内容的提问来源于stack exchange,提问作者zeusbella

