如何用BeautifulSoup提取指定<ul>下<li>嵌套的<span>文本?
Fixing Your BeautifulSoup Extraction Issue
Let’s break down why your code isn’t working and get you the content you need:
Key Problems in Your Original Code
- Incorrect URL Parameter: Your URL uses
&instead of&for thenodeIdparameter. While Amazon might redirect you, this can lead to unexpected behavior when fetching the page. - Parser Limitations: The built-in
html.parserstruggles with malformed HTML (like the unclosed<br>in your target<li>). Switching to a more robust parser likelxmlwill handle this better. - Missing Check for Target Element: Your first code’s
try-exceptonly catches errors, not the case where the target<ul>isn’t found. Iffind_allreturns an empty list, the loop just doesn’t run—no error is raised, so you get no output. - Logical Error in Second Code: You initialize
ulsas an empty list and loop over it (which does nothing) before trying to append elements. This structure won’t collect any results.
Corrected Code
First, install the lxml parser if you haven’t already:
pip install lxml
Then use this code:
from urllib.request import urlopen from bs4 import BeautifulSoup import sys # Fixed URL: replaced & with & page_url = 'https://www.amazon.com/gp/help/customer/display.html/ref=hp_left_v4_sib?ie=UTF8&nodeId=G54HPVAW86CHYHKS' try: page = urlopen(page_url) except Exception as e: sys.exit(f"Error fetching page: {e}") # Use lxml parser for better handling of messy HTML soup = BeautifulSoup(page, 'lxml') # Target UL ID (exact match as you found in the page source) target_ul_id = 'GUID-8B03C49D-3A98-45F1-9128-392E55823F61__UL_E0490B159DE04E22AD519CE2E7D7A35B' target_ul = soup.find('ul', id=target_ul_id) if not target_ul: print("Couldn't find the 'Here’s what’s new' section.") else: print("Extracted content:") for li in target_ul.find_all('li'): # Get the span with the actual content content_span = li.find('span', class_='a-list-item') if content_span: # Clean up text: remove extra whitespace and fix smart quotes clean_text = content_span.get_text(strip=True).replace('�', "'") # Remove the "Read Now:" label from the first item if clean_text.startswith('Read Now:'): clean_text = clean_text[len('Read Now:'):].strip() print(f"- {clean_text}")
What This Code Does
- Fixed URL: Uses
&instead of&to ensure we fetch the correct page. - Robust Parser:
lxmlhandles malformed HTML (like the unclosed<br>) much better thanhtml.parser. - Existence Check: Uses
find(since IDs are unique) and checks if the target<ul>exists before proceeding. - Clean Content Extraction: Pulls text from the
<span class="a-list-item">elements, cleans up whitespace and garbled smart quotes, and removes the unnecessary "Read Now:" label from the first item.
Expected Output
Extracted content: - In the coming weeks, you will be able to read items that you own with a single click from the 'Before You Go' dialog. - Performance improvements, bug fixes, and other general enhancements.
内容的提问来源于stack exchange,提问作者cashread
相关产品推荐
相关产品推荐

