You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取指定<ul>下<li>嵌套的<span>文本?

Fixing Your BeautifulSoup Extraction Issue

Let’s break down why your code isn’t working and get you the content you need:

Key Problems in Your Original Code

  1. Incorrect URL Parameter: Your URL uses &amp; instead of & for the nodeId parameter. While Amazon might redirect you, this can lead to unexpected behavior when fetching the page.
  2. Parser Limitations: The built-in html.parser struggles with malformed HTML (like the unclosed <br> in your target <li>). Switching to a more robust parser like lxml will handle this better.
  3. Missing Check for Target Element: Your first code’s try-except only catches errors, not the case where the target <ul> isn’t found. If find_all returns an empty list, the loop just doesn’t run—no error is raised, so you get no output.
  4. Logical Error in Second Code: You initialize uls as an empty list and loop over it (which does nothing) before trying to append elements. This structure won’t collect any results.

Corrected Code

First, install the lxml parser if you haven’t already:

pip install lxml

Then use this code:

from urllib.request import urlopen
from bs4 import BeautifulSoup
import sys

# Fixed URL: replaced &amp; with &
page_url = 'https://www.amazon.com/gp/help/customer/display.html/ref=hp_left_v4_sib?ie=UTF8&nodeId=G54HPVAW86CHYHKS'

try:
    page = urlopen(page_url)
except Exception as e:
    sys.exit(f"Error fetching page: {e}")

# Use lxml parser for better handling of messy HTML
soup = BeautifulSoup(page, 'lxml')

# Target UL ID (exact match as you found in the page source)
target_ul_id = 'GUID-8B03C49D-3A98-45F1-9128-392E55823F61__UL_E0490B159DE04E22AD519CE2E7D7A35B'
target_ul = soup.find('ul', id=target_ul_id)

if not target_ul:
    print("Couldn't find the 'Here’s what’s new' section.")
else:
    print("Extracted content:")
    for li in target_ul.find_all('li'):
        # Get the span with the actual content
        content_span = li.find('span', class_='a-list-item')
        if content_span:
            # Clean up text: remove extra whitespace and fix smart quotes
            clean_text = content_span.get_text(strip=True).replace('�', "'")
            # Remove the "Read Now:" label from the first item
            if clean_text.startswith('Read Now:'):
                clean_text = clean_text[len('Read Now:'):].strip()
            print(f"- {clean_text}")

What This Code Does

  1. Fixed URL: Uses & instead of &amp; to ensure we fetch the correct page.
  2. Robust Parser: lxml handles malformed HTML (like the unclosed <br>) much better than html.parser.
  3. Existence Check: Uses find (since IDs are unique) and checks if the target <ul> exists before proceeding.
  4. Clean Content Extraction: Pulls text from the <span class="a-list-item"> elements, cleans up whitespace and garbled smart quotes, and removes the unnecessary "Read Now:" label from the first item.

Expected Output

Extracted content:
- In the coming weeks, you will be able to read items that you own with a single click from the 'Before You Go' dialog.
- Performance improvements, bug fixes, and other general enhancements.

内容的提问来源于stack exchange,提问作者cashread

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:25:51