You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用正则表达式从网站抓取指定数据?Python网页爬虫新手的目标文本提取问题求助

Hey there! Let's work through fixing your web scraping script to grab exactly those entries under the bold headings you mentioned. Here's a step-by-step solution tailored to your needs:

Solution for Scraping Cuba Restricted List Entries

1. Target the Content Container Correctly

Your current code uses find_all() to get the entry-content div, but this returns a list. Since there's only one such container on the page, we can simplify this by using find() to grab it directly:

entry_content = soup.find('div', class_='entry-content')

2. Traverse Elements to Capture Headings and Entries

The page's section headings are wrapped in <strong> tags, and the entries we want come right after each heading. We'll use state tracking to start collecting when we hit the "Ministries" heading, and stop after we've captured the "Additional Subentities of Habaguanex" section.

We also need to handle two types of content elements: standalone <p> tags and <li> items inside unordered lists.

Full Working Code

import requests
from bs4 import BeautifulSoup

URL = "https://www.state.gov/cuba-restricted-list/list-of-restricted-entities-and-subentities-associated-with-cuba-effective-january-8-2021/"
page = requests.get(URL)
soup = BeautifulSoup(page.text, "lxml")

# Grab the main content container
entry_content = soup.find('div', class_='entry-content')

# List to store our final results
restricted_entries = []
# Flags to control when we start/stop collecting
is_collecting = False
stop_section = "Additional Subentities of Habaguanex"
stop_triggered = False

for element in entry_content.children:
    # Check for section headings
    if element.name == 'strong':
        heading_text = element.get_text(strip=True)
        
        # Start collecting when we reach "Ministries"
        if heading_text == "Ministries":
            is_collecting = True
        # Mark that we need to stop after processing this section
        if heading_text == stop_section:
            stop_triggered = True
            is_collecting = True
        
        # Optional: Add section headers to your list for clarity
        if is_collecting:
            restricted_entries.append(f"=== {heading_text} ===")
    
    # Collect entry text when in active mode
    if is_collecting and element.name in ['p', 'li']:
        item_text = element.get_text(strip=True)
        # Skip empty strings and header duplicates
        if item_text and not item_text.startswith("==="):
            restricted_entries.append(item_text)
    
    # Stop collecting once we've passed the final target section
    if stop_triggered and element.name == 'strong' and element.get_text(strip=True) != stop_section:
        is_collecting = False

# Print or save your results
for item in restricted_entries:
    print(item)

Key Notes

  • State Tracking: The is_collecting and stop_triggered flags ensure we only capture content between your desired start and end sections.
  • Content Compatibility: The code handles both paragraph and list item elements, so you won't miss any entries regardless of how they're formatted on the page.
  • Clean Text: get_text(strip=True) removes extra whitespace and ensures we get clean, usable text entries.

If you don't need the section headers in your final list, just remove the line where we append the === {heading_text} === string.

内容的提问来源于stack exchange,提问作者Raja Manikam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 02:17:29