如何使用正则表达式从网站抓取指定数据?Python网页爬虫新手的目标文本提取问题求助
Hey there! Let's work through fixing your web scraping script to grab exactly those entries under the bold headings you mentioned. Here's a step-by-step solution tailored to your needs:
1. Target the Content Container Correctly
Your current code uses find_all() to get the entry-content div, but this returns a list. Since there's only one such container on the page, we can simplify this by using find() to grab it directly:
entry_content = soup.find('div', class_='entry-content')
2. Traverse Elements to Capture Headings and Entries
The page's section headings are wrapped in <strong> tags, and the entries we want come right after each heading. We'll use state tracking to start collecting when we hit the "Ministries" heading, and stop after we've captured the "Additional Subentities of Habaguanex" section.
We also need to handle two types of content elements: standalone <p> tags and <li> items inside unordered lists.
Full Working Code
import requests from bs4 import BeautifulSoup URL = "https://www.state.gov/cuba-restricted-list/list-of-restricted-entities-and-subentities-associated-with-cuba-effective-january-8-2021/" page = requests.get(URL) soup = BeautifulSoup(page.text, "lxml") # Grab the main content container entry_content = soup.find('div', class_='entry-content') # List to store our final results restricted_entries = [] # Flags to control when we start/stop collecting is_collecting = False stop_section = "Additional Subentities of Habaguanex" stop_triggered = False for element in entry_content.children: # Check for section headings if element.name == 'strong': heading_text = element.get_text(strip=True) # Start collecting when we reach "Ministries" if heading_text == "Ministries": is_collecting = True # Mark that we need to stop after processing this section if heading_text == stop_section: stop_triggered = True is_collecting = True # Optional: Add section headers to your list for clarity if is_collecting: restricted_entries.append(f"=== {heading_text} ===") # Collect entry text when in active mode if is_collecting and element.name in ['p', 'li']: item_text = element.get_text(strip=True) # Skip empty strings and header duplicates if item_text and not item_text.startswith("==="): restricted_entries.append(item_text) # Stop collecting once we've passed the final target section if stop_triggered and element.name == 'strong' and element.get_text(strip=True) != stop_section: is_collecting = False # Print or save your results for item in restricted_entries: print(item)
Key Notes
- State Tracking: The
is_collectingandstop_triggeredflags ensure we only capture content between your desired start and end sections. - Content Compatibility: The code handles both paragraph and list item elements, so you won't miss any entries regardless of how they're formatted on the page.
- Clean Text:
get_text(strip=True)removes extra whitespace and ensures we get clean, usable text entries.
If you don't need the section headers in your final list, just remove the line where we append the === {heading_text} === string.
内容的提问来源于stack exchange,提问作者Raja Manikam

