如何用BeautifulSoup去除提取内容中的无关标签?加拿大央行爬虫求助
Fixing Your Bank of Canada Web Scraper
Hey there! Let's tweak your code to get the exact date-and-announcement format you're after. Right now, your script is only pulling the link text from each media-body div, but we need to extract both the publication date and the summary content.
First, let's fix a small syntax issue in your imports (you had two statements on one line) and add a few key steps to grab the missing data:
from bs4 import BeautifulSoup import urllib.request # Fetch and parse the page url = 'https://www.bankofcanada.ca/content_type/publications/mpr/?post_type%5B0%5D=post&post_type%5B1%5D=page' page = urllib.request.urlopen(url).read() soup = BeautifulSoup(page, 'html.parser') # Specify parser to avoid warnings # Extract announcements in your desired format announcements = [] for media_body in soup.find_all("div", class_="media-body"): # Grab the date (adjust the tag/class if your inspection shows something different) date = media_body.find("small").get_text(strip=True) # Grab the announcement summary summary = media_body.a.get_text(strip=True) # Combine into your target format announcements.append(f"{date}: {summary}") # Print the results for item in announcements: print(item)
Key Changes Explained:
- Fixed Imports: Split the two import statements onto separate lines to avoid syntax errors.
- Specified Parser: Added
'html.parser'to the BeautifulSoup call to eliminate default parser warnings. - Extracted Dates: Used
find("small")to pull the publication date (this is based on common media list structures—if your browser's dev tools show the date in a different tag/class, replace"small"with that, e.g.,media_body.find("span", class_="publication-date")). - Cleaned Text: Added
strip=Trueto remove extra whitespace/newlines from the date and summary. - Formatted Output: Combined date and summary into the exact string format you requested.
If the date isn't in a <small> tag, just right-click the date on the Bank of Canada page, select "Inspect", and look at the HTML element containing the date—adjust the find() call to match that element's tag or class.
内容的提问来源于stack exchange,提问作者Abhishek Kulkarni
相关产品推荐
相关产品推荐

