如何用BeautifulSoup遍历新闻网站左侧分类并获取对应内容?
Hey there! Awesome work getting to the point where you can pull article titles with BeautifulSoup—you’re already halfway there. Let’s break down how to loop through those left-hand categories and scrape content across each section.
Step 1: Inspect the Category HTML Structure
First, pop open your browser’s dev tools (F12) and look at the left-side categories. You’ll notice they’re probably wrapped in a container (like a <div> or <ul>) with links (<a> tags) pointing to each category’s search results. For example, the CSS selector might look something like #left-nav .category-link—you’ll need to tweak this to match the actual structure of the NYT page.
Step 2: Extract Category Links with BeautifulSoup
Once you know the right selector, you can pull all category URLs and their names. Here’s a code snippet to get you started:
import requests from bs4 import BeautifulSoup import time # Base URL for the NYT search site base_search_url = "http://query.nytimes.com/search/sitesearch/#/" # Fetch the main search page response = requests.get(base_search_url) soup = BeautifulSoup(response.text, "html.parser") # Extract category links (adjust the selector to match the page's actual HTML) category_links = soup.select("#left-navigation .category-item a") # Store category names and full URLs in a list categories = [] for link in category_links: cat_name = link.get_text(strip=True) # Handle relative URLs by combining with the base URL cat_href = link["href"] cat_url = f"{base_search_url}{cat_href}" if cat_href.startswith("/") else cat_href categories.append((cat_name, cat_url)) print(f"Found {len(categories)} categories to scrape!")
Step 3: Iterate Over Each Category and Scrape Titles
Now that you have all category URLs, loop through each one and use the title-scraping code you already know. Here’s how to tie it all together:
for cat_name, cat_url in categories: print(f"\n--- Scraping {cat_name} ---") # Fetch the category's search results page cat_response = requests.get(cat_url) cat_soup = BeautifulSoup(cat_response.text, "html.parser") # Use your existing selector for article titles (replace with your actual selector) article_titles = cat_soup.select(".right-column-results .article-headline") for title in article_titles: print(f"- {title.get_text(strip=True)}") # Add a delay to avoid overwhelming the server (be nice!) time.sleep(2)
Handling Dynamic Content (If Needed)
If you notice that the left-side categories or article titles don’t show up when using requests, that means the content is loaded dynamically with JavaScript. In that case, you’ll need to use a tool like Selenium to simulate a browser. Here’s a quick example:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time driver = webdriver.Chrome() wait = WebDriverWait(driver, 10) driver.get(base_search_url) # Wait for categories to load category_elements = wait.until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "#left-navigation .category-item a")) ) categories = [] for elem in category_elements: cat_name = elem.text.strip() cat_url = elem.get_attribute("href") categories.append((cat_name, cat_url)) # Scrape each category for cat_name, cat_url in categories: print(f"\n--- Scraping {cat_name} ---") driver.get(cat_url) # Wait for article titles to load title_elements = wait.until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".right-column-results .article-headline")) ) for elem in title_elements: print(f"- {elem.text.strip()}") time.sleep(2) driver.quit()
Important Notes
- Respect
robots.txt: Checkhttp://query.nytimes.com/robots.txtto make sure scraping these sections is allowed. - Rate Limiting: Always add delays between requests (like the
time.sleep(2)in the examples) to avoid getting your IP blocked. - Update Selectors: NYT might change their page structure over time, so double-check your CSS selectors if the code stops working.
内容的提问来源于stack exchange,提问作者A.S

