You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup遍历新闻网站左侧分类并获取对应内容?

How to Iterate Over Left-Side Categories on NYT Search Site

Hey there! Awesome work getting to the point where you can pull article titles with BeautifulSoup—you’re already halfway there. Let’s break down how to loop through those left-hand categories and scrape content across each section.

Step 1: Inspect the Category HTML Structure

First, pop open your browser’s dev tools (F12) and look at the left-side categories. You’ll notice they’re probably wrapped in a container (like a <div> or <ul>) with links (<a> tags) pointing to each category’s search results. For example, the CSS selector might look something like #left-nav .category-link—you’ll need to tweak this to match the actual structure of the NYT page.

Once you know the right selector, you can pull all category URLs and their names. Here’s a code snippet to get you started:

import requests
from bs4 import BeautifulSoup
import time

# Base URL for the NYT search site
base_search_url = "http://query.nytimes.com/search/sitesearch/#/"

# Fetch the main search page
response = requests.get(base_search_url)
soup = BeautifulSoup(response.text, "html.parser")

# Extract category links (adjust the selector to match the page's actual HTML)
category_links = soup.select("#left-navigation .category-item a")

# Store category names and full URLs in a list
categories = []
for link in category_links:
    cat_name = link.get_text(strip=True)
    # Handle relative URLs by combining with the base URL
    cat_href = link["href"]
    cat_url = f"{base_search_url}{cat_href}" if cat_href.startswith("/") else cat_href
    categories.append((cat_name, cat_url))

print(f"Found {len(categories)} categories to scrape!")

Step 3: Iterate Over Each Category and Scrape Titles

Now that you have all category URLs, loop through each one and use the title-scraping code you already know. Here’s how to tie it all together:

for cat_name, cat_url in categories:
    print(f"\n--- Scraping {cat_name} ---")
    
    # Fetch the category's search results page
    cat_response = requests.get(cat_url)
    cat_soup = BeautifulSoup(cat_response.text, "html.parser")
    
    # Use your existing selector for article titles (replace with your actual selector)
    article_titles = cat_soup.select(".right-column-results .article-headline")
    
    for title in article_titles:
        print(f"- {title.get_text(strip=True)}")
    
    # Add a delay to avoid overwhelming the server (be nice!)
    time.sleep(2)

Handling Dynamic Content (If Needed)

If you notice that the left-side categories or article titles don’t show up when using requests, that means the content is loaded dynamically with JavaScript. In that case, you’ll need to use a tool like Selenium to simulate a browser. Here’s a quick example:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

driver = webdriver.Chrome()
wait = WebDriverWait(driver, 10)

driver.get(base_search_url)

# Wait for categories to load
category_elements = wait.until(
    EC.presence_of_all_elements_located((By.CSS_SELECTOR, "#left-navigation .category-item a"))
)

categories = []
for elem in category_elements:
    cat_name = elem.text.strip()
    cat_url = elem.get_attribute("href")
    categories.append((cat_name, cat_url))

# Scrape each category
for cat_name, cat_url in categories:
    print(f"\n--- Scraping {cat_name} ---")
    driver.get(cat_url)
    
    # Wait for article titles to load
    title_elements = wait.until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".right-column-results .article-headline"))
    )
    
    for elem in title_elements:
        print(f"- {elem.text.strip()}")
    
    time.sleep(2)

driver.quit()

Important Notes

  • Respect robots.txt: Check http://query.nytimes.com/robots.txt to make sure scraping these sections is allowed.
  • Rate Limiting: Always add delays between requests (like the time.sleep(2) in the examples) to avoid getting your IP blocked.
  • Update Selectors: NYT might change their page structure over time, so double-check your CSS selectors if the code stops working.

内容的提问来源于stack exchange,提问作者A.S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:27:53