You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫求助:从首页提取文章URL并跳转解析页面

Fixing the Article Page Parsing Issue for Your News Scraper

Hey there! It sounds like you’ve already nailed the first big part—grabbing the headline and URL from the homepage. Let’s get you over the hump of parsing the actual article page. Here are the most common fixes and a full working example to make it happen:

1. First, Check if Your URL is Complete

A super common gotcha for new scrapers: sometimes the URL you extract from the homepage is a relative path (like /2024/05/top-story.html) instead of a full, usable URL (like https://example-news-site.com/2024/05/top-story.html). If you try to request a relative path directly, it’ll fail because your script doesn’t know the base domain to attach it to.

Fix this by using urllib.parse.urljoin to combine the homepage’s base URL with the relative path:

from urllib.parse import urljoin

base_homepage_url = "https://example-news-site.com"
relative_article_url = "/2024/05/top-story.html"
full_article_url = urljoin(base_homepage_url, relative_article_url)

2. Add Request Headers to Avoid Being Blocked

Most news sites block requests that don’t look like they’re coming from a real browser. You can fix this by adding a User-Agent header to your requests to mimic a standard browser visit.

3. Full Working Script Example

Here’s a complete, commented script that ties it all together using requests and BeautifulSoup (the go-to tools for web scraping in Python):

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

# Step 1: Scrape the homepage to get the first article's URL
homepage_url = "https://example-news-site.com"
# Mimic a browser to avoid being blocked
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"
}

# Fetch and parse the homepage
try:
    homepage_response = requests.get(homepage_url, headers=headers)
    homepage_response.raise_for_status()  # Throw an error if the request fails
except requests.exceptions.RequestException as e:
    print(f"Failed to load homepage: {e}")
    exit()

homepage_soup = BeautifulSoup(homepage_response.text, "html.parser")

# Extract the first article's link (adjust the selector to match your target site!)
# Use your browser's dev tools (Right-click → Inspect) to find the correct selector
first_article_link = homepage_soup.select_one("h2 a")
if not first_article_link:
    print("Couldn't find the first article link on the homepage.")
    exit()

relative_url = first_article_link["href"]
full_article_url = urljoin(homepage_url, relative_url)
print(f"Found first article URL: {full_article_url}")

# Step 2: Scrape and parse the article page
try:
    article_response = requests.get(full_article_url, headers=headers)
    article_response.raise_for_status()
except requests.exceptions.RequestException as e:
    print(f"Failed to load article page: {e}")
    exit()

article_soup = BeautifulSoup(article_response.text, "html.parser")

# Extract article content (adjust selectors to match your target site!)
article_title = article_soup.select_one("h1").get_text(strip=True)
# Grab all paragraph tags inside the article body
article_body = "\n".join([p.get_text(strip=True) for p in article_soup.select("article p")])

print("\n--- Article Content ---")
print(f"Title: {article_title}")
print(f"\nBody:\n{article_body}")

4. Quick Troubleshooting Tips

  • 403 Forbidden Error?: Double-check your User-Agent header—some sites require extra headers like Accept-Language or Referer to let you in.
  • Empty Article Content?: Use your browser’s dev tools to find the exact CSS selectors for the article title and body. The selectors in the example are generic—you’ll need to tweak them to match your target site’s HTML structure.
  • URL Still Broken?: Print out full_article_url and paste it into your browser. If it works there but not in your script, you might be missing a header or dealing with a minor formatting issue.

内容的提问来源于stack exchange,提问作者E.C.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:30:33