You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬虫技术请求:爬取特定链接及替换摘要为文章正文

Solution: Replace Summary with Article Body Content

Hey there! Let's adjust your code to pull the full article body instead of the repeated homepage summary. The key is to add a step where we fetch each article's detail page, parse its content, and replace the old Summary field with that content.

Modified Code

Here's the updated script with explanations in comments:

import pandas as pd
import requests
from bs4 import BeautifulSoup
import time  # Add this to avoid hitting the site too quickly

source = requests.get('https://www.vanglaini.org/').text
soup = BeautifulSoup(source, 'lxml')

list_with_headlines = []
list_with_contents = []  # Renamed from list_with_summaries for clarity
list_with_links = []

for article in soup.find_all('article'):
    if article.a is None:
        continue
    
    headline = article.a.text.strip()
    link = "https://www.vanglaini.org" + article.a['href']
    
    # Fetch the article detail page
    try:
        # Add a small delay to be respectful to the site's server
        time.sleep(1)
        article_source = requests.get(link).text
        article_soup = BeautifulSoup(article_source, 'lxml')
        
        # Extract the body content - adjust the selector based on the site's structure
        # I've used a common selector for article content; you might need to tweak this
        body_content = article_soup.find('div', class_='entry-content')  # Replace with actual selector if needed
        if body_content:
            # Clean up the text: remove extra newlines and spaces
            cleaned_content = ' '.join([p.text.strip() for p in body_content.find_all('p')])
        else:
            # Fallback if content can't be found
            cleaned_content = "Content unavailable"
            
    except requests.exceptions.RequestException as e:
        # Handle any errors when fetching the article
        print(f"Error fetching {link}: {e}")
        cleaned_content = "Failed to fetch content"
    
    # Append data to lists
    list_with_headlines.append(headline)
    list_with_contents.append(cleaned_content)
    list_with_links.append(link)

# Create DataFrame with updated column name
news_csv = pd.DataFrame({
    'Headline': list_with_headlines,
    'Article Body': list_with_contents,  # Renamed column to match new content
    'Link': list_with_links,
})

print(news_csv)
news_csv.to_csv('vanglaini_articles.csv', index=False)  # Added index=False to clean up the CSV

Key Changes Explained

  • Added delay: Using time.sleep(1) between requests prevents overwhelming the site's server, which helps avoid getting blocked.
  • Detail page fetching: For each article link, we send a new request to get the full page content.
  • Content extraction: The script looks for a div with class entry-content (a common class for article bodies). If your target site uses a different structure, right-click the article content in your browser, select "Inspect", and find the correct HTML tag/class to use in article_soup.find().
  • Error handling: A try-except block catches issues like broken links or network errors, so the script keeps running instead of crashing.
  • Cleaner CSV: Added index=False to exclude the default pandas index from the output CSV.

Quick Tip

If you find that the content selector isn't working, inspect the article page's HTML (using your browser's dev tools) to identify the exact container holding the article text. For example, if the content is inside a <section> with id main-content, replace the line with:

body_content = article_soup.find('section', id='main-content')

内容的提问来源于stack exchange,提问作者user12205480

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:37:18