网页爬虫技术请求:爬取特定链接及替换摘要为文章正文
Solution: Replace Summary with Article Body Content
Hey there! Let's adjust your code to pull the full article body instead of the repeated homepage summary. The key is to add a step where we fetch each article's detail page, parse its content, and replace the old Summary field with that content.
Modified Code
Here's the updated script with explanations in comments:
import pandas as pd import requests from bs4 import BeautifulSoup import time # Add this to avoid hitting the site too quickly source = requests.get('https://www.vanglaini.org/').text soup = BeautifulSoup(source, 'lxml') list_with_headlines = [] list_with_contents = [] # Renamed from list_with_summaries for clarity list_with_links = [] for article in soup.find_all('article'): if article.a is None: continue headline = article.a.text.strip() link = "https://www.vanglaini.org" + article.a['href'] # Fetch the article detail page try: # Add a small delay to be respectful to the site's server time.sleep(1) article_source = requests.get(link).text article_soup = BeautifulSoup(article_source, 'lxml') # Extract the body content - adjust the selector based on the site's structure # I've used a common selector for article content; you might need to tweak this body_content = article_soup.find('div', class_='entry-content') # Replace with actual selector if needed if body_content: # Clean up the text: remove extra newlines and spaces cleaned_content = ' '.join([p.text.strip() for p in body_content.find_all('p')]) else: # Fallback if content can't be found cleaned_content = "Content unavailable" except requests.exceptions.RequestException as e: # Handle any errors when fetching the article print(f"Error fetching {link}: {e}") cleaned_content = "Failed to fetch content" # Append data to lists list_with_headlines.append(headline) list_with_contents.append(cleaned_content) list_with_links.append(link) # Create DataFrame with updated column name news_csv = pd.DataFrame({ 'Headline': list_with_headlines, 'Article Body': list_with_contents, # Renamed column to match new content 'Link': list_with_links, }) print(news_csv) news_csv.to_csv('vanglaini_articles.csv', index=False) # Added index=False to clean up the CSV
Key Changes Explained
- Added delay: Using
time.sleep(1)between requests prevents overwhelming the site's server, which helps avoid getting blocked. - Detail page fetching: For each article link, we send a new request to get the full page content.
- Content extraction: The script looks for a
divwith classentry-content(a common class for article bodies). If your target site uses a different structure, right-click the article content in your browser, select "Inspect", and find the correct HTML tag/class to use inarticle_soup.find(). - Error handling: A
try-exceptblock catches issues like broken links or network errors, so the script keeps running instead of crashing. - Cleaner CSV: Added
index=Falseto exclude the default pandas index from the output CSV.
Quick Tip
If you find that the content selector isn't working, inspect the article page's HTML (using your browser's dev tools) to identify the exact container holding the article text. For example, if the content is inside a <section> with id main-content, replace the line with:
body_content = article_soup.find('section', id='main-content')
内容的提问来源于stack exchange,提问作者user12205480
相关产品推荐
相关产品推荐

