网页批量抓取含Read More内容并转Excel/Word文档技术问询
Alright, let's put together a full, hands-on solution for your web scraping + data conversion workflow. Since you already have the basics of URL navigation and button clicks, we'll build from there with practical code snippets and step-by-step logic.
Since you need to simulate clicks and navigate back to the list page, Selenium is the perfect tool here—it mimics real user interactions, which works great for dynamic content that requires button clicks to load full details.
Step 1: Install Dependencies
First, install the required packages:
pip install selenium webdriver-manager
Step 2: Core Scraping Code
Here's a complete script that handles the loop of clicking "Read More", extracting data, navigating back, and repeating:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from webdriver_manager.chrome import ChromeDriverManager import time # Initialize browser (Chrome in this example) driver = webdriver.Chrome(ChromeDriverManager().install()) driver.get("YOUR_TARGET_LIST_PAGE_URL") # Store scraped data in a list of dictionaries scraped_data = [] try: # Get the total number of list items (adjust selector to match your page) list_items = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "YOUR_LIST_ITEM_SELECTOR")) ) total_items = len(list_items) for i in range(total_items): # Re-fetch list items each time (DOM might refresh after navigating back) list_items = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "YOUR_LIST_ITEM_SELECTOR")) ) # Locate the "Read More" button for the current item (adjust selector) read_more_btn = list_items[i].find_element(By.XPATH, './/button[text()="Read More"]') read_more_btn.click() # Wait for detail page to load and extract content (adjust selectors) time.sleep(2) # Or use WebDriverWait for more reliability title = driver.find_element(By.CSS_SELECTOR, "YOUR_TITLE_SELECTOR").text.strip() full_content = driver.find_element(By.CSS_SELECTOR, "YOUR_CONTENT_SELECTOR").text.strip() # Add to scraped data scraped_data.append({ "Title": title, "Full Content": full_content }) # Navigate back to list page driver.back() # Wait for list page to reload time.sleep(2) finally: # Close browser when done driver.quit() # Print preview of scraped data print(f"Scraped {len(scraped_data)} items successfully!")
Key Notes for This Step:
- Replace all
YOUR_*_SELECTORplaceholders with actual CSS/XPath selectors from your target site (use browser dev tools to inspect elements). - Use
WebDriverWaitinstead of hardcodedtime.sleep()whenever possible—it's more efficient and avoids race conditions. - Wrap the loop in a
try/finallyblock to ensure the browser closes even if an error occurs.
We'll use Pandas for this—it's simple and handles Excel exports seamlessly.
Step 1: Install Pandas
pip install pandas openpyxl
Step 2: Convert Data to Excel
Add this code after the scraping script:
import pandas as pd # Convert scraped data to a DataFrame df = pd.DataFrame(scraped_data) # Export to Excel (index=False removes the default row numbers) df.to_excel("scraped_content.xlsx", index=False, engine="openpyxl") print("Data exported to scraped_content.xlsx successfully!")
For this, we'll use python-docx to generate a formatted Word document from the Excel data.
Step 1: Install Dependencies
pip install python-docx pandas
Step 2: Generate Word Document
Here's the script to convert the Excel file to a Word doc:
from docx import Document from docx.shared import Pt from docx.enum.text import WD_ALIGN_PARAGRAPH import pandas as pd # Load Excel data df = pd.read_excel("scraped_content.xlsx") # Initialize Word document doc = Document() # Add title to the document title = doc.add_heading("Scraped Content", level=1) title.alignment = WD_ALIGN_PARAGRAPH.CENTER # Loop through each row in the DataFrame for idx, row in df.iterrows(): # Add section title (bold) doc.add_heading(row["Title"], level=2) # Add full content as a paragraph content_paragraph = doc.add_paragraph(row["Full Content"]) # Adjust font size for readability content_paragraph.style.font.size = Pt(12) # Add a blank line between items doc.add_paragraph() # Save the Word document doc.save("scraped_content.docx") print("Word document generated successfully!")
- Handle Anti-Scraping Measures: Some sites might block automated clicks. Add a random delay between actions (use
time.sleep(random.uniform(1,3))), and set a custom user-agent for your browser. - Clean Extracted Data: Use string methods like
.strip(),.replace("\n", " ")to remove extra whitespace or line breaks from your content. - Error Handling: Add
try/exceptblocks inside the scraping loop to catch issues like missing elements or failed clicks—this way, one bad item won't crash the entire script. - Test with a Small Batch: Before running the full scrape, test with 2-3 items to make sure your selectors work and the data is extracted correctly.
内容的提问来源于stack exchange,提问作者JonerZ

