You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页批量抓取含Read More内容并转Excel/Word文档技术问询

Alright, let's put together a full, hands-on solution for your web scraping + data conversion workflow. Since you already have the basics of URL navigation and button clicks, we'll build from there with practical code snippets and step-by-step logic.

1. Automate "Read More" Scraping with Selenium

Since you need to simulate clicks and navigate back to the list page, Selenium is the perfect tool here—it mimics real user interactions, which works great for dynamic content that requires button clicks to load full details.

Step 1: Install Dependencies

First, install the required packages:

pip install selenium webdriver-manager

Step 2: Core Scraping Code

Here's a complete script that handles the loop of clicking "Read More", extracting data, navigating back, and repeating:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from webdriver_manager.chrome import ChromeDriverManager
import time

# Initialize browser (Chrome in this example)
driver = webdriver.Chrome(ChromeDriverManager().install())
driver.get("YOUR_TARGET_LIST_PAGE_URL")

# Store scraped data in a list of dictionaries
scraped_data = []

try:
    # Get the total number of list items (adjust selector to match your page)
    list_items = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, "YOUR_LIST_ITEM_SELECTOR"))
    )
    total_items = len(list_items)

    for i in range(total_items):
        # Re-fetch list items each time (DOM might refresh after navigating back)
        list_items = WebDriverWait(driver, 10).until(
            EC.presence_of_all_elements_located((By.CSS_SELECTOR, "YOUR_LIST_ITEM_SELECTOR"))
        )
        # Locate the "Read More" button for the current item (adjust selector)
        read_more_btn = list_items[i].find_element(By.XPATH, './/button[text()="Read More"]')
        read_more_btn.click()

        # Wait for detail page to load and extract content (adjust selectors)
        time.sleep(2)  # Or use WebDriverWait for more reliability
        title = driver.find_element(By.CSS_SELECTOR, "YOUR_TITLE_SELECTOR").text.strip()
        full_content = driver.find_element(By.CSS_SELECTOR, "YOUR_CONTENT_SELECTOR").text.strip()

        # Add to scraped data
        scraped_data.append({
            "Title": title,
            "Full Content": full_content
        })

        # Navigate back to list page
        driver.back()
        # Wait for list page to reload
        time.sleep(2)

finally:
    # Close browser when done
    driver.quit()

# Print preview of scraped data
print(f"Scraped {len(scraped_data)} items successfully!")

Key Notes for This Step:

  • Replace all YOUR_*_SELECTOR placeholders with actual CSS/XPath selectors from your target site (use browser dev tools to inspect elements).
  • Use WebDriverWait instead of hardcoded time.sleep() whenever possible—it's more efficient and avoids race conditions.
  • Wrap the loop in a try/finally block to ensure the browser closes even if an error occurs.
2. Export Scraped Data to Excel

We'll use Pandas for this—it's simple and handles Excel exports seamlessly.

Step 1: Install Pandas

pip install pandas openpyxl

Step 2: Convert Data to Excel

Add this code after the scraping script:

import pandas as pd

# Convert scraped data to a DataFrame
df = pd.DataFrame(scraped_data)

# Export to Excel (index=False removes the default row numbers)
df.to_excel("scraped_content.xlsx", index=False, engine="openpyxl")
print("Data exported to scraped_content.xlsx successfully!")
3. Convert Excel to Word Document

For this, we'll use python-docx to generate a formatted Word document from the Excel data.

Step 1: Install Dependencies

pip install python-docx pandas

Step 2: Generate Word Document

Here's the script to convert the Excel file to a Word doc:

from docx import Document
from docx.shared import Pt
from docx.enum.text import WD_ALIGN_PARAGRAPH
import pandas as pd

# Load Excel data
df = pd.read_excel("scraped_content.xlsx")

# Initialize Word document
doc = Document()

# Add title to the document
title = doc.add_heading("Scraped Content", level=1)
title.alignment = WD_ALIGN_PARAGRAPH.CENTER

# Loop through each row in the DataFrame
for idx, row in df.iterrows():
    # Add section title (bold)
    doc.add_heading(row["Title"], level=2)
    # Add full content as a paragraph
    content_paragraph = doc.add_paragraph(row["Full Content"])
    # Adjust font size for readability
    content_paragraph.style.font.size = Pt(12)
    # Add a blank line between items
    doc.add_paragraph()

# Save the Word document
doc.save("scraped_content.docx")
print("Word document generated successfully!")
4. Critical Tips to Avoid Common Pitfalls
  • Handle Anti-Scraping Measures: Some sites might block automated clicks. Add a random delay between actions (use time.sleep(random.uniform(1,3))), and set a custom user-agent for your browser.
  • Clean Extracted Data: Use string methods like .strip(), .replace("\n", " ") to remove extra whitespace or line breaks from your content.
  • Error Handling: Add try/except blocks inside the scraping loop to catch issues like missing elements or failed clicks—this way, one bad item won't crash the entire script.
  • Test with a Small Batch: Before running the full scrape, test with 2-3 items to make sure your selectors work and the data is extracted correctly.

内容的提问来源于stack exchange,提问作者JonerZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:33:45