使用Python+BeautifulSoup/Selenium按季度和国家提取PDF链接
Hey there! Let's tackle this PDF scraping and organizing challenge together. It sounds like you're trying to target links grouped by quarters (like Q4 2014 – Q3 2015) and countries (Malaysia, Indonesia, etc.), then save those PDFs into nested folders (quarter > country). Here's a step-by-step approach to make this work:
Step 1: Parse HTML Structure with BeautifulSoup
First, let's assume your site's HTML follows a common grouped structure (adjust selectors to match your actual snippet):
<div class="quarter-section"> <h2>Q4 2014 – Q3 2015</h2> <div class="country-container"> <div class="country-item"> <span class="country-name">Malaysia</span> <a href="/assets/reports/my_q42014_q32015.pdf" class="pdf-link">Download</a> </div> <div class="country-item"> <span class="country-name">Indonesia</span> <a href="/assets/reports/id_q42014_q32015.pdf" class="pdf-link">Download</a> </div> </div> </div>
Use this code to extract quarter-country-PDF mappings:
from bs4 import BeautifulSoup import requests # Fetch static page content target_url = "your-site-url-here" page_response = requests.get(target_url) soup = BeautifulSoup(page_response.text, "html.parser") # Grab all quarter sections quarter_sections = soup.find_all("div", class_="quarter-section") for section in quarter_sections: # Clean up quarter name (remove extra spaces/newlines) quarter_title = section.find("h2").get_text(strip=True) # Get all country entries inside this quarter country_items = section.find_all("div", class_="country-item") for item in country_items: country_name = item.find("span", class_="country-name").get_text(strip=True) pdf_relative_link = item.find("a", class_="pdf-link")["href"] # Convert relative link to absolute if needed absolute_pdf_link = f"{target_url.split('/')[0]}//{target_url.split('/')[2]}{pdf_relative_link}" print(f"Quarter: {quarter_title} | Country: {country_name} | PDF: {absolute_pdf_link}")
Step 2: Handle Dynamic Content with Selenium
If the PDF links load dynamically (via JavaScript), use Selenium to wait for elements to render before parsing:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # Initialize browser (use Chrome/Firefox driver) driver = webdriver.Chrome() driver.get(target_url) # Wait for quarter sections to load (adjust selector to match your site) WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CLASS_NAME, "quarter-section")) ) # Pass rendered page source to BeautifulSoup soup = BeautifulSoup(driver.page_source, "html.parser") driver.quit() # Reuse the parsing logic from Step 1 here
Step 3: Organize & Download PDFs to Nested Folders
Once you have the mappings, create nested folders and save PDFs with this code:
import os import requests for section in quarter_sections: quarter_title = section.find("h2").get_text(strip=True) # Replace invalid folder characters (like slashes/dashes) safe_quarter_name = quarter_title.replace(" – ", "_").replace("/", "-") country_items = section.find_all("div", class_="country-item") for item in country_items: country_name = item.find("span", class_="country-name").get_text(strip=True) pdf_link = item.find("a", class_="pdf-link")["href"] if not pdf_link.startswith("http"): pdf_link = f"{target_url.split('/')[0]}//{target_url.split('/')[2]}{pdf_link}" # Create nested folder path: Quarter > Country folder_path = os.path.join(safe_quarter_name, country_name) os.makedirs(folder_path, exist_ok=True) # Download and save the PDF pdf_response = requests.get(pdf_link) pdf_filename = os.path.basename(pdf_link) # Or use a custom name: f"{country_name}_{safe_quarter_name}.pdf" with open(os.path.join(folder_path, pdf_filename), "wb") as pdf_file: pdf_file.write(pdf_response.content) print(f"Saved PDF: {os.path.join(folder_path, pdf_filename)}")
Step 4: Fix find_all() Issues
If your find_all() calls aren't returning results, try these fixes:
- Verify selectors: Use
soup.prettify()to print the parsed HTML and confirm class names/tags match exactly (no typos!). - Check dynamic loading: If elements don't appear in BeautifulSoup, use your browser's dev tools to see if content loads via JS—switch to Selenium or check for API endpoints that fetch the data.
- Clean text/attributes: Use
.get_text(strip=True)to remove extra whitespace, and ensure you're accessing attributes like["href"]correctly (make sure the<a>tag actually has that attribute). - Traverse nested elements: Make sure you're finding country items inside the quarter section, not globally (avoids mixing up data across quarters).
内容的提问来源于stack exchange,提问作者Funkeh-Monkeh

