You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+BeautifulSoup/Selenium按季度和国家提取PDF链接

Hey there! Let's tackle this PDF scraping and organizing challenge together. It sounds like you're trying to target links grouped by quarters (like Q4 2014 – Q3 2015) and countries (Malaysia, Indonesia, etc.), then save those PDFs into nested folders (quarter > country). Here's a step-by-step approach to make this work:

Step 1: Parse HTML Structure with BeautifulSoup

First, let's assume your site's HTML follows a common grouped structure (adjust selectors to match your actual snippet):

<div class="quarter-section">
  <h2>Q4 2014 – Q3 2015</h2>
  <div class="country-container">
    <div class="country-item">
      <span class="country-name">Malaysia</span>
      <a href="/assets/reports/my_q42014_q32015.pdf" class="pdf-link">Download</a>
    </div>
    <div class="country-item">
      <span class="country-name">Indonesia</span>
      <a href="/assets/reports/id_q42014_q32015.pdf" class="pdf-link">Download</a>
    </div>
  </div>
</div>

Use this code to extract quarter-country-PDF mappings:

from bs4 import BeautifulSoup
import requests

# Fetch static page content
target_url = "your-site-url-here"
page_response = requests.get(target_url)
soup = BeautifulSoup(page_response.text, "html.parser")

# Grab all quarter sections
quarter_sections = soup.find_all("div", class_="quarter-section")

for section in quarter_sections:
    # Clean up quarter name (remove extra spaces/newlines)
    quarter_title = section.find("h2").get_text(strip=True)
    # Get all country entries inside this quarter
    country_items = section.find_all("div", class_="country-item")
    
    for item in country_items:
        country_name = item.find("span", class_="country-name").get_text(strip=True)
        pdf_relative_link = item.find("a", class_="pdf-link")["href"]
        # Convert relative link to absolute if needed
        absolute_pdf_link = f"{target_url.split('/')[0]}//{target_url.split('/')[2]}{pdf_relative_link}"
        
        print(f"Quarter: {quarter_title} | Country: {country_name} | PDF: {absolute_pdf_link}")

Step 2: Handle Dynamic Content with Selenium

If the PDF links load dynamically (via JavaScript), use Selenium to wait for elements to render before parsing:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# Initialize browser (use Chrome/Firefox driver)
driver = webdriver.Chrome()
driver.get(target_url)

# Wait for quarter sections to load (adjust selector to match your site)
WebDriverWait(driver, 10).until(
    EC.presence_of_all_elements_located((By.CLASS_NAME, "quarter-section"))
)

# Pass rendered page source to BeautifulSoup
soup = BeautifulSoup(driver.page_source, "html.parser")
driver.quit()

# Reuse the parsing logic from Step 1 here

Step 3: Organize & Download PDFs to Nested Folders

Once you have the mappings, create nested folders and save PDFs with this code:

import os
import requests

for section in quarter_sections:
    quarter_title = section.find("h2").get_text(strip=True)
    # Replace invalid folder characters (like slashes/dashes)
    safe_quarter_name = quarter_title.replace(" – ", "_").replace("/", "-")
    country_items = section.find_all("div", class_="country-item")
    
    for item in country_items:
        country_name = item.find("span", class_="country-name").get_text(strip=True)
        pdf_link = item.find("a", class_="pdf-link")["href"]
        if not pdf_link.startswith("http"):
            pdf_link = f"{target_url.split('/')[0]}//{target_url.split('/')[2]}{pdf_link}"
        
        # Create nested folder path: Quarter > Country
        folder_path = os.path.join(safe_quarter_name, country_name)
        os.makedirs(folder_path, exist_ok=True)
        
        # Download and save the PDF
        pdf_response = requests.get(pdf_link)
        pdf_filename = os.path.basename(pdf_link)
        # Or use a custom name: f"{country_name}_{safe_quarter_name}.pdf"
        with open(os.path.join(folder_path, pdf_filename), "wb") as pdf_file:
            pdf_file.write(pdf_response.content)
        
        print(f"Saved PDF: {os.path.join(folder_path, pdf_filename)}")

Step 4: Fix find_all() Issues

If your find_all() calls aren't returning results, try these fixes:

  • Verify selectors: Use soup.prettify() to print the parsed HTML and confirm class names/tags match exactly (no typos!).
  • Check dynamic loading: If elements don't appear in BeautifulSoup, use your browser's dev tools to see if content loads via JS—switch to Selenium or check for API endpoints that fetch the data.
  • Clean text/attributes: Use .get_text(strip=True) to remove extra whitespace, and ensure you're accessing attributes like ["href"] correctly (make sure the <a> tag actually has that attribute).
  • Traverse nested elements: Make sure you're finding country items inside the quarter section, not globally (avoids mixing up data across quarters).

内容的提问来源于stack exchange,提问作者Funkeh-Monkeh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:14:51