You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3 Beautiful Soup 4循环提取多dt/dd标签的表格化数据?

Hey there! Let's work through this USGS report scraping problem together. I've messed around with similar government pub sites before, so I can walk you through how to loop through those <dt> and <dd> tags and get your data into a clean table format.

USGS Open-File Report Scraping with BeautifulSoup 4

1. First, Get Your Tools Ready

Make sure you have the necessary libraries installed. Run this in your terminal:

pip install beautifulsoup4 requests pandas

2. The Core Idea

Looking at the USGS page structure, each report's metadata is paired up in <dt> (field names like Report Number, Title) and <dd> (the actual values). We'll grab all these pairs, map them into dictionaries, then convert everything into a pandas DataFrame for table output.

3. Full Code Example

import requests
from bs4 import BeautifulSoup
import pandas as pd
import time

def scrape_usgs_year(year):
    # Build the URL for the target year
    url = f"https://pubs.er.usgs.gov/browse/Report/USGS%20Numbered%20Series/Open-File%20Report/{year}/"
    # Add a user-agent to avoid getting blocked
    headers = {
        "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    # Fetch the page
    try:
        response = requests.get(url, headers=headers)
        response.raise_for_status()  # Throw an error if request fails
        soup = BeautifulSoup(response.text, 'html.parser')
    except Exception as e:
        print(f"Couldn't fetch page for {year}: {str(e)}")
        return pd.DataFrame()
    
    # Grab all report containers (adjust the class based on actual page structure)
    report_containers = soup.find_all('dl', class_='pub-list-item')
    all_reports = []
    
    for container in report_containers:
        report_data = {}
        # Get all dt (labels) and dd (values) in this container
        labels = container.find_all('dt')
        values = container.find_all('dd')
        
        # Pair each label with its corresponding value
        for label, value in zip(labels, values):
            clean_label = label.get_text(strip=True)
            clean_value = value.get_text(strip=True)
            # Extract links if present (like report URLs)
            link = value.find('a')
            if link:
                # Combine text and URL for clarity
                clean_value = f"{clean_value} | {link['href']}"
            report_data[clean_label] = clean_value
        
        all_reports.append(report_data)
    
    # Convert to a DataFrame (table format)
    return pd.DataFrame(all_reports)

# Test with 2018 first
df_2018 = scrape_usgs_year(2018)
print("Sample 2018 Data:")
print(df_2018.head())
# Save to CSV
df_2018.to_csv('usgs_open_file_reports_2018.csv', index=False)

# Loop through all 117 years (adjust the year range to match your needs)
start_year = 1907  # Replace with actual start year
end_year = 2023    # Replace with actual end year

for year in range(start_year, end_year + 1):
    df = scrape_usgs_year(year)
    if not df.empty:
        df.to_csv(f'usgs_open_file_reports_{year}.csv', index=False)
        print(f"Successfully scraped data for {year}")
    else:
        print(f"No data found for {year}")
    # Add a delay to avoid overwhelming the server
    time.sleep(2)

4. Key Tips to Fix Your Previous Issues

  • Why next_sibling didn't work: Often, there are hidden whitespace or <br> tags between <dt> and <dd> elements. Using zip(labels, values) from the same container is way more reliable because it pairs elements directly based on their position in the group.
  • Check page structure: Use your browser's dev tools (F12) to confirm the actual class of the container holding each report's dt/dd pairs. I used pub-list-item as a guess—you might need to tweak this.
  • Avoid getting blocked: Always include a valid User-Agent header and add delays between requests. USGS servers will block you if you hit them too fast.
  • Handle edge cases: Some years might have no data, or unexpected HTML structure. The try-except blocks and empty DataFrame check help skip those gracefully.

5. Final Output

Each year's data will be saved as a CSV file with columns matching the <dt> labels (like Report Number, Title, Publication Date) and rows for each report. You can also combine all CSVs into one big table if needed using pandas.

内容的提问来源于stack exchange,提问作者geoastro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:30:46