如何用Python 3 Beautiful Soup 4循环提取多dt/dd标签的表格化数据?
Hey there! Let's work through this USGS report scraping problem together. I've messed around with similar government pub sites before, so I can walk you through how to loop through those <dt> and <dd> tags and get your data into a clean table format.
1. First, Get Your Tools Ready
Make sure you have the necessary libraries installed. Run this in your terminal:
pip install beautifulsoup4 requests pandas
2. The Core Idea
Looking at the USGS page structure, each report's metadata is paired up in <dt> (field names like Report Number, Title) and <dd> (the actual values). We'll grab all these pairs, map them into dictionaries, then convert everything into a pandas DataFrame for table output.
3. Full Code Example
import requests from bs4 import BeautifulSoup import pandas as pd import time def scrape_usgs_year(year): # Build the URL for the target year url = f"https://pubs.er.usgs.gov/browse/Report/USGS%20Numbered%20Series/Open-File%20Report/{year}/" # Add a user-agent to avoid getting blocked headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # Fetch the page try: response = requests.get(url, headers=headers) response.raise_for_status() # Throw an error if request fails soup = BeautifulSoup(response.text, 'html.parser') except Exception as e: print(f"Couldn't fetch page for {year}: {str(e)}") return pd.DataFrame() # Grab all report containers (adjust the class based on actual page structure) report_containers = soup.find_all('dl', class_='pub-list-item') all_reports = [] for container in report_containers: report_data = {} # Get all dt (labels) and dd (values) in this container labels = container.find_all('dt') values = container.find_all('dd') # Pair each label with its corresponding value for label, value in zip(labels, values): clean_label = label.get_text(strip=True) clean_value = value.get_text(strip=True) # Extract links if present (like report URLs) link = value.find('a') if link: # Combine text and URL for clarity clean_value = f"{clean_value} | {link['href']}" report_data[clean_label] = clean_value all_reports.append(report_data) # Convert to a DataFrame (table format) return pd.DataFrame(all_reports) # Test with 2018 first df_2018 = scrape_usgs_year(2018) print("Sample 2018 Data:") print(df_2018.head()) # Save to CSV df_2018.to_csv('usgs_open_file_reports_2018.csv', index=False) # Loop through all 117 years (adjust the year range to match your needs) start_year = 1907 # Replace with actual start year end_year = 2023 # Replace with actual end year for year in range(start_year, end_year + 1): df = scrape_usgs_year(year) if not df.empty: df.to_csv(f'usgs_open_file_reports_{year}.csv', index=False) print(f"Successfully scraped data for {year}") else: print(f"No data found for {year}") # Add a delay to avoid overwhelming the server time.sleep(2)
4. Key Tips to Fix Your Previous Issues
- Why
next_siblingdidn't work: Often, there are hidden whitespace or<br>tags between<dt>and<dd>elements. Usingzip(labels, values)from the same container is way more reliable because it pairs elements directly based on their position in the group. - Check page structure: Use your browser's dev tools (F12) to confirm the actual class of the container holding each report's dt/dd pairs. I used
pub-list-itemas a guess—you might need to tweak this. - Avoid getting blocked: Always include a valid
User-Agentheader and add delays between requests. USGS servers will block you if you hit them too fast. - Handle edge cases: Some years might have no data, or unexpected HTML structure. The try-except blocks and empty DataFrame check help skip those gracefully.
5. Final Output
Each year's data will be saved as a CSV file with columns matching the <dt> labels (like Report Number, Title, Publication Date) and rows for each report. You can also combine all CSVs into one big table if needed using pandas.
内容的提问来源于stack exchange,提问作者geoastro

