Python网页爬虫仅写入首行至CSV,循环遍历存CSV失败求助
Hey there! Super common issue when you're just getting started with web scraping and CSV exports—let's break this down step by step.
First, let's recap your problem: you're successfully scraping all the data into an array of objects, but when you try to write to CSV, only the header and first entry show up. Chances are, the issue is in how you're handling the CSV writing loop or file access.
Common Causes & Fixes
Let's start with the most likely culprits, then walk through a complete working example using your code.
1. You're opening the file inside your loop (or overwriting it accidentally)
If you have code like this:
# ❌ Bad: Opens file every time, overwriting content for entry in team_data: with open('stats.csv', 'w') as f: writer = csv.writer(f) writer.writerow(entry.values())
Each time you open the file with 'w' mode, it clears the entire file. Even if you intended to write multiple entries, this approach would only leave the last one (or in your case, maybe you only called the write function once outside the loop).
2. You're not using the csv module correctly (or skipping the loop entirely)
Manual string concatenation for CSVs is error-prone, and it's easy to accidentally only write one entry. The built-in csv module handles all the formatting (like commas, quotes, newlines) for you, so it's the safest way to export data.
Complete Working Example
Here's how to adjust your code to scrape the Air Force football stats and write all entries to CSV:
from urllib.request import urlopen as uReq from bs4 import BeautifulSoup as soup import csv # Don't forget this essential module! my_url = 'https://www.sports-reference.com/cfb/schools/air-force/' # Step 1: Grab and parse the page uClient = uReq(my_url) page_html = uClient.read() uClient.close() page_soup = soup(page_html, "html.parser") # Step 2: Extract data into an array of dictionaries # Targeting the yearly stats table (adjust if you're scraping different data) stats_table = page_soup.find('table', {'id': 'school_yearly'}) rows = stats_table.find_all('tr')[1:] # Skip the table's header row team_data = [] for row in rows: cols = row.find_all('td') # Skip empty rows that might be in the table if cols: year_entry = { 'Year': cols[0].text.strip(), 'Wins': cols[1].text.strip(), 'Losses': cols[2].text.strip(), 'Ties': cols[3].text.strip(), 'Points For': cols[4].text.strip(), 'Points Against': cols[5].text.strip() } team_data.append(year_entry) # Step 3: Write ALL data to CSV if team_data: # Make sure we have data to write # Get headers from the first entry's keys headers = team_data[0].keys() # Open file with newline='' to avoid extra blank lines in CSV with open('air_force_football_stats.csv', 'w', newline='', encoding='utf-8') as csv_file: # Use DictWriter to map dictionary keys to CSV columns writer = csv.DictWriter(csv_file, fieldnames=headers) writer.writeheader() # Write the header row # Loop through EVERY entry in your array for entry in team_data: writer.writerow(entry) # Optional: Print to verify you're processing all entries print(f"Written entry: {entry['Year']}") print("Done! All data saved to CSV.") else: print("No data found to write.")
How to Troubleshoot Your Existing Code
If you want to fix your current code instead of starting fresh:
- First, check how many items are in your array with
print(len(your_array_name)). If it returns1, the problem is in your scraping code (you're not actually collecting all entries). - If the length is greater than 1, check your CSV writing loop:
- Add a print statement inside the loop (like
print(entry)) to confirm you're iterating over every item. - Ensure you're opening the file once before the loop (not inside it) using
'w'mode—this way, you don't overwrite the file each time you write an entry.
- Add a print statement inside the loop (like
Let me know if you hit any roadblocks with your specific code—I'm happy to help tweak it further!
内容的提问来源于stack exchange,提问作者Eric Baker

