使用BeautifulSoup爬取数据存JSON仅保留最后一项的问题排查
Hey there! I’ve run into this exact issue dozens of times—your script is probably overwriting your data variable every time it loops through an item, instead of collecting all entries into a list. Let’s break down how to fix this quickly.
The Root Cause
When you print each item in the terminal, it works because you’re outputting each entry as you process it. But if you’re initializing your data structure (like a dictionary) inside the loop or replacing it each time, you’ll only end up with the last item when you save to JSON.
Corrected Code Example
Here’s how to adjust your script to collect and save all items properly:
from urllib.request import urlopen from bs4 import BeautifulSoup as soup import json otl_url = 'https://open.umn.edu/opentextbooks/SearchResults.aspx?subjectAre...' # 1. Initialize an empty list to hold ALL scraped items all_textbooks = [] # 2. Fetch and parse the webpage uClient = urlopen(otl_url) page_html = uClient.read() uClient.close() page_soup = soup(page_html, "html.parser") # 3. Loop through each textbook item on the page # (Replace the selector below with whatever you're using to target items) for textbook in page_soup.find_all("div", class_="textbook-entry"): # Extract your desired fields (adjust selectors to match your page) title = textbook.find("h2").text.strip() author = textbook.find("p", class_="author-name").text.strip() url = textbook.find("a")["href"] # 4. Create a dictionary for this item and add it to the list textbook_entry = { "title": title, "author": author, "url": url } all_textbooks.append(textbook_entry) # Verify in terminal (this should still print every item) print(textbook_entry) # 5. Save the entire list to JSON with open("open_textbooks.json", "w") as outfile: json.dump(all_textbooks, outfile, indent=4) print(f"Successfully saved {len(all_textbooks)} textbooks to open_textbooks.json!")
Key Fixes Explained
- List initialization outside the loop:
all_textbooks = []creates a persistent container to collect every item, instead of overwriting a single variable each iteration. - Append items to the list:
all_textbooks.append(textbook_entry)adds each new dictionary to our collection, rather than replacing the previous one. - Dump the full list to JSON: When we call
json.dump(all_textbooks, outfile), we’re saving the entire set of items, not just the last one processed.
Quick Check for Your Original Code
If your original script looked something like this (the mistake):
# Wrong: Overwrites 'data' every loop data = {} for item in page_soup.find_all(...): data["title"] = item.find("h3").text # Only saves the last item json.dump(data, outfile)
Swapping that single dictionary for a list and appending each entry will fix the problem instantly.
Let me know if you need help tweaking the selectors to match the exact structure of the UMN page!
内容的提问来源于stack exchange,提问作者brawlins4

