You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取数据存JSON仅保留最后一项的问题排查

Fix: Saving All Scraped Items to JSON Instead of Just the Last One

Hey there! I’ve run into this exact issue dozens of times—your script is probably overwriting your data variable every time it loops through an item, instead of collecting all entries into a list. Let’s break down how to fix this quickly.

The Root Cause

When you print each item in the terminal, it works because you’re outputting each entry as you process it. But if you’re initializing your data structure (like a dictionary) inside the loop or replacing it each time, you’ll only end up with the last item when you save to JSON.

Corrected Code Example

Here’s how to adjust your script to collect and save all items properly:

from urllib.request import urlopen
from bs4 import BeautifulSoup as soup
import json

otl_url = 'https://open.umn.edu/opentextbooks/SearchResults.aspx?subjectAre...'

# 1. Initialize an empty list to hold ALL scraped items
all_textbooks = []

# 2. Fetch and parse the webpage
uClient = urlopen(otl_url)
page_html = uClient.read()
uClient.close()
page_soup = soup(page_html, "html.parser")

# 3. Loop through each textbook item on the page
# (Replace the selector below with whatever you're using to target items)
for textbook in page_soup.find_all("div", class_="textbook-entry"):
    # Extract your desired fields (adjust selectors to match your page)
    title = textbook.find("h2").text.strip()
    author = textbook.find("p", class_="author-name").text.strip()
    url = textbook.find("a")["href"]
    
    # 4. Create a dictionary for this item and add it to the list
    textbook_entry = {
        "title": title,
        "author": author,
        "url": url
    }
    all_textbooks.append(textbook_entry)
    
    # Verify in terminal (this should still print every item)
    print(textbook_entry)

# 5. Save the entire list to JSON
with open("open_textbooks.json", "w") as outfile:
    json.dump(all_textbooks, outfile, indent=4)

print(f"Successfully saved {len(all_textbooks)} textbooks to open_textbooks.json!")

Key Fixes Explained

  • List initialization outside the loop: all_textbooks = [] creates a persistent container to collect every item, instead of overwriting a single variable each iteration.
  • Append items to the list: all_textbooks.append(textbook_entry) adds each new dictionary to our collection, rather than replacing the previous one.
  • Dump the full list to JSON: When we call json.dump(all_textbooks, outfile), we’re saving the entire set of items, not just the last one processed.

Quick Check for Your Original Code

If your original script looked something like this (the mistake):

# Wrong: Overwrites 'data' every loop
data = {}
for item in page_soup.find_all(...):
    data["title"] = item.find("h3").text
# Only saves the last item
json.dump(data, outfile)

Swapping that single dictionary for a list and appending each entry will fix the problem instantly.

Let me know if you need help tweaking the selectors to match the exact structure of the UMN page!

内容的提问来源于stack exchange,提问作者brawlins4

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:25:24