如何用BeautifulSoup爬取信息并生成格式规范的JSON文件
问题:生成的JSON结构不符合预期,重复title且存在多根元素
原执行代码
from bs4 import BeautifulSoup import requests import json url = "https://stock.adobe.com/search?k=interstellar%20movie" page = requests.get(url).text soup = BeautifulSoup(page, "lxml") jfile = open('images.json', 'w') title = soup.find('em', class_='gravel-text').text for images in soup.find_all('div', class_='thumb-frame'): image = images.a['href'] j = [{'title':title, 'image':image}] jstring = json.dumps(j) jfile.write(jstring) jfile.close()
当前生成的错误JSON输出
[ { "title":"interstellar movie", "image":"https://stock.adobe.com/images/gargantua-galaxy-design-graphic-3d-illustration-red-wormhole-or-black-hole-shine-in-space-inspiration-from-interstellar-movie-night-sky-background/90184980" } ][ { "title":"interstellar movie", "image":"https://stock.adobe.com/images/wanderlust-explorer-discovering-icelandic-natural-wonders/395993532" } ][ { "title":"interstellar movie", "image":"https://stock.adobe.com/images/panoramic-beautiful-night-sky-and-star-abstract-background-elements-of-this-image-furnished-by-nasa/223412156" } ]
期望的JSON格式
[{ "title": "interstellar movie", "image": [ "https://stock.adobe.com/images/gargantua-galaxy-design-graphic-3d-illustration-red-wormhole-or-black-hole-shine-in-space-inspiration-from-interstellar-movie-night-sky-background/90184980", "https://stock.adobe.com/images/wanderlust-explorer-discovering-icelandic-natural-wonders/395993532", "https://stock.adobe.com/images/panoramic-beautiful-night-sky-and-star-abstract-background-elements-of-this-image-furnished-by-nasa/223412156" ] }]
修正方案
原代码的问题在于每次循环都会创建一个新的单元素列表并写入文件,导致最终文件是多个JSON结构拼接,同时重复写入title。正确的做法是先构建一个包含title和空图片列表的字典,在循环中收集所有图片链接,最后一次性生成完整的JSON并写入:
from bs4 import BeautifulSoup import requests import json url = "https://stock.adobe.com/search?k=interstellar%20movie" page = requests.get(url).text soup = BeautifulSoup(page, "lxml") title = soup.find('em', class_='gravel-text').text # 初始化目标数据结构:包含title和空图片列表 result = [{'title': title, 'image': []}] for images in soup.find_all('div', class_='thumb-frame'): image = images.a['href'] # 将图片链接添加到列表中 result[0]['image'].append(image) # 用with语句管理文件,一次性写入完整JSON结构 with open('images.json', 'w') as jfile: json.dump(result, jfile, indent=4)
说明
- 使用
with语句管理文件,无需手动调用close(),更安全简洁 - 先初始化好目标结构,循环中仅做链接收集,避免重复创建对象和多次写入
json.dump直接将对象写入文件,比先转字符串再写入更高效,indent参数可生成格式化的JSON,提升可读性
内容的提问来源于stack exchange,提问作者Ryan
相关产品推荐
相关产品推荐

