You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

已完成Excel作者列拆分,求Python将数据保存为JSON的方法

How to Save Split Author Data to JSON File

Hey there! Great work getting the author data extracted and split into first/last names—you're almost there. Saving this to JSON is straightforward with Python's built-in json module, and I'll walk you through modifying your code to do it.

Step-by-Step Modifications

  1. Import the json module: It’s part of Python’s standard library, so no extra installation is needed.
  2. Create a list to store structured author data: Instead of just printing each name pair, we’ll collect them into a list of dictionaries—this structure translates perfectly to JSON.
  3. Write the data to a JSON file: Use json.dump() to write the list to a file, with optional settings to keep the output readable and preserve special characters.

Modified Full Code

import pandas as pd
import json  # Add this required import

def unique(list1):
    unique_list = []
    for x in list1:
        if x not in unique_list:
            unique_list.append(x)
    return unique_list

tbr = pd.read_excel('TBR.xlsx')
idx_of_column = 3-1
authors = tbr.iloc[:,idx_of_column]

authors_list = authors.values.tolist()
cleaned_author_List = [x for x in authors_list if str(x) != 'nan']
unique_cleaned_author_list = unique(cleaned_author_List)

# Initialize an empty list to hold our structured author data
authors_data = []

for fullname in unique_cleaned_author_list:
    firstname = fullname.strip().split(' ')[0]
    lastname = ' '.join((fullname + ' ').split(' ')[1:]).strip()
    # Add each author's data as a dictionary to the list
    authors_data.append({
        "firstname": firstname,
        "lastname": lastname
    })

# Write the data to a JSON file
with open('authors.json', 'w', encoding='utf-8') as f:
    json.dump(authors_data, f, ensure_ascii=False, indent=4)

print("Author data saved to authors.json successfully!")

Key Details

  • ensure_ascii=False: Preserves non-ASCII characters (like accented letters in author names) in the JSON output instead of escaping them.
  • indent=4: Formats the JSON with 4-space indentation, making it easy to read and edit manually.
  • with open(...): This context manager ensures the file is properly closed after writing, which is safe and clean practice.

Optional: Optimize the Unique List

Your unique function works perfectly, but if you don’t need to preserve the original order of authors, you can simplify it using Python’s set:

unique_cleaned_author_list = list(set(cleaned_author_List))

If order matters (Python 3.7+), use dict.fromkeys() to keep the original sequence while removing duplicates:

unique_cleaned_author_list = list(dict.fromkeys(cleaned_author_List))

内容的提问来源于stack exchange,提问作者Aniket Paul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 10:52:28