已完成Excel作者列拆分,求Python将数据保存为JSON的方法
Hey there! Great work getting the author data extracted and split into first/last names—you're almost there. Saving this to JSON is straightforward with Python's built-in json module, and I'll walk you through modifying your code to do it.
Step-by-Step Modifications
- Import the
jsonmodule: It’s part of Python’s standard library, so no extra installation is needed. - Create a list to store structured author data: Instead of just printing each name pair, we’ll collect them into a list of dictionaries—this structure translates perfectly to JSON.
- Write the data to a JSON file: Use
json.dump()to write the list to a file, with optional settings to keep the output readable and preserve special characters.
Modified Full Code
import pandas as pd import json # Add this required import def unique(list1): unique_list = [] for x in list1: if x not in unique_list: unique_list.append(x) return unique_list tbr = pd.read_excel('TBR.xlsx') idx_of_column = 3-1 authors = tbr.iloc[:,idx_of_column] authors_list = authors.values.tolist() cleaned_author_List = [x for x in authors_list if str(x) != 'nan'] unique_cleaned_author_list = unique(cleaned_author_List) # Initialize an empty list to hold our structured author data authors_data = [] for fullname in unique_cleaned_author_list: firstname = fullname.strip().split(' ')[0] lastname = ' '.join((fullname + ' ').split(' ')[1:]).strip() # Add each author's data as a dictionary to the list authors_data.append({ "firstname": firstname, "lastname": lastname }) # Write the data to a JSON file with open('authors.json', 'w', encoding='utf-8') as f: json.dump(authors_data, f, ensure_ascii=False, indent=4) print("Author data saved to authors.json successfully!")
Key Details
ensure_ascii=False: Preserves non-ASCII characters (like accented letters in author names) in the JSON output instead of escaping them.indent=4: Formats the JSON with 4-space indentation, making it easy to read and edit manually.with open(...): This context manager ensures the file is properly closed after writing, which is safe and clean practice.
Optional: Optimize the Unique List
Your unique function works perfectly, but if you don’t need to preserve the original order of authors, you can simplify it using Python’s set:
unique_cleaned_author_list = list(set(cleaned_author_List))
If order matters (Python 3.7+), use dict.fromkeys() to keep the original sequence while removing duplicates:
unique_cleaned_author_list = list(dict.fromkeys(cleaned_author_List))
内容的提问来源于stack exchange,提问作者Aniket Paul
相关产品推荐
相关产品推荐

