如何在Python中将CSV内TMDB数据集的JSON分类列转成字符串
genres Column into Semicolon-Separated Strings Hey there! No worries at all—we all start somewhere, and this is a super common task when working with movie datasets like TMDB. Let's break this down step by step so you can get that genres column formatted exactly how you want it.
Step 1: Add the json module to your imports
First, we need a way to turn those JSON string entries into actual Python data structures (lists of dictionaries). Python's built-in json module does this perfectly, so add it to your existing imports:
import numpy as np import matplotlib.pyplot as plt import pandas as pd import json # New import to handle JSON parsing!
Step 2: Create a helper function to parse each genre entry
We'll write a small function that takes one JSON string from the genres column, extracts all the genre names, and joins them with semicolons (plus an extra semicolon at the end, just like your example shows):
def parse_genres(genres_str): # Handle empty/missing values (in case some rows have no genres listed) if pd.isna(genres_str): return '' # Convert the JSON string to a Python list of dictionaries genres_list = json.loads(genres_str) # Extract just the 'name' value from each genre dictionary genre_names = [genre['name'] for genre in genres_list] # Join the names with semicolons and add a final semicolon to match your desired format return ';'.join(genre_names) + ';'
Step 3: Apply the function to your genres column
Now we'll use Pandas' apply() method to run this function on every row in the genres column. You can either replace the original column or create a new one—let's replace the original for simplicity:
# Load your dataset (same as your original code) reviews = pd.read_csv("C:/Users/HP/Desktop/data science project/tmpd second attempt/2. prepared data/tmdb_5000_movies.csv") # Apply the parsing function to clean up the genres column reviews["genres"] = reviews["genres"].apply(parse_genres) # Check the result to make sure it's working! print(reviews["genres"])
What this does (in plain terms):
json.loads(genres_str)turns the messy string like[{"id":28,"name":"Action"},...]into a list of dictionaries we can easily work with.- The list comprehension
[genre['name'] for genre in genres_list]grabs just the genre names (like "Action", "Adventure") from each dictionary in the list. ';'.join(genre_names)sticks those names together into a single string separated by semicolons, and we add an extra;at the end to match your requested output.
Example Output:
Running this code will turn your original genres entries into exactly what you wanted:
0 Action;Adventure;Fantasy;Science Fiction; 1 Adventure;Fantasy;Action; 2 Action;Adventure;Crime; 3 Action;Crime;Drama;Thriller; ...
No need to stress about being a beginner—this is a great first step in cleaning messy dataset columns, and you're already doing awesome by tackling this project!
内容的提问来源于stack exchange,提问作者Ankit Kumar

