You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中将CSV内TMDB数据集的JSON分类列转成字符串

How to Parse TMDB's JSON-formatted genres Column into Semicolon-Separated Strings

Hey there! No worries at all—we all start somewhere, and this is a super common task when working with movie datasets like TMDB. Let's break this down step by step so you can get that genres column formatted exactly how you want it.

Step 1: Add the json module to your imports

First, we need a way to turn those JSON string entries into actual Python data structures (lists of dictionaries). Python's built-in json module does this perfectly, so add it to your existing imports:

import numpy as np
import matplotlib.pyplot as plt
import pandas as pd
import json  # New import to handle JSON parsing!

Step 2: Create a helper function to parse each genre entry

We'll write a small function that takes one JSON string from the genres column, extracts all the genre names, and joins them with semicolons (plus an extra semicolon at the end, just like your example shows):

def parse_genres(genres_str):
    # Handle empty/missing values (in case some rows have no genres listed)
    if pd.isna(genres_str):
        return ''
    
    # Convert the JSON string to a Python list of dictionaries
    genres_list = json.loads(genres_str)
    
    # Extract just the 'name' value from each genre dictionary
    genre_names = [genre['name'] for genre in genres_list]
    
    # Join the names with semicolons and add a final semicolon to match your desired format
    return ';'.join(genre_names) + ';'

Step 3: Apply the function to your genres column

Now we'll use Pandas' apply() method to run this function on every row in the genres column. You can either replace the original column or create a new one—let's replace the original for simplicity:

# Load your dataset (same as your original code)
reviews = pd.read_csv("C:/Users/HP/Desktop/data science project/tmpd second attempt/2. prepared data/tmdb_5000_movies.csv")

# Apply the parsing function to clean up the genres column
reviews["genres"] = reviews["genres"].apply(parse_genres)

# Check the result to make sure it's working!
print(reviews["genres"])

What this does (in plain terms):

  • json.loads(genres_str) turns the messy string like [{"id":28,"name":"Action"},...] into a list of dictionaries we can easily work with.
  • The list comprehension [genre['name'] for genre in genres_list] grabs just the genre names (like "Action", "Adventure") from each dictionary in the list.
  • ';'.join(genre_names) sticks those names together into a single string separated by semicolons, and we add an extra ; at the end to match your requested output.

Example Output:

Running this code will turn your original genres entries into exactly what you wanted:

0    Action;Adventure;Fantasy;Science Fiction;
1    Adventure;Fantasy;Action;
2    Action;Adventure;Crime;
3    Action;Crime;Drama;Thriller;
...

No need to stress about being a beginner—this is a great first step in cleaning messy dataset columns, and you're already doing awesome by tackling this project!

内容的提问来源于stack exchange,提问作者Ankit Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:07:56