You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从train_df的title name列移除年份,保留纯电影名称?

问题描述

I have a DataFrame named train_df with two columns: "gross" and "title name". The dataset is shown below:

gross       title name
760507625.0 Avatar (2009)
658672302.0 Titanic (1997)
652270625.0 Jurassic World (2015)
623357910.0 The Avengers (2012)
534858444.0 The Dark Knight (2008)
532177324.0 Rogue One (2016)
474544677.0 Star Wars: Episode I - The Phantom Menace (1999)
459005868.0 Avengers: Age of Ultron (2015)
448139099.0 The Dark Knight Rises (2012)
436471036.0 Shrek 2 (2004)
424668047.0 The Hunger Games: Catching Fire (2013)
423315812.0 Pirates of the Caribbean: Dead Man's Chest (2006)
415004880.0 Toy Story 3 (2010)
409013994.0 Iron Man 3 (2013)
408084349.0 Captain America: Civil War (2016)
408010692.0 The Hunger Games (2012)
403706375.0 Spider-Man (2002)
402453882.0 Jurassic Park (1993)
402111870.0 Transformers: Revenge of the Fallen (2009)
400738009.0 Frozen (2013)
381011219.0 Harry Potter and the Deathly Hallows: Part 2 (2011)
380843261.0 Finding Nemo (2003)
380262555.0 Star Wars: Episode III - Revenge of the Sith (2005)
373585825.0 Spider-Man 2 (2004)
370782930.0 The Passion of the Christ (2004)

I need to remove the year information (in the format (YYYY)) from the "title name" column, retaining only the movie title. The "gross" column should stay unchanged. The expected output is:

gross   title name
760507625.0 Avatar
658672302.0 Titanic
652270625.0 Jurassic World
623357910.0 The Avengers
534858444.0 The Dark Knight

解决方案

You can use pandas' string manipulation methods to clean the "title name" column. Here are two reliable approaches:

Method 1: Replace the year pattern with an empty string

Use str.replace() with a regular expression to target and remove the (YYYY) segment (including any leading whitespace):

import pandas as pd

# Clean the "title name" column
train_df['title name'] = train_df['title name'].str.replace(r'\s*\(\d{4}\)', '', regex=True)

Explanation:

  • \s*: Matches zero or more whitespace characters before the year parentheses
  • \(\d{4}\): Matches exactly 4 digits enclosed in parentheses (the year format)
  • Replacing this pattern with an empty string removes the year entirely.

Method 2: Extract the title before the year parentheses

Use str.extract() to capture all text before the (YYYY) segment, and handle edge cases where no year exists with fillna():

train_df['title name'] = train_df['title name'].str.extract(r'(.+?)\s*\(', expand=False).fillna(train_df['title name'])

Explanation:

  • (.+?): Non-greedily captures all characters until the next part of the pattern
  • \s*\(: Matches leading whitespace followed by an opening parenthesis
  • fillna(train_df['title name']): Ensures rows without a year retain their original title (adds robustness even if your dataset doesn't have such cases)

验证结果

After running either method, the first 5 rows of train_df will match your expected output:

gross       title name
760507625.0 Avatar
658672302.0 Titanic
652270625.0 Jurassic World
623357910.0 The Avengers
534858444.0 The Dark Knight

内容的提问来源于stack exchange,提问作者user21006068

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 02:10:51