如何从train_df的title name列移除年份,保留纯电影名称?
I have a DataFrame named train_df with two columns: "gross" and "title name". The dataset is shown below:
gross title name 760507625.0 Avatar (2009) 658672302.0 Titanic (1997) 652270625.0 Jurassic World (2015) 623357910.0 The Avengers (2012) 534858444.0 The Dark Knight (2008) 532177324.0 Rogue One (2016) 474544677.0 Star Wars: Episode I - The Phantom Menace (1999) 459005868.0 Avengers: Age of Ultron (2015) 448139099.0 The Dark Knight Rises (2012) 436471036.0 Shrek 2 (2004) 424668047.0 The Hunger Games: Catching Fire (2013) 423315812.0 Pirates of the Caribbean: Dead Man's Chest (2006) 415004880.0 Toy Story 3 (2010) 409013994.0 Iron Man 3 (2013) 408084349.0 Captain America: Civil War (2016) 408010692.0 The Hunger Games (2012) 403706375.0 Spider-Man (2002) 402453882.0 Jurassic Park (1993) 402111870.0 Transformers: Revenge of the Fallen (2009) 400738009.0 Frozen (2013) 381011219.0 Harry Potter and the Deathly Hallows: Part 2 (2011) 380843261.0 Finding Nemo (2003) 380262555.0 Star Wars: Episode III - Revenge of the Sith (2005) 373585825.0 Spider-Man 2 (2004) 370782930.0 The Passion of the Christ (2004)
I need to remove the year information (in the format (YYYY)) from the "title name" column, retaining only the movie title. The "gross" column should stay unchanged. The expected output is:
gross title name 760507625.0 Avatar 658672302.0 Titanic 652270625.0 Jurassic World 623357910.0 The Avengers 534858444.0 The Dark Knight
You can use pandas' string manipulation methods to clean the "title name" column. Here are two reliable approaches:
Method 1: Replace the year pattern with an empty string
Use str.replace() with a regular expression to target and remove the (YYYY) segment (including any leading whitespace):
import pandas as pd # Clean the "title name" column train_df['title name'] = train_df['title name'].str.replace(r'\s*\(\d{4}\)', '', regex=True)
Explanation:
\s*: Matches zero or more whitespace characters before the year parentheses\(\d{4}\): Matches exactly 4 digits enclosed in parentheses (the year format)- Replacing this pattern with an empty string removes the year entirely.
Method 2: Extract the title before the year parentheses
Use str.extract() to capture all text before the (YYYY) segment, and handle edge cases where no year exists with fillna():
train_df['title name'] = train_df['title name'].str.extract(r'(.+?)\s*\(', expand=False).fillna(train_df['title name'])
Explanation:
(.+?): Non-greedily captures all characters until the next part of the pattern\s*\(: Matches leading whitespace followed by an opening parenthesisfillna(train_df['title name']): Ensures rows without a year retain their original title (adds robustness even if your dataset doesn't have such cases)
After running either method, the first 5 rows of train_df will match your expected output:
gross title name 760507625.0 Avatar 658672302.0 Titanic 652270625.0 Jurassic World 623357910.0 The Avengers 534858444.0 The Dark Knight
内容的提问来源于stack exchange,提问作者user21006068

