You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中从CSV的title列提取年份创建新列?

从CSV的title列提取年份并创建新列的解决方案

方法一:基于原生csv模块实现

基于你当前使用的Python原生csv模块,修改代码即可完成需求。核心是通过正则精准匹配title中的四位年份,再将年份添加到每行数据末尾,最终写入新文件:

import csv
import re
from itertools import islice

# 提取年份的工具函数
def get_year_from_title(title):
    # 匹配括号内的四位数字(年份格式)
    year_match = re.search(r'\((\d{4})\)', title)
    return year_match.group(1) if year_match else ''

# 读取原文件并写入带年份列的新文件
with open('/home/raymondossai/movies.csv', mode='r') as input_file, \
     open('/home/raymondossai/movies_with_year.csv', mode='w', newline='') as output_file:
    csv_reader = csv.reader(input_file)
    csv_writer = csv.writer(output_file)
    
    # 处理表头:添加year列
    header_row = next(csv_reader)
    header_row.append('year')
    csv_writer.writerow(header_row)
    
    # 逐行处理数据
    for row in csv_reader:
        movie_title = row[1]
        extracted_year = get_year_from_title(movie_title)
        row.append(extracted_year)
        csv_writer.writerow(row)

# 验证结果:打印新文件的前11行(含表头)
with open('/home/raymondossai/movies_with_year.csv', mode='r') as file:
    for row in islice(csv.reader(file), 11):
        print(row)

运行后会生成新的movies_with_year.csv,其中新增的year列就是从title中提取的年份。正则表达式r'\((\d{4})\)'能精准匹配括号中的四位数字,避免title中其他括号内容的干扰。


方法二:使用pandas库(更简洁高效)

如果可以安装pandas库,处理这类表格数据会更简便,一行代码就能完成年份提取:

import pandas as pd

# 读取CSV文件
df = pd.read_csv('/home/raymondossai/movies.csv')

# 从title列提取年份,生成新的year列
df['year'] = df['title'].str.extract(r'\((\d{4})\)', expand=False)

# 保存包含年份列的新CSV文件(不保留索引)
df.to_csv('/home/raymondossai/movies_with_year.csv', index=False)

# 打印前10行数据验证
print(df.head(10))

pandas的str.extract方法会自动遍历整列数据,提取符合正则规则的内容,非常适合批量处理表格数据。


内容的提问来源于stack exchange,提问作者Raymond Ossai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 07:10:28