You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的re模块正则表达式提取TXT文件中的电影分类?

提取电影分类列表的Python解决方案

需求说明

从给定格式的TXT文件中提取所有电影分类,输出为无重复的列表格式(类似['action', 'comedy', 'crime', 'drama', 'thriller'])。

输入文件示例

m0 +++$+++ 10 things i hate about you +++$+++ 1999 +++$+++ 6.90 +++$+++ 62847 +++$+++ ['comedy', 'romance']
m1 +++$+++ 1492: conquest of paradise +++$+++ 1992 +++$+++ 6.20 +++$+++ 10421 +++$+++ ['adventure', 'biography', 'drama', 'history']

用户原代码

import re

file = open('datasets/movie_titles_metadata.txt')

def extract_categories(file):

    for line in file:
        line: str = line.rstrip()
        if re.search(" ", line):
            line = re.sub(r"[0-9]", "", line)
            line = re.sub(r"[$ + : . ]", "", line)
            return line
        
      

extract_categories(file) 

原代码问题分析

  • 无差别正则替换直接破坏了分类的结构,无法保留有效分类信息
  • 循环仅处理第一行就执行return,无法遍历所有行提取分类
  • 没有针对性地定位到每行末尾的分类字段

解决方案

方法一:按固定分隔符拆分(推荐)

利用文件中固定的分隔符+++$+++拆分每行,直接提取最后一个字段的分类内容:

import ast

def extract_categories(file_path):
    all_categories = []
    # 使用with语句自动管理文件资源
    with open(file_path, 'r', encoding='utf-8') as file:
        for line in file:
            line = line.strip()
            if not line:
                continue
            # 按分隔符拆分每行内容
            parts = line.split('+++$+++')
            # 提取最后一个字段并去除前后空格
            category_str = parts[-1].strip()
            # 将字符串形式的列表转为真实Python列表
            categories = ast.literal_eval(category_str)
            # 将当前行的分类加入总列表
            all_categories.extend(categories)
    # 去重并保持原有顺序(Python 3.7+字典默认有序)
    unique_categories = list(dict.fromkeys(all_categories))
    return unique_categories

# 调用函数并打印结果
result = extract_categories('datasets/movie_titles_metadata.txt')
print(result)

方法二:正则匹配分类字段

如果偏好使用正则,可以直接匹配每行末尾的['...']格式内容:

import re
import ast

def extract_categories(file_path):
    all_categories = []
    # 匹配末尾的列表格式字符串
    pattern = r"\['.*?'\]$"
    with open(file_path, 'r', encoding='utf-8') as file:
        for line in file:
            line = line.strip()
            if not line:
                continue
            match = re.search(pattern, line)
            if match:
                category_str = match.group()
                categories = ast.literal_eval(category_str)
                all_categories.extend(categories)
    # 去重处理
    unique_categories = list(dict.fromkeys(all_categories))
    return unique_categories

result = extract_categories('datasets/movie_titles_metadata.txt')
print(result)

关键要点

  • 使用with语句打开文件,避免手动关闭文件的遗漏问题
  • 固定分隔符拆分比正则更高效,适合格式明确的文本
  • ast.literal_eval是安全解析字符串形式列表的方法,避免eval()的安全风险
  • 加入去重逻辑,确保最终列表中每个分类只出现一次

内容的提问来源于stack exchange,提问作者mongol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 15:55:33