You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:从原始文本生成键值对字典的文本处理与构建问题

嘿,作为Python新手碰到这种文本解析和字典构建的问题太正常啦!我来一步步帮你搞定~

解决Python文本解析与字典构建的问题

一、先搞定文本解析:拆解你的电影信息

你的示例文本格式是这样的:

Ready Player One (2018) 140 min - Action | Adventure | Sci-Fi Us...

我们需要把里面的电影名、年份、时长、类型这些字段逐一拆出来,这里给你两种方法,新手可以先从基础操作入手,熟练后再用更高效的方式。

方法1:用字符串基础操作逐步拆分

适合刚接触Python的你,每一步都清晰可控:

# 先拿一行示例文本测试,后续替换成逐行读取的内容
line = "Ready Player One (2018) 140 min - Action | Adventure | Sci-Fi"

# 第一步:分割基础信息和类型部分
basic_info, genres_part = line.split(" - ", 1)  # 只分割第一次,避免类型里的特殊字符干扰

# 第二步:从基础信息里提取电影名和年份
year_start = basic_info.find("(")
year_end = basic_info.find(")")
year = basic_info[year_start+1:year_end]
title = basic_info[:year_start].strip()  # 去掉前后多余空格

# 第三步:提取时长
duration = basic_info[year_end+1:].strip()

# 第四步:拆分类型为列表
genres = genres_part.split(" | ")

# 组装成字典
movie_dict = {
    "title": title,
    "year": int(year),  # 转成整数更规范
    "duration": duration,
    "genres": genres
}

print(movie_dict)

运行后会得到规整的字典:

{'title': 'Ready Player One', 'year': 2018, 'duration': '140 min', 'genres': ['Action', 'Adventure', 'Sci-Fi']}

方法2:用正则表达式(灵活适配统一格式的文本)

如果你的所有文本行格式都和示例一致,正则能一步提取所有字段,代码更简洁:

import re

line = "Ready Player One (2018) 140 min - Action | Adventure | Sci-Fi"

# 编写匹配规则,对应每个字段
pattern = r"(.*?) \((\d{4})\) (\d+ min) - (.*)"
match_result = re.match(pattern, line)

if match_result:
    title = match_result.group(1).strip()
    year = int(match_result.group(2))
    duration = match_result.group(3)
    genres = match_result.group(4).split(" | ")
    
    movie_dict = {
        "title": title,
        "year": year,
        "duration": duration,
        "genres": genres
    }
    print(movie_dict)

二、规范管理字典:批量处理与结构化存储

如果是处理整个文本文件,建议把每个电影的字典存到一个列表里,方便后续筛选、统计等操作:

import re

# 用来存储所有电影字典的列表
movies_collection = []

# 逐行读取文本文件
with open("你的电影文本文件名.txt", "r", encoding="utf-8") as f:
    for line in f:
        line = line.strip()  # 去掉换行符和前后空格
        if not line:  # 跳过空行
            continue
        
        # 用正则解析每一行
        pattern = r"(.*?) \((\d{4})\) (\d+ min) - (.*)"
        match_result = re.match(pattern, line)
        if match_result:
            title = match_result.group(1).strip()
            year = int(match_result.group(2))
            duration = match_result.group(3)
            genres = match_result.group(4).split(" | ")
            
            # 将字典添加到列表
            movies_collection.append({
                "title": title,
                "year": year,
                "duration": duration,
                "genres": genres
            })

# 现在movies_collection里就是所有电影的结构化数据啦
print(movies_collection)

小提醒:规范字典键名

尽量用统一的英文键名(比如title、year),不要混用中文和英文,也不要随意更改键名,这样后续处理数据时会更顺畅。如果遇到格式奇怪的行,可以加个try-except捕获错误,避免程序崩溃:

try:
    # 解析代码放在这里
except:
    print(f"跳过无法解析的行:{line}")

内容的提问来源于stack exchange,提问作者spider22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:07:34