Python新手求助:从原始文本生成键值对字典的文本处理与构建问题
嘿,作为Python新手碰到这种文本解析和字典构建的问题太正常啦!我来一步步帮你搞定~
解决Python文本解析与字典构建的问题
一、先搞定文本解析:拆解你的电影信息
你的示例文本格式是这样的:
Ready Player One (2018) 140 min - Action | Adventure | Sci-Fi Us...
我们需要把里面的电影名、年份、时长、类型这些字段逐一拆出来,这里给你两种方法,新手可以先从基础操作入手,熟练后再用更高效的方式。
方法1:用字符串基础操作逐步拆分
适合刚接触Python的你,每一步都清晰可控:
# 先拿一行示例文本测试,后续替换成逐行读取的内容 line = "Ready Player One (2018) 140 min - Action | Adventure | Sci-Fi" # 第一步:分割基础信息和类型部分 basic_info, genres_part = line.split(" - ", 1) # 只分割第一次,避免类型里的特殊字符干扰 # 第二步:从基础信息里提取电影名和年份 year_start = basic_info.find("(") year_end = basic_info.find(")") year = basic_info[year_start+1:year_end] title = basic_info[:year_start].strip() # 去掉前后多余空格 # 第三步:提取时长 duration = basic_info[year_end+1:].strip() # 第四步:拆分类型为列表 genres = genres_part.split(" | ") # 组装成字典 movie_dict = { "title": title, "year": int(year), # 转成整数更规范 "duration": duration, "genres": genres } print(movie_dict)
运行后会得到规整的字典:
{'title': 'Ready Player One', 'year': 2018, 'duration': '140 min', 'genres': ['Action', 'Adventure', 'Sci-Fi']}
方法2:用正则表达式(灵活适配统一格式的文本)
如果你的所有文本行格式都和示例一致,正则能一步提取所有字段,代码更简洁:
import re line = "Ready Player One (2018) 140 min - Action | Adventure | Sci-Fi" # 编写匹配规则,对应每个字段 pattern = r"(.*?) \((\d{4})\) (\d+ min) - (.*)" match_result = re.match(pattern, line) if match_result: title = match_result.group(1).strip() year = int(match_result.group(2)) duration = match_result.group(3) genres = match_result.group(4).split(" | ") movie_dict = { "title": title, "year": year, "duration": duration, "genres": genres } print(movie_dict)
二、规范管理字典:批量处理与结构化存储
如果是处理整个文本文件,建议把每个电影的字典存到一个列表里,方便后续筛选、统计等操作:
import re # 用来存储所有电影字典的列表 movies_collection = [] # 逐行读取文本文件 with open("你的电影文本文件名.txt", "r", encoding="utf-8") as f: for line in f: line = line.strip() # 去掉换行符和前后空格 if not line: # 跳过空行 continue # 用正则解析每一行 pattern = r"(.*?) \((\d{4})\) (\d+ min) - (.*)" match_result = re.match(pattern, line) if match_result: title = match_result.group(1).strip() year = int(match_result.group(2)) duration = match_result.group(3) genres = match_result.group(4).split(" | ") # 将字典添加到列表 movies_collection.append({ "title": title, "year": year, "duration": duration, "genres": genres }) # 现在movies_collection里就是所有电影的结构化数据啦 print(movies_collection)
小提醒:规范字典键名
尽量用统一的英文键名(比如title、year),不要混用中文和英文,也不要随意更改键名,这样后续处理数据时会更顺畅。如果遇到格式奇怪的行,可以加个try-except捕获错误,避免程序崩溃:
try: # 解析代码放在这里 except: print(f"跳过无法解析的行:{line}")
内容的提问来源于stack exchange,提问作者spider22
相关产品推荐
相关产品推荐

