You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按多条件筛选CSV电影数据集?解决类型转换问题

问题说明

现有film.csv数据集结构如下:

Year;Length;Title;Subject;Actor;Actress;Director;Popularity;Awards;*Image
1990;111;Tie Me Up! Tie Me Down!;Comedy;Banderas, Antonio;Abril, Victoria;Almodóvar, Pedro;68;No;NicholasCage.png
1991;113;High Heels;Comedy;Bosé, Miguel;Abril, Victoria;Almodóvar, Pedro;68;No;NicholasCage.png
1983;104;Dead Zone, The;Horror;Walken, Christopher;Adams, Brooke;Cronenberg, David;79;No;NicholasCage.png
1979;122;Cuba;Action;Connery, Sean;Adams, Brooke;Lester, Richard;6;No;seanConnery.png
1978;94;Days of Heaven;Drama;Gere, Richard;Adams, Brooke;Malick, Terrence;14;No;NicholasCage.png
1983;140;Octopussy;Action;Moore, Roger;Adams, Maud;Glen, John;68;No;NicholasCage.png

需要筛选同时满足以下三个条件的电影标题:

  • Actor字段包含"Richard"
  • Year小于1985
  • Awards等于"Y"

现有代码仅能实现Awards条件的筛选,添加Year条件时出现字符串与整数类型不兼容错误,且不清楚如何组合多条件,需修正代码。

现有代码:

file_name = "film.csv"
lines = (line for line in open(file_name,encoding='cp1252')) #generator to capture lines
lists = (s.rstrip().split(";")) for s in lines) #generators to capture lists containing values from lines

#browse lists and index them per header values, then filter all movies that have been awarded
#using a new generator object

cols=next(lists) #obtains only the header
print(cols)
collections = (dict(zip(cols,data)) for data in lists)
    
filtered = (col["Title"] for col in collections if col["Awards"][0] == "Y")
                                                
for item in filtered:
        print(item)
    #   input()

修正后的代码

file_name = "film.csv"
# 生成器读取文件行
lines = (line for line in open(file_name, encoding='cp1252'))
# 分割每行数据为列表(修复原代码的语法错误:多余的右括号)
lists = (s.rstrip().split(";") for s in lines)

# 获取表头字段
cols = next(lists)
# 将每行数据转为字典格式,方便按字段名调用
collections = (dict(zip(cols, data)) for data in lists)

# 多条件组合筛选
filtered = (
    col["Title"] 
    for col in collections 
    if col["Awards"] == "Y"  # 直接匹配奖项字段值
    and int(col["Year"]) < 1985  # 将年份字符串转为整数后比较,解决类型不兼容问题
    and "Richard" in col["Actor"]  # 检查演员字段是否包含目标关键词
)

# 输出筛选结果
for item in filtered:
    print(item)

关键修复点

  1. 修复语法错误:原代码中lists定义多了一个右括号,导致运行报错,已修正
  2. 类型转换处理:从CSV读取的Year是字符串类型,必须用int()转为整数后,才能和数值1985做大小比较
  3. 多条件逻辑组合:用and连接三个筛选条件,确保同时满足所有要求
  4. 简化奖项判断:直接用col["Awards"] == "Y"即可,无需取索引[0](给定数据集中奖项字段值为明确的"No"/"Y")

内容的提问来源于stack exchange,提问作者Tristan Houghton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 05:46:09