Prolog读取CSV时maplist(atom_string)失效 如何筛选mateIn1标签行
问题背景
以下为sorted.csv文件示例内容:
01Vbe,2r5/3Qnk1p/8/4B2b/Pp2p3/1P2P3/5PPP/3R2K1 w - - 3 32,d1d6 c8c1 d6d1 c1d1 d7d1 h5d1,600,80,100,633,advantage endgame fork long,https://lichess.org/b334W1Ga#63 01GC2,3b4/pp2kprp/8/1Bp5/4R3/1P6/P4PPP/1K6 b - - 0 22,e7f8 e4e8,600,82,89,374,endgame mate mateIn1 oneMove,https://lichess.org/6c75Zk3T/black#44
file_line为逐行读取文件的谓词,用于读取CSV文件内容。
当前编写的Prolog实现代码如下:
stuff(Lines) :- findall(Tags, ( file_line("data/athousand_sorted.csv", Line), csv_fen(Line, Fen, Moves, Id, Tags), maplist(atom_string, TagsAtom, Tags), tags_with(TagsAtom, mateIn1) ), Lines). csv_fen(Line, Fen, Moves, Id, Tags) :- split_string(Line, ",", "", [Id, Fen, Moves, _, _, _, _, Tags| _]). tags_with(Tags, Tag) :- split_string(Tags, " ", "", Ls), member(Tag, Ls).
故障表现
从文件读取得到的字符串传入逻辑时,maplist(atom_string, XYZ)无法正常运行,但直接在REPL中执行?- maplist(atom_string, Tags, ["mateIn1"]).可正常工作,预期功能为筛选出CSV文件中包含mateIn1标签的行。
故障原因
- CSV字段拆分规则错误:每行CSV在
Moves(走法序列)和Tags(标签)之间共有4个数值字段,原代码仅设置3个占位符跳过字段,导致Tags变量实际取到的是数值字段内容,不是列表类型,不符合maplist要求传入列表的参数规则。 - 类型转换逻辑冗余且不匹配:
Tags从CSV中读取到的是空格分隔的标签整串,不需要通过maplist转换为原子列表;且后续tags_with谓词调用的split_string要求第一个入参为字符串类型,传入原子会触发类型错误。
修复代码
修正字段拆分规则,删除冗余的类型转换逻辑,统一使用字符串做标签匹配,避免类型不兼容问题:
% 筛选所有包含mateIn1标签的行 stuff(Lines) :- findall(Line, ( file_line("data/athousand_sorted.csv", Line), csv_fen(Line, _Fen, _Moves, _Id, Tags), tags_with(Tags, "mateIn1") ), Lines). % 解析CSV行字段:Moves后共4个数值字段,依次跳过,提取标签字段后忽略后续URL等内容 csv_fen(Line, Fen, Moves, Id, Tags) :- split_string(Line, ",", "", [Id, Fen, Moves, _, _, _, _, Tags, _|_]). % 判断标签串是否包含目标标签 tags_with(TagsStr, TargetTag) :- split_string(TagsStr, " ", "", TagList), member(TargetTag, TagList).
如果需要返回指定字段而非整行内容,只需修改findall的第一个参数,组合需要的字段即可。
内容的提问来源于stack exchange,提问作者eguneys
相关产品推荐
相关产品推荐

