使用字符串数组过滤ASCII列表中的数字字符串问题
问题:过滤ASCII列表中的数字字符无效
我想用自定义字符串数组过滤包含ASCII字符的列表,移除所有数字字符串,但写的Python代码没达到效果,输出里还是有数字。
原代码
import pandas as pd with open('ASCII.txt') as f: data = f.read().replace('\t', ',') print(data, file=open('my_file.csv', 'w')) df = list(data) test = ['0','1','2','3','4','5','6','7','8','9'] for x in df: try: df = int(df) for i in range(0,9): while any(test) in df: df.remove('i') print(df) except: continue print(df)
当前输出
['3', '3', ',', '0', '4', '1', ',', '2', '1', ',', '!', ',', '\n', '3', '4', ',', '0', '4', ...]
原代码问题分析
- 错误转换列表为整数:
df = int(df)试图把整个列表转成整数,完全不符合逻辑,直接触发异常进入except分支,后面的过滤代码根本没执行。 - 错误的存在性判断:
any(test) in df写法错误,any(test)本身是布尔值(永远为True),这里应该判断当前字符是否在test数组里。 - 移除错误的字符:
df.remove('i')是移除字符'i',而不是变量i对应的数字字符,而且循环range(0,9)只到8,漏了数字9。 - 宽泛的异常捕获:
except:捕获所有异常,导致代码逻辑出错后直接跳过,你根本看不到哪里出问题。 - 遍历列表时修改原列表:即使前面逻辑对,遍历
df的时候直接修改原列表,会导致遍历漏项。
修正后的代码
直接用列表推导式过滤掉数字字符,简单高效:
import pandas as pd # 读取文件并替换制表符为逗号,保存为csv with open('ASCII.txt') as f: data = f.read().replace('\t', ',') with open('my_file.csv', 'w') as f: f.write(data) # 过滤数字字符 test = {'0','1','2','3','4','5','6','7','8','9'} # 用集合查询更快 filtered_list = [char for char in data if char not in test] print(filtered_list)
或者用字符串的isdigit()方法,不用自定义数组:
filtered_list = [char for char in data if not char.isdigit()]
这样输出就会移除所有数字字符,只保留其他ASCII字符。
内容的提问来源于stack exchange,提问作者Harry Critchfiled
相关产品推荐
相关产品推荐

