使用pandas read_table时如何跳过行数不定的C风格注释行
实现逻辑
先逐行扫描文件定位注释结束位置,再将对应行号传入skiprows参数即可,具体代码如下:
1. 编写注释行统计函数
def count_comment_lines(file_path, encoding='utf-8'): in_comment = False end_line = 0 with open(file_path, 'r', encoding=encoding) as f: for idx, line in enumerate(f): stripped = line.strip() if stripped.startswith('/*'): in_comment = True if in_comment: end_line = idx + 1 if stripped.endswith('*/'): break return end_line
函数返回的是注释块占用的总行数,刚好可以直接传给skiprows参数,因为skiprows传入整数时就代表跳过前N行。
2. 读取.tab文件
import pandas as pd tab_path = "你的目标文件路径.tab" skip_n = count_comment_lines(tab_path) df = pd.read_table(tab_path, skiprows=skip_n)
特殊场景适配
如果文件存在多段分散的/* */注释块,修改统计函数收集所有注释行的行号即可:
def get_all_comment_lines(file_path, encoding='utf-8'): in_comment = False skip_lines = [] with open(file_path, 'r', encoding=encoding) as f: for idx, line in enumerate(f): stripped = line.strip() if stripped.startswith('/*'): in_comment = True if in_comment: skip_lines.append(idx) if stripped.endswith('*/'): in_comment = False return skip_lines
使用时直接把返回的行号列表传给skiprows参数即可。
内容的提问来源于stack exchange,提问作者Yongwu Xiu
相关产品推荐
相关产品推荐

