匹配文本与注释并创建Pandas DataFrame的技术需求
解决CSV文本行与注释行匹配并生成Pandas DataFrame的问题
需求说明
需要从CSV文件中识别包含关键词$$n[]:或$$n[<任意字符>]:的注释行,将这些注释行与它上方最近的、不含该关键词的文本行进行匹配,最终生成一个Pandas DataFrame展示匹配结果——没有对应注释的文本行用None填充。
示例CSV内容
This is text line 1 $$n[]: This is note for text line 1 This is text line 2 This is text line 3 $$n[custom]: This is note for text line 3 This is text line 4
修正后的实现代码
import pandas as pd import re # 定义匹配注释行的正则表达式,兼容带自定义标识和不带标识的两种格式 note_pattern = re.compile(r'^\$\$n(\[.*?\])?:') def match_text_note(csv_path): text_note_pairs = [] current_text = None with open(csv_path, 'r', encoding='utf-8') as f: for line in f: line = line.strip() if not line: continue # 跳过空行 # 判断当前行是否为注释行 if note_pattern.match(line): if current_text is not None: # 提取注释内容(移除开头的关键词标识) note_content = re.sub(note_pattern, '', line).strip() text_note_pairs.append({'text': current_text, 'note': note_content}) current_text = None # 匹配完成后重置当前文本行,避免重复配对 else: # 如果存在未匹配的文本行,先将其以无注释的形式存入列表 if current_text is not None: text_note_pairs.append({'text': current_text, 'note': None}) # 更新当前跟踪的文本行 current_text = line # 处理文件末尾剩余的未匹配文本行 if current_text is not None: text_note_pairs.append({'text': current_text, 'note': None}) # 转换为Pandas DataFrame df = pd.DataFrame(text_note_pairs) return df # 示例调用(替换为你的CSV文件路径) result_df = match_text_note('your_target.csv') print(result_df)
代码逻辑说明
- 使用正则表达式精准匹配两种格式的注释行开头
- 遍历CSV时维护
current_text变量,跟踪当前未配对的文本行 - 遇到注释行时,立即与最近的文本行配对;遇到新文本行时,先处理之前未配对的文本行(标记为无注释),再更新跟踪的文本行
- 最后处理文件末尾可能剩余的未配对文本行
预期输出
text note 0 This is text line 1 This is note for text line 1 1 This is text line 2 None 2 This is text line 3 This is note for text line 3 3 This is text line 4 None
内容的提问来源于stack exchange,提问作者edge selcuk
相关产品推荐
相关产品推荐

