You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas循环遍历中为DataFrame新增文本文件匹配列的实现问题

问题解决方案

你代码的核心问题是dfl['Match'] = s1.strip()属于整列统一赋值,每次循环都会覆盖所有行的Match列值,需要改成仅针对当前遍历到的行赋值。同时原代码存在字符串拼接语法错误、未处理无匹配/多匹配场景的问题。
修改后的完整代码如下:

import re
import pandas as pd

# 提前初始化Match列,默认填充空值
dfl['Match'] = ''

for index, row in dfl.iterrows():
    words = row['Term']
    field_name = row['Type']
    # 注意:原代码的row['chart']和你给出的DataFrame结构的TextID字段不匹配,如实际字段为chart可自行修改
    path = "C:\\Users\\myfolder\\"+str(row['TextID'])+".txt"
    match_res = []
    with open(path,"r", encoding='utf-8') as myfile:
       target_string = myfile.read()
       # 执行正则匹配
       for match in re.finditer(words, str(target_string), re.IGNORECASE):
           print(match.start(), match.end(), match.group())
           # 拼接单条匹配结果
           single_match = f"{match.start()},{match.end()},{field_name}"
           match_res.append(single_match)
    # 多个匹配用分号分隔,赋值给当前行
    dfl.loc[index, 'Match'] = ';'.join(match_res).strip()

关键修改点

  • 提前初始化Match列,避免后续按行赋值时报错
  • 把整列赋值改为按行索引赋值dfl.loc[index, 'Match'],仅修改当前遍历行的值
  • 修正了字符串拼接的语法错误,用f-string简化拼接逻辑
  • 补充了多匹配场景的处理,多个匹配结果用分号分隔,无匹配时留空
  • 补充了文件读取的编码参数,避免读取中文文件时报错
  • 修正了路径拼接的字段名,和你给出的DataFrame的TextID字段对齐,如实际字段名不同可自行调整

内容的提问来源于stack exchange,提问作者RData

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 09:36:02