循环将NER引擎输出插入追加到pandas DataFrame时数据被覆盖如何解决
你的代码存在两处问题,是导致数据覆盖、结果不符合预期的核心原因:
- 语法错误:给match列赋值时遗漏了
.loc索引器,写成了非法的dfner.[dfindex,'match'],导致赋值逻辑失效 - 逻辑错误:你要求match列存储识别到的实体文本,实际代码写的是
ent.label_(实体类型字段),需要替换为ent.text
修正后的完整代码如下:
import pandas as pd # 提前声明列名更规范,避免赋值时的类型推断问题 dfner = pd.DataFrame(columns=['id', 'match']) dfindex = 0 for index, row in dfl.iterrows(): id = row['id'] path = "C:\\Users\\myfolder\\"+str(row['chart'])+".txt" # 建议指定编码,避免读取文件时出现乱码 with open(path,"r", encoding='utf-8') as myfile: target_string = myfile.read() doc = nlp(target_string) for ent in doc.ents: dfner.loc[dfindex,'id'] = str(id) dfner.loc[dfindex,'match'] = ent.text dfindex += 1
运行上述代码即可得到你预期的dfner结构:
| id | match |
|---|---|
| 1 | john |
| 1 | doe |
| 2 | chris |
| 2 | clark |
内容的提问来源于stack exchange,提问作者RData
相关产品推荐
相关产品推荐

