如何用for循环处理DataFrame每行,运行BERT NER UMLS手术实体识别脚本
修改后可处理多行CSV的完整脚本
%%capture pip install https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.4.0/en_core_sci_scibert-0.4.0.tar.gz import pandas as pd import spacy from scispacy.umls_linking import UmlsEntityLinker %%capture # 模型仅初始化一次,不要放在循环内浪费算力 nlp = spacy.load("en_core_sci_scibert") nlp.add_pipe("scispacy_linker", config={"resolve_abbreviations": True, "linker_name": "umls"}) linker = nlp.get_pipe("scispacy_linker") # 单条医疗记录手术实体判断函数 def has_surgery_entity(record): # 空记录直接返回无手术实体 if pd.isna(record) or not str(record).strip(): return 'No' doc = nlp(str(record)) for entity in doc.ents: # 跳过未匹配到UMLS的实体,避免报错 if not entity._.umls_ents: continue # 取置信度最高的UMLS匹配项 highest_cui = entity._.umls_ents[0][0] umls_data = linker.umls.cui_to_entity[highest_cui] # 匹配手术类型T061则返回存在 if umls_data.types and umls_data.types[0] == 'T061': return 'Yes' return 'No' # -------------------------- # 替换为你本地的实际文件路径和列名 # 此处默认医疗记录在CSV中的列名为medical_record,可根据实际情况修改 df = pd.read_csv("你的输入文件路径.csv") # 新增标记列,保留所有原始数据 df['has_surgery'] = df['medical_record'].apply(has_surgery_entity) # 输出结果到新CSV,避免覆盖原始文件 df.to_csv("输出结果路径.csv", index=False, encoding='utf-8-sig') # 需预览结果时取消注释下行 # print(df)
改动说明
- 把原单条记录的处理逻辑封装为独立函数,比手动写for循环更简洁、可复用性更高
- 新增空值、无UMLS匹配实体的异常处理,避免批量运行时意外中断
- 模型初始化逻辑保留在循环外,大幅提升多记录处理速度
- 输出自动保留所有原始CSV列,仅新增手术实体标记列,符合需求
注意事项
- 代码中
你的输入文件路径.csv、你的输出文件路径.csv需要替换为你本地的实际文件路径 - 如果你的CSV中存储医疗记录的列名不是
medical_record,需要修改为对应的列名 - 如需保留所有识别到的手术实体详情,可在判断函数中追加对应存储逻辑
内容的提问来源于stack exchange,提问作者Konar
相关产品推荐
相关产品推荐

