DataFrame文本分句迭代提取结节尺寸值的问题求助
问题分析与解决方案
我来帮你拆解下问题核心,然后给出可直接运行的修复方案:
你的代码主要卡在循环赋值逻辑错误和函数返回值类型不匹配这两个点上:
- 循环里直接执行
rf["Nod_size"] = result_calcified_nod会覆盖所有行的结果,而不是对应到当前处理的那一行文本,最后只剩最后一次循环的结果(如果最后一次没匹配到就会是空值)。 GetNodule函数返回的是pandas Series对象,不是单个字符串,直接赋值给DataFrame列会导致类型不兼容,最终显示空值。- 原函数没有处理数值提取的边界情况,也没有把结果转为易使用的字符串格式。
修复后的完整代码
1. 优化结节尺寸提取函数
调整函数逻辑,确保返回字符串格式的结果(或None表示无匹配):
def GetNodule(sentence): sentence = re.sub('-', ' ', sentence) token_words = nltk.word_tokenize(sentence) df = pd.DataFrame(token_words, columns=["token"]) # 检查句子是否同时包含结节关键词和长度单位关键词 has_nodule = df["token"].str.lower().isin(nodule_keywords).any() has_length = df["token"].str.lower().isin(nodule_length_keyword).any() if has_nodule and has_length: # 找到所有长度单位的位置,取前一个token作为尺寸数值 length_positions = df[df["token"].str.lower().isin(nodule_length_keyword)].index # 取第一个匹配的尺寸,如需所有匹配可改为合并多个结果 nodule_size = df.loc[length_positions[0] - 1, "token"] return nodule_size else: return None
2. 用pandas原生方法处理每行文本
放弃手动循环,改用apply逐行处理,确保结果对应到正确的行:
import pandas as pd import numpy as np import nltk import re from nltk.tokenize import sent_tokenize, word_tokenize # 初始化你的DataFrame rf = pd.DataFrame([ {"Text": "CHEST CA lung. -Increased sizes of nodules in RLL. There is further increased size and solid component of part-solid nodule associated with internal bubbly lucency and pleural tagging at apicoposterior segment of the LUL (SE 3; IM 38-50), now measuring about 2.9x1.7 cm in greatest transaxial dimension (previously size 2.5x1.3 cm in 2015).", "Stage": "T2aN2M0"}, {"Text": "CHEST CA lung. Post LL lobectomy. As compared to study obtained on 30/10/2018, -Top normal heart size. -Increased sizes of nodules in RLL.", "Stage": "T2aN2M0"} ]) nodule_keywords = ["nodules","nodule"] nodule_length_keyword = ["cm","mm", "centimeters", "milimeters"] # 定义单行文本处理函数:遍历分句,返回第一个匹配的结节尺寸 def process_single_text(text): sentences = sent_tokenize(str(text)) for sent in sentences: size = GetNodule(sent) if size is not None: return size return None # 所有句子都无匹配时返回None # 应用到DataFrame的每行 rf["Nod_size"] = rf["Text"].apply(process_single_text) # 查看结果 print(rf)
运行结果验证
你会得到预期的输出:
| Text | Stage | Nod_size |
|---|---|---|
| CHEST CA lung. -Increased sizes of nodules in RLL. There is furth... | T2aN2M0 | 2.9x1.7 |
| CHEST CA lung. Post LL lobectomy. As compared to study obtained ... | T2aN2M0 | NaN |
如果想提取文本中所有的结节尺寸(比如同时获取2.9x1.7和2.5x1.3),可以修改GetNodule函数,收集所有匹配的数值后用逗号分隔返回:
def GetNodule(sentence): sentence = re.sub('-', ' ', sentence) token_words = nltk.word_tokenize(sentence) df = pd.DataFrame(token_words, columns=["token"]) has_nodule = df["token"].str.lower().isin(nodule_keywords).any() has_length = df["token"].str.lower().isin(nodule_length_keyword).any() if has_nodule and has_length: length_positions = df[df["token"].str.lower().isin(nodule_length_keyword)].index nodule_sizes = [df.loc[pos - 1, "token"] for pos in length_positions] return ", ".join(nodule_sizes) else: return None
修改后第一行的Nod_size会变成2.9x1.7, 2.5x1.3,更完整地保留文本中的尺寸信息。
内容的提问来源于stack exchange,提问作者khushbu
相关产品推荐
相关产品推荐

