You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从Pandas文档列提取计费码并保持索引一致的技术求助

Solution to Extract ICD Codes from Pandas DataFrame

It looks like your current approach has a couple of key issues causing partial results and index mismatches. Let's break down the problems and fix them step by step:

Key Issues in Your Current Code

  1. Slow Membership Checks: Using a list for icd_codes means every word lookup is O(n), which gets very slow with large code lists and can lead to incomplete processing.
  2. Returning List-Wrapped NaN: Returning [np.NaN] (a list) instead of a scalar np.nan can cause unexpected behavior in your DataFrame, especially with index alignment.
  3. Broad Word Matching: Splitting all words with \w+ checks every word in the document, which is inefficient and might pick up non-ICD terms that accidentally match your code list.

Corrected Approach

Here's an optimized solution that fixes these issues and ensures full extraction across all rows:

Step 1: Prep Your Valid Codes

First, convert your ICD code list to a set for near-instant membership checks:

# Convert list to set for fast lookups
icd_codes_set = set(icd_codes)
# Optional: Convert to uppercase to handle case insensitivity (adjust if your codes are case-sensitive)
icd_codes_set = {code.upper() for code in icd_codes}

Step 2: Targeted Regex Pattern

Create a regex that specifically matches ICD code patterns (letter followed by 1-6 alphanumeric characters) to avoid checking irrelevant words:

import re
import numpy as np
import pandas as pd

# Regex to match ICD codes: word boundary, letter, 1-6 alphanumerics, word boundary
icd_pattern = re.compile(r'\b[A-Za-z][A-Za-z0-9]{1,6}\b')

Step 3: Improved Extraction Function

Rewrite the function to first extract only potential ICD candidates, then filter against your valid set:

def extract_icd_codes(text, valid_codes):
    # Extract all patterns that match the ICD code structure
    candidates = icd_pattern.findall(text)
    # Filter candidates that exist in your valid code set (case-insensitive)
    valid_matches = [code.upper() for code in candidates if code.upper() in valid_codes]
    # Remove duplicate codes from the same document
    valid_matches = list(set(valid_matches))
    # Return matches list, or np.nan if no valid codes found
    return valid_matches if valid_matches else np.nan

Step 4: Apply to Entire DataFrame

Apply the function to your Diagnosis column—this will preserve your original index and process all rows:

df['ICD'] = df['Diagnosis'].apply(lambda doc: extract_icd_codes(doc, icd_codes_set))

Why This Works

  • Set Lookups: Converting your code list to a set reduces membership checks from O(n) to O(1), making the function fast enough to handle thousands of documents.
  • Targeted Regex: Only checking words that fit the ICD code structure cuts down on unnecessary processing and reduces false positives.
  • Proper Missing Values: Returning np.nan (a scalar) instead of a list ensures your DataFrame handles missing values correctly, maintaining index alignment.
  • Case Insensitivity: Converting both candidates and valid codes to uppercase ensures you don't miss matches due to case differences (adjust this if your codes are case-sensitive).

Example Output

If your Diagnosis column has a row like:

"Patient presents with symptoms matching ICD codes M545 and I10, no other conditions noted."

And your valid codes include M545 and I10, the ICD column will show ['M545', 'I10'] for that row. Rows with no matches will show NaN.

内容的提问来源于stack exchange,提问作者StephenD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:59:06