从Pandas文档列提取计费码并保持索引一致的技术求助
It looks like your current approach has a couple of key issues causing partial results and index mismatches. Let's break down the problems and fix them step by step:
Key Issues in Your Current Code
- Slow Membership Checks: Using a list for
icd_codesmeans every word lookup is O(n), which gets very slow with large code lists and can lead to incomplete processing. - Returning List-Wrapped NaN: Returning
[np.NaN](a list) instead of a scalarnp.nancan cause unexpected behavior in your DataFrame, especially with index alignment. - Broad Word Matching: Splitting all words with
\w+checks every word in the document, which is inefficient and might pick up non-ICD terms that accidentally match your code list.
Corrected Approach
Here's an optimized solution that fixes these issues and ensures full extraction across all rows:
Step 1: Prep Your Valid Codes
First, convert your ICD code list to a set for near-instant membership checks:
# Convert list to set for fast lookups icd_codes_set = set(icd_codes) # Optional: Convert to uppercase to handle case insensitivity (adjust if your codes are case-sensitive) icd_codes_set = {code.upper() for code in icd_codes}
Step 2: Targeted Regex Pattern
Create a regex that specifically matches ICD code patterns (letter followed by 1-6 alphanumeric characters) to avoid checking irrelevant words:
import re import numpy as np import pandas as pd # Regex to match ICD codes: word boundary, letter, 1-6 alphanumerics, word boundary icd_pattern = re.compile(r'\b[A-Za-z][A-Za-z0-9]{1,6}\b')
Step 3: Improved Extraction Function
Rewrite the function to first extract only potential ICD candidates, then filter against your valid set:
def extract_icd_codes(text, valid_codes): # Extract all patterns that match the ICD code structure candidates = icd_pattern.findall(text) # Filter candidates that exist in your valid code set (case-insensitive) valid_matches = [code.upper() for code in candidates if code.upper() in valid_codes] # Remove duplicate codes from the same document valid_matches = list(set(valid_matches)) # Return matches list, or np.nan if no valid codes found return valid_matches if valid_matches else np.nan
Step 4: Apply to Entire DataFrame
Apply the function to your Diagnosis column—this will preserve your original index and process all rows:
df['ICD'] = df['Diagnosis'].apply(lambda doc: extract_icd_codes(doc, icd_codes_set))
Why This Works
- Set Lookups: Converting your code list to a set reduces membership checks from O(n) to O(1), making the function fast enough to handle thousands of documents.
- Targeted Regex: Only checking words that fit the ICD code structure cuts down on unnecessary processing and reduces false positives.
- Proper Missing Values: Returning
np.nan(a scalar) instead of a list ensures your DataFrame handles missing values correctly, maintaining index alignment. - Case Insensitivity: Converting both candidates and valid codes to uppercase ensures you don't miss matches due to case differences (adjust this if your codes are case-sensitive).
Example Output
If your Diagnosis column has a row like:
"Patient presents with symptoms matching ICD codes M545 and I10, no other conditions noted."
And your valid codes include M545 and I10, the ICD column will show ['M545', 'I10'] for that row. Rows with no matches will show NaN.
内容的提问来源于stack exchange,提问作者StephenD

