Python difflib diff结果提取问题:按规则筛选指定行
Hey there! Let's fix this step by step. The main issue with your current code is that you're checking if a single line starts with both - and ? (which can't happen) instead of looking at the current line and the next one. Let's break down the solution to match your exact requirements.
First, let's recap your rules clearly to make sure we're on the same page:
- Updated list: Lines starting with
-where the very next line starts with?(we'll extract thetextvalue from these) - Deleted list: Lines starting with
+where the next line doesn't start with?(extracttextvalue) - Inserted list: Lines starting with
-where the next line doesn't start with?(extracttextvalue)
Corrected Code
import difflib import re # Assume f1_text and f2_text are your list of lines from the two JSON files diff = difflib.Differ() diffile = [] # First, collect only the relevant diff lines (-/+/?) for line in diff.compare(f1_text, f2_text): # Difflib uses "- ", "+ ", "? " as prefixes (note the space after each symbol) if line.startswith(('- ', '+ ', '? ')): diffile.append(line) # Initialize our result lists updated = [] deleted = [] inserted = [] # Regex to extract the text content from lines like '"text": "abc..."' # This ignores header lines as you requested text_pattern = re.compile(r'"text":\s*"([^"]+)"') # Iterate through each line with its index to check the next line for i in range(len(diffile)): current_line = diffile[i] # Handle lines starting with "- " (removed from first file) if current_line.startswith('- '): # Check if there's a next line and it's a diff marker (?) next_is_question = (i + 1 < len(diffile)) and diffile[i+1].startswith('? ') if next_is_question: # Extract the text value if the line has a "text" field match = text_pattern.search(current_line) if match: updated.append(match.group(1)) else: # No ? line after, add to inserted match = text_pattern.search(current_line) if match: inserted.append(match.group(1)) # Handle lines starting with "+ " (added in second file) elif current_line.startswith('+ '): # Check if there's no next line OR next line isn't a ? marker next_not_question = (i + 1 >= len(diffile)) or not diffile[i+1].startswith('? ') if next_not_question: # Extract the text value if present match = text_pattern.search(current_line) if match: deleted.append(match.group(1))
Why This Works
- Collecting Diff Lines: We filter only lines with
-,+, or?prefixes (difflib's standard format) to avoid irrelevant lines. - Regex for Text Extraction: The regex
r'"text":\s*"([^"]+)"'grabs the value inside the quotes for any line with a"text"key, ignoring header lines as you asked. - Checking Next Line: For each line, we use its index to look ahead to the next line. This lets us correctly apply your rules:
- For
-lines: If next line is?, add toupdated; else add toinserted. - For
+lines: If next line isn't?(or there is no next line), add todeleted.
- For
Example Walkthrough with Your Diff
Let's apply this to your sample diff:
- Line
- "text": "abc xyz efg "has no?line after it → added toinserted. - Line
- "header": true,has a?line after it, but it's a header → ignored. - Line
+ "text": "some text"is the last line (no next line) → added todeleted. - Any
-line with a?line after it that has a"text"field → added toupdated.
This should exactly match what you need as a Python beginner!
内容的提问来源于stack exchange,提问作者Ritz3854
相关产品推荐
相关产品推荐

