You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python difflib diff结果提取问题:按规则筛选指定行

Hey there! Let's fix this step by step. The main issue with your current code is that you're checking if a single line starts with both - and ? (which can't happen) instead of looking at the current line and the next one. Let's break down the solution to match your exact requirements.

First, let's recap your rules clearly to make sure we're on the same page:

  1. Updated list: Lines starting with - where the very next line starts with ? (we'll extract the text value from these)
  2. Deleted list: Lines starting with + where the next line doesn't start with ? (extract text value)
  3. Inserted list: Lines starting with - where the next line doesn't start with ? (extract text value)

Corrected Code

import difflib
import re

# Assume f1_text and f2_text are your list of lines from the two JSON files
diff = difflib.Differ()
diffile = []

# First, collect only the relevant diff lines (-/+/?)
for line in diff.compare(f1_text, f2_text):
    # Difflib uses "- ", "+ ", "? " as prefixes (note the space after each symbol)
    if line.startswith(('- ', '+ ', '? ')):
        diffile.append(line)

# Initialize our result lists
updated = []
deleted = []
inserted = []

# Regex to extract the text content from lines like '"text": "abc..."'
# This ignores header lines as you requested
text_pattern = re.compile(r'"text":\s*"([^"]+)"')

# Iterate through each line with its index to check the next line
for i in range(len(diffile)):
    current_line = diffile[i]
    
    # Handle lines starting with "- " (removed from first file)
    if current_line.startswith('- '):
        # Check if there's a next line and it's a diff marker (?)
        next_is_question = (i + 1 < len(diffile)) and diffile[i+1].startswith('? ')
        
        if next_is_question:
            # Extract the text value if the line has a "text" field
            match = text_pattern.search(current_line)
            if match:
                updated.append(match.group(1))
        else:
            # No ? line after, add to inserted
            match = text_pattern.search(current_line)
            if match:
                inserted.append(match.group(1))
    
    # Handle lines starting with "+ " (added in second file)
    elif current_line.startswith('+ '):
        # Check if there's no next line OR next line isn't a ? marker
        next_not_question = (i + 1 >= len(diffile)) or not diffile[i+1].startswith('? ')
        
        if next_not_question:
            # Extract the text value if present
            match = text_pattern.search(current_line)
            if match:
                deleted.append(match.group(1))

Why This Works

  1. Collecting Diff Lines: We filter only lines with - , + , or ? prefixes (difflib's standard format) to avoid irrelevant lines.
  2. Regex for Text Extraction: The regex r'"text":\s*"([^"]+)"' grabs the value inside the quotes for any line with a "text" key, ignoring header lines as you asked.
  3. Checking Next Line: For each line, we use its index to look ahead to the next line. This lets us correctly apply your rules:
    • For - lines: If next line is ?, add to updated; else add to inserted.
    • For + lines: If next line isn't ? (or there is no next line), add to deleted.

Example Walkthrough with Your Diff

Let's apply this to your sample diff:

  • Line - "text": "abc xyz efg " has no ? line after it → added to inserted.
  • Line - "header": true, has a ? line after it, but it's a header → ignored.
  • Line + "text": "some text" is the last line (no next line) → added to deleted.
  • Any - line with a ? line after it that has a "text" field → added to updated.

This should exactly match what you need as a Python beginner!

内容的提问来源于stack exchange,提问作者Ritz3854

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 13:22:51