如何用Python提取Diff块头中的数字(如74,6、73,7)
用Python提取Diff块头的行号信息
示例Diff内容
@@ -74,6 +73,7 @@ <dependency> <groupId>com.jolbox</groupId> <artifactId>bonecp-test - commons</artifactId> + <classifier>${project.classifier}</classifier> <scope>test</scope>
需求
从Diff的块头(开头的@@ -74,6 +73,7 @@行)中提取74,6和73,7这类数字组合。
实现方案
利用正则匹配Diff块头的固定格式即可,以下是具体代码:
单块Diff提取
import re # 读取Diff内容到字符串(实际场景可从文件读取) diff_content = """@@ -74,6 +73,7 @@ <dependency> <groupId>com.jolbox</groupId> <artifactId>bonecp-test - commons</artifactId> + <classifier>${project.classifier}</classifier> <scope>test</scope>""" # 匹配块头的正则,捕获两组数字组合 pattern = r'@@ -(\d+,\d+) \+(\d+,\d+) @@' match_result = re.search(pattern, diff_content) if match_result: old_file_info = match_result.group(1) # 输出:74,6 new_file_info = match_result.group(2) # 输出:73,7 print(f"旧文件起始行+行数:{old_file_info}") print(f"新文件起始行+行数:{new_file_info}")
多块Diff提取(处理包含多个块的Diff内容)
import re multi_diff_content = """@@ -1,3 +1,2 @@ foo -bar +baz @@ -74,6 +73,7 @@ <dependency> <groupId>com.jolbox</groupId> <artifactId>bonecp-test - commons</artifactId> + <classifier>${project.classifier}</classifier> <scope>test</scope>""" pattern = r'@@ -(\d+,\d+) \+(\d+,\d+) @@' all_results = re.findall(pattern, multi_diff_content) for block_num, (old_info, new_info) in enumerate(all_results, 1): print(f"第{block_num}块 - 旧文件:{old_info},新文件:{new_info}")
说明
正则表达式@@ -(\d+,\d+) \+(\d+,\d+) @@的作用:
@@和固定符号-、+用来定位Diff块头的特征格式(\d+,\d+)精准捕获数字+逗号+数字的组合,对应Diff块头里的行号信息
内容的提问来源于stack exchange,提问作者waled
相关产品推荐
相关产品推荐

