正则If-then、前瞻断言与括号应用:提取Python指定import语句需求
Matching Both Parenthesized and Line-Continued
from ... import ... Statements with Regex + Grep Got it, let's solve this problem of matching both styles of Python import statements using a single regex with grep. First, let's recap the two formats we need to cover:
Parenthesized imports (with optional newlines inside)
Example:from mymodule import ( func_a, ClassB, CONSTANT_C )Line-continued imports (using
\, no internal newlines without the line-continuation character)
Example:from another_module import func_x, ClassY, \ CONSTANT_Z
The Combined Regular Expression
Here's a regex that handles both cases, optimized for use with grep (using Perl-compatible mode):
^from\s+([\w.]+)\s+import\s+(?:\((?:\s|\n)*([\w\s,]+)(?:\s|\n)*\)|([\w,]+(?:\\\n[\w, ]+)*))$
Breakdown of the Regex
Let's break this down piece by piece so you can tweak it if needed:
^from\s+: Anchors to the start of a line, matches the literalfromfollowed by one or more spaces.([\w.]+): Captures the module name (supports dots for nested packages likemy.package.module).\s+import\s+: Matches theimportkeyword surrounded by spaces.(?:...): A non-capturing group that wraps our two import style options, separated by|(OR):- Parenthesized branch:
\((?:\s|\n)*([\w\s,]+)(?:\s|\n)*\)\(: Matches the opening parenthesis.(?:\s|\n)*: Allows any number of spaces or newlines inside the parentheses (before, between, and after imports).([\w\s,]+): Captures the list of imports (words, spaces, commas).\): Matches the closing parenthesis.
- Line-continued branch:
([\w,]+(?:\\\n[\w, ]+)*)[\w,]+: Matches the first part of the import list (words and commas).(?:\\\n[\w, ]+)*: Allows any number of line-continued segments (matches the\followed by a newline, then more imports with spaces/commas).
- Parenthesized branch:
$: Anchors to the end of the line (works with grep's multiline mode when using-P).
Using with Grep
To run this on your Python file and extract only the matching import statements, use this command:
grep -Pzo '^from\s+([\w.]+)\s+import\s+(?:\((?:\s|\n)*[\w\s,]+\s*\)|[\w,]+(?:\\\n[\w, ]+)*)$' your_script.py
Let's explain the flags:
-P: Enables Perl-compatible regular expressions (required for non-capturing groups and proper newline handling).-z: Treats the input as a single NUL-terminated string, which lets the regex match across newlines (critical for parenthesized imports with internal line breaks).-o: Only outputs the parts of the file that match the regex, not the entire lines.
Quick Adjustment Tips
- If you don't need to capture the module name specifically, you can replace
([\w.]+)with\S+to simplify. - For stricter matching (e.g., disallowing extra spaces around commas), tweak the
[\w\s,]+parts to something like[\w]+(?:,\s*[\w]+)*.
内容的提问来源于stack exchange,提问作者Brad Solomon
相关产品推荐
相关产品推荐

