如何用Python正则表达式提取换行后的文本?提取指定内容遇阻
Hey there! Let's sort out that regex problem you're facing when trying to extract content after "REQUIRED QUALIFICATIONS:".
The Root Cause
The issue is that by default, the regex wildcard . doesn't match newline characters (like \r or \n). When you used pattern = re.compile(r"REQUIRED QUALIFICATIONS: .*"), the .* only matched up to the immediate newline character (\r) after the heading, since that's the first character it can't handle.
Your earlier match for "JOB RESPONSIBILITIES:" worked because that heading's content was on the same line—no newline to stop the .* from grabbing the text.
Solutions to Try
1. Use the re.DOTALL Flag
This flag makes the . wildcard match all characters, including newlines. Here's how to adjust your code:
import re # Compile the pattern with DOTALL pattern = re.compile(r"REQUIRED QUALIFICATIONS: (.*)", re.DOTALL) matches = pattern.finditer(gh) # Extract the cleaned matched content for match in matches: print(match.group(1).strip()) # Strip extra whitespace/newlines
If you want to avoid grabbing content beyond the "REQUIRED QUALIFICATIONS" section (e.g., if there's another uppercase heading after it), use non-greedy matching with .*? and a lookahead to stop at the next section:
pattern = re.compile(r"REQUIRED QUALIFICATIONS: (.*?)(?=\n[A-Z\s]+:|$)", re.DOTALL)
The (?=\n[A-Z\s]+:|$) part tells the regex to stop when it hits a new line starting with uppercase letters/spaces followed by a colon, or the end of the string.
2. Replace .* with [\s\S]*
Instead of using a flag, you can use a character set that explicitly matches all characters. [\s\S] matches any whitespace character (\s) or non-whitespace character (\S)—which covers every possible character, including newlines:
pattern = re.compile(r"REQUIRED QUALIFICATIONS: ([\s\S]*)")
Again, swap [\s\S]* for [\s\S]*? if you need non-greedy matching to limit the range.
Quick Test Tip
After adjusting your pattern, print out the full matched content to verify—this helps you spot if you're grabbing extra text or stopping too early.
内容的提问来源于stack exchange,提问作者nabskim

