如何用正则表达式筛选含两处数字的企业地址网页爬取内容?
Hey Bob, let's tweak your code to filter texts that have at least two separate groups of numbers. The original regex for lowercase letters ([a-z]*) was actually matching empty strings, so it wasn't properly checking for the presence of lowercase characters—let's fix that too if you still need the uppercase/lowercase check.
Modified Code (With Case Validation)
Here's the updated version that targets texts with at least two digit sequences, plus validates the presence of both uppercase and lowercase letters (fixed the lowercase check from * to + to ensure there's at least one actual lowercase character):
x=['174 WEST 4TH ST, NYC','All contents © Copyright 2018 Propela'] import re def is_location(text): """Checks if text has at least two digit sequences, plus both lowercase and uppercase letters""" # Find all consecutive digit groups (e.g., 174 and 4 in your example) digit_groups = re.findall(r'\d+', text) has_multiple_digits = len(digit_groups) >= 2 # Ensure at least one lowercase letter (fixed from * to + to avoid matching empty strings) has_lower = bool(re.search(r'[a-z]+', text)) # Ensure at least one uppercase letter has_upper = bool(re.search(r'[A-Z]+', text)) return has_multiple_digits and has_lower and has_upper print(list(filter(is_location, x)))
If You Don't Need the Lowercase Requirement
Wait, I noticed your example "174 WEST 4TH ST, NYC" doesn't have lowercase letters. If you only care about having at least two digit sequences and uppercase letters (no lowercase required), adjust the function like this:
def is_location(text): """Checks if text has at least two digit sequences and uppercase letters""" digit_groups = re.findall(r'\d+', text) has_multiple_digits = len(digit_groups) >= 2 has_upper = bool(re.search(r'[A-Z]+', text)) return has_multiple_digits and has_upper
How It Works
re.findall(r'\d+', text)grabs all blocks of consecutive digits in the text. We just need to confirm there are 2 or more of these blocks.r'[a-z]+'ensures there's at least one lowercase letter (unlike*which would match empty strings and always return true).r'[A-Z]+'checks for at least one uppercase letter to filter out non-address text like copyright lines.
The second version will correctly return ['174 WEST 4TH ST, NYC'] as you wanted, while excluding the copyright string that only has one digit group.
内容的提问来源于stack exchange,提问作者Bob

