如何不使用正则表达式过滤Python列表中movie后带数字的URL
Answer: Filter links without regex
Absolutely! You don’t need regular expressions to achieve this filtering. Simple string splitting and digit-checking methods work perfectly for your use case.
Approach
The key idea is to isolate the segment immediately after movie/ and verify if it consists entirely of digits before the next /:
- For each link, check if it contains the substring
movie/(to skip any unrelated links). - Split the link to extract the part right after
movie/. - Split that extracted part again at the next
/to get the segment we need to validate. - Use Python’s built-in
isdigit()method to check if this segment is all numeric.
Code Implementation
Using List Comprehension (Python 3.8+)
This uses the walrus operator (:=) for concise, readable code:
all_links = [ 'afisha.ru/movie/y2010-2019/vybor-afishi/', 'afisha.ru/movie/257915/', 'afisha.ru/movie/257574/', 'afisha.ru/movie/258600/', 'afisha.ru/movie/257467/', 'afisha.ru/movie/246562/', 'afisha.ru/movie/changed_world/' ] filtered_links = [ link for link in all_links if 'movie/' in link and (id_segment := link.split('movie/')[1].split('/')[0]).isdigit() ] print(filtered_links)
Compatible with Older Python Versions (Pre-3.8)
If you’re using an older Python version, you can rewrite it without the walrus operator:
all_links = [ 'afisha.ru/movie/y2010-2019/vybor-afishi/', 'afisha.ru/movie/257915/', 'afisha.ru/movie/257574/', 'afisha.ru/movie/258600/', 'afisha.ru/movie/257467/', 'afisha.ru/movie/246562/', 'afisha.ru/movie/changed_world/' ] filtered_links = [] for link in all_links: if 'movie/' not in link: continue post_movie_part = link.split('movie/')[1] id_segment = post_movie_part.split('/')[0] if id_segment.isdigit(): filtered_links.append(link) print(filtered_links)
Expected Output
Both code snippets will produce your desired result:
['afisha.ru/movie/257915/', 'afisha.ru/movie/257574/', 'afisha.ru/movie/258600/', 'afisha.ru/movie/257467/', 'afisha.ru/movie/246562/']
Notes
- This method assumes each link contains
movie/exactly once (as in your example). If you have links with multiplemovie/substrings, you may need to adjust the splitting logic to target the correct occurrence. - The
isdigit()method works here because your valid segments are positive integers. If you ever need to handle negative numbers or decimals, you’d need a slight adjustment, but this fits your use case perfectly.
内容的提问来源于stack exchange,提问作者name_of_user
相关产品推荐
相关产品推荐

