从网页文本中提取电影标题失败,求技术解决方案
Let’s break down how to pull out those movie titles from your messy input string. First, let’s look at the pattern in your text: after the header content ("2018年3月23日-25日周末 标题 周末票房 上映周数"), each movie entry follows this consistent structure:
- Movie Title (can include spaces, colons, and other non-$/digit characters)
- Weekend gross (starts with
$, like$28.1M) - Secondary gross figure (also
$-prefixed, like$17M) - Release week number (a plain digit, like
1)
The key insight here is that movie titles are the only segments that don’t start with a $ or a digit, and they’re always immediately followed by a $-prefixed box office value. We can use regular expressions to capitalize on this pattern.
Example Code (Python)
Here’s a practical implementation that works perfectly for your input:
import re raw_text = "2018年3月23日-25日周末 标题 周末票房 上映周数 Pacific Rim: Uprising $28.1M $17M 1 Spider Man: Home Coming $37.8M $12M 3" # Regex to target titles: sequences that don't start with $/digit, end right before a $ value movie_titles = re.findall(r'(?<=\s)([^\$\d].*?)(?=\s\$)', raw_text) print(movie_titles) # Output: ['Pacific Rim: Uprising', 'Spider Man: Home Coming']
How the Regex Works
Let’s unpack the pattern (?<=\s)([^\$\d].*?)(?=\s\$):
(?<=\s): Positive lookbehind to ensure the title is preceded by a space (prevents matching header text)[^\$\d].*?: Matches any sequence that doesn’t start with$or a digit, using non-greedy matching (*?) to stop as soon as the next box office value begins(?=\s\$): Positive lookahead to confirm the title is followed by a space and$(the start of the box office number)
Edge Case Adjustment
If you ever encounter titles that include digits (e.g., Fast & Furious 9), you can tweak the regex to allow digits inside the title while still stopping at the $ value:
# Allows digits in titles, still stops before the $ box office value movie_titles = re.findall(r'(?<=\s)([^\$].*?)(?=\s\$)', raw_text)
内容的提问来源于stack exchange,提问作者Ahmad Egbaria

