You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从网页文本中提取电影标题失败,求技术解决方案

Extracting Movie Titles from Unstructured Box Office Text

Let’s break down how to pull out those movie titles from your messy input string. First, let’s look at the pattern in your text: after the header content ("2018年3月23日-25日周末 标题 周末票房 上映周数"), each movie entry follows this consistent structure:

  • Movie Title (can include spaces, colons, and other non-$/digit characters)
  • Weekend gross (starts with $, like $28.1M)
  • Secondary gross figure (also $-prefixed, like $17M)
  • Release week number (a plain digit, like 1)

The key insight here is that movie titles are the only segments that don’t start with a $ or a digit, and they’re always immediately followed by a $-prefixed box office value. We can use regular expressions to capitalize on this pattern.

Example Code (Python)

Here’s a practical implementation that works perfectly for your input:

import re

raw_text = "2018年3月23日-25日周末 标题 周末票房 上映周数 Pacific Rim: Uprising $28.1M $17M 1 Spider Man: Home Coming $37.8M $12M 3"

# Regex to target titles: sequences that don't start with $/digit, end right before a $ value
movie_titles = re.findall(r'(?<=\s)([^\$\d].*?)(?=\s\$)', raw_text)

print(movie_titles)
# Output: ['Pacific Rim: Uprising', 'Spider Man: Home Coming']

How the Regex Works

Let’s unpack the pattern (?<=\s)([^\$\d].*?)(?=\s\$):

  • (?<=\s): Positive lookbehind to ensure the title is preceded by a space (prevents matching header text)
  • [^\$\d].*?: Matches any sequence that doesn’t start with $ or a digit, using non-greedy matching (*?) to stop as soon as the next box office value begins
  • (?=\s\$): Positive lookahead to confirm the title is followed by a space and $ (the start of the box office number)

Edge Case Adjustment

If you ever encounter titles that include digits (e.g., Fast & Furious 9), you can tweak the regex to allow digits inside the title while still stopping at the $ value:

# Allows digits in titles, still stops before the $ box office value
movie_titles = re.findall(r'(?<=\s)([^\$].*?)(?=\s\$)', raw_text)

内容的提问来源于stack exchange,提问作者Ahmad Egbaria

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:28:50