You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用正则表达式(re)提取以[[Image、[[Media:或[[:file:开头的字符串

Alright, let's figure out how to extract those specific wiki-style media links using Python's regex module. Here's a straightforward solution tailored to your needs:

Solution

First, let's break down what we need to match:

  • Links starting with [[Image (like [[Image:Houstonia...]])
  • Links starting with [[Media: (like [[Media:Example.jpg]])
  • Links starting with [[:file: (like [[:File:Example.jpg]])

All these links end with ]], so we can build a regex pattern that targets these start patterns and captures everything up to the closing ]].

Step-by-Step Implementation

1. Regex Pattern Explanation

We'll use this pattern (with case-insensitivity to handle uppercase/lowercase variants like [[:File:]]):

\[\[(?:Image|Media:|\[:file:).*?]]

Let's break it down:

  • \[\[: Matches the opening double brackets (we escape [ because it's a special regex character)
  • (?:Image|Media:|\[:file:): A non-capturing group that matches any of our three target start sequences:
    • Image: For links starting with [[Image
    • Media:: For links starting with [[Media:
    • \[:file:: For links starting with [[:file:
  • .*?: Non-greedy match for any character (until the first closing ]]—this ensures we don't accidentally capture multiple links if they're close together)
  • ]]: Matches the closing double brackets

2. Python Code Example

Here's how to put this into practice with the re module:

import re

# Your input strings (can be a list or a single concatenated string)
input_links = [
    "[[:File:Example.jpg]]",
    "[[:File:Example.jpg|this example]]",
    "[[Media:Example.jpg]]",
    "[[Georgia (U.S. state)|Georgia]]",
    "[[Arkansas]]",
    "[[Canada]]",
    "[[Virginia]]",
    "[[Image:Houstonia longifolia - Long Leaf Bluet 2.jpg|thumb|left]]"
]

# Combine all strings into one (or process each item individually if needed)
full_text = " ".join(input_links)

# Compile the regex pattern with case-insensitive flag
media_pattern = re.compile(r'\[\[(?:Image|Media:|\[:file:).*?]]', re.IGNORECASE)

# Extract all matching links
matched_links = media_pattern.findall(full_text)

# Print or use the results
for link in matched_links:
    print(link)

3. Output

Running this code will output exactly the links you want:

[[:File:Example.jpg]]
[[:File:Example.jpg|this example]]
[[Media:Example.jpg]]
[[Image:Houstonia longifolia - Long Leaf Bluet 2.jpg|thumb|left]]

内容的提问来源于stack exchange,提问作者Dor Cohen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 21:12:42