You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写Acestream正则表达式?从文本提取Acestream URL技术问询

Extracting Acestream URLs with Regular Expressions

Great question! Acestream URLs follow a consistent structure centered around a 40-character hexadecimal SHA-1 ID, so crafting a regex to pull them out of text is straightforward once you know the pattern.

Key Background

A standard Acestream URL looks like this:
acestream://a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0
Occasionally, you might see a suffix like /playlist.m3u8 added for direct playlist access.

Regex Patterns for Different Use Cases

1. Match Full Standard Acestream URLs (Protocol + 40-char ID)

Use this to capture the complete, basic URL:

acestream://[0-9a-fA-F]{40}
  • Breakdown:
    • acestream://: Exact match for the Acestream protocol prefix
    • [0-9a-fA-F]{40}: Matches exactly 40 hexadecimal characters (0-9, A-F, a-f) — this is the unique stream ID

2. Match URLs with Optional Playlist Suffix

If you need to account for URLs ending in /playlist.m3u8:

acestream://[0-9a-fA-F]{40}(/playlist\.m3u8)?
  • The (/playlist\.m3u8)? part makes the playlist suffix optional (the ? means "zero or one occurrence")

3. Extract Only the Stream ID (Skip the Protocol Prefix)

If you just want the 40-character ID without acestream://, use a positive lookbehind:

(?<=acestream://)[0-9a-fA-F]{40}
  • (?<=acestream://): Ensures the ID is preceded by the protocol prefix, but doesn't include the prefix in the match result

Example Implementation (Python)

Here's how you'd use these regexes in practice to extract URLs/IDs from a sample text:

import re

sample_text = """
Check out this sports stream: acestream://a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0
Another one with playlist: acestream://f0e1d2c3b4a5f6e7d8c9b0a1f2e3d4c5b6a7f8e9/playlist.m3u8
"""

# Capture all full URLs (including optional playlist suffix)
full_urls = re.findall(r'acestream://[0-9a-fA-F]{40}(/playlist\.m3u8)?', sample_text, re.IGNORECASE)
# Extract only the stream IDs
stream_ids = re.findall(r'(?<=acestream://)[0-9a-fA-F]{40}', sample_text, re.IGNORECASE)

print("Full Acestream URLs:", full_urls)
print("Extracted Stream IDs:", stream_ids)

Pro Tips

  • Add re.IGNORECASE (or the equivalent flag in your regex engine) to handle both uppercase and lowercase hex characters
  • If your text has URLs followed by punctuation (like commas or periods), add a word boundary \b or a lookahead (?=\s|$|,|\.) to avoid capturing extra characters:
    acestream://[0-9a-fA-F]{40}(/playlist\.m3u8)?(?=\s|$|,|\.)
    

内容的提问来源于stack exchange,提问作者BrunoRamalho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:04:06