You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:从文本文件提取话题标签及@标签的格式问题求助

Got it, this is a common issue when dealing with unstructured text! The split() method falls short here because it relies on explicit separators, but regex is perfect for this scenario since it can pattern-match each tag directly—even when they’re stuck together like #socality#thisismycommunity. Here’s how to solve it:

Step 1: Use Regular Expressions to Extract Tags

Instead of splitting, we’ll use regex to actively find every instance of #hashtags and @mentions. The key is to define patterns that stop capturing as soon as they hit another tag start or whitespace.

Here’s a Python implementation:

import re
from collections import Counter

# Read your text file
with open('your_text_file.txt', 'r', encoding='utf-8') as file:
    text_content = file.read()

# Extract all hashtags (handles concatenated ones)
hashtags = re.findall(r'#[^\s#]+', text_content)

# Extract all mentions
mentions = re.findall(r'@[^\s@]+', text_content)

What the regex patterns mean:

  • r'#[^\s#]+': Looks for a # followed by one or more characters that are not whitespace (\s) or another #. This ensures it splits concatenated hashtags like #a#b#c into ['#a', '#b', '#c'].
  • r'@[^\s@]+': Works the same way for mentions—captures @ followed by any characters until the next @ or space.

Step 2: Sort the Tags

Now that you have your list of tags, you can sort them in two common ways:

Option A: Sort by Frequency (Most Common First)

If you want to prioritize tags that appear most often:

# Count hashtag occurrences and sort descending
hashtag_counts = Counter(hashtags)
sorted_hashtags = sorted(hashtag_counts.items(), key=lambda x: x[1], reverse=True)

# Do the same for mentions
mention_counts = Counter(mentions)
sorted_mentions = sorted(mention_counts.items(), key=lambda x: x[1], reverse=True)

# Example output: [('#mountaincultureelevated', 5), ('#modernoutdoors', 3)]

Option B: Sort Alphabetically (Descending Order)

If you just want unique tags sorted from Z to A:

# Unique hashtags sorted descending
unique_sorted_hashtags = sorted(list(set(hashtags)), reverse=True)

# Unique mentions sorted descending
unique_sorted_mentions = sorted(list(set(mentions)), reverse=True)

Edge Cases Handled

  • Concatenated tags (like #x#y#z) are split correctly
  • Double hashes/mentions (like ##hello or @@user) are captured as valid tags (#hello, @user)
  • Empty tags (like ###) are ignored since the regex requires at least one character after #/@

内容的提问来源于stack exchange,提问作者Rui Torres

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:58:45