Python:从文本文件提取话题标签及@标签的格式问题求助
Got it, this is a common issue when dealing with unstructured text! The split() method falls short here because it relies on explicit separators, but regex is perfect for this scenario since it can pattern-match each tag directly—even when they’re stuck together like #socality#thisismycommunity. Here’s how to solve it:
Step 1: Use Regular Expressions to Extract Tags
Instead of splitting, we’ll use regex to actively find every instance of #hashtags and @mentions. The key is to define patterns that stop capturing as soon as they hit another tag start or whitespace.
Here’s a Python implementation:
import re from collections import Counter # Read your text file with open('your_text_file.txt', 'r', encoding='utf-8') as file: text_content = file.read() # Extract all hashtags (handles concatenated ones) hashtags = re.findall(r'#[^\s#]+', text_content) # Extract all mentions mentions = re.findall(r'@[^\s@]+', text_content)
What the regex patterns mean:
r'#[^\s#]+': Looks for a#followed by one or more characters that are not whitespace (\s) or another#. This ensures it splits concatenated hashtags like#a#b#cinto['#a', '#b', '#c'].r'@[^\s@]+': Works the same way for mentions—captures@followed by any characters until the next@or space.
Step 2: Sort the Tags
Now that you have your list of tags, you can sort them in two common ways:
Option A: Sort by Frequency (Most Common First)
If you want to prioritize tags that appear most often:
# Count hashtag occurrences and sort descending hashtag_counts = Counter(hashtags) sorted_hashtags = sorted(hashtag_counts.items(), key=lambda x: x[1], reverse=True) # Do the same for mentions mention_counts = Counter(mentions) sorted_mentions = sorted(mention_counts.items(), key=lambda x: x[1], reverse=True) # Example output: [('#mountaincultureelevated', 5), ('#modernoutdoors', 3)]
Option B: Sort Alphabetically (Descending Order)
If you just want unique tags sorted from Z to A:
# Unique hashtags sorted descending unique_sorted_hashtags = sorted(list(set(hashtags)), reverse=True) # Unique mentions sorted descending unique_sorted_mentions = sorted(list(set(mentions)), reverse=True)
Edge Cases Handled
- Concatenated tags (like
#x#y#z) are split correctly - Double hashes/mentions (like
##helloor@@user) are captured as valid tags (#hello,@user) - Empty tags (like
###) are ignored since the regex requires at least one character after#/@
内容的提问来源于stack exchange,提问作者Rui Torres

