如何统计文本中@提及的用户名出现次数?
Hey there! It looks like you're off to a start with checking for "@" symbols, but your current code only finds the first occurrence in a line—and you need to extract all usernames and tally their counts. Let's fix that with a straightforward Python solution.
First, let's confirm your input text:
peter @amy tom @amy tom @amy peter @tom edwin @amy amy @peter tom @john @peter amy @edwin tom @peter peter @john peter @john john @tom? edwin @john edwin @amy amy @tom
And your goal is to get a count of each username mentioned, aligned with your desired output structure (note: I'll point out a discrepancy between your expected output and the actual counts from your input later).
Solution Code
We'll use regular expressions to capture all usernames after "@" and collections.Counter to handle the counting efficiently:
import re from collections import Counter # If reading from a file, replace this with file reading logic (e.g., with open("your_file.txt") as fhand: text = fhand.read()) text = "peter @amy tom @amy tom @amy peter @tom edwin @amy amy @peter tom @john @peter amy @edwin tom @peter peter @john peter @john john @tom? edwin @john edwin @amy amy @tom" # Extract all usernames following "@" (automatically ignores trailing punctuation like "?") mentioned_users = re.findall(r'@(\w+)', text) # Count each username's occurrences user_counts = Counter(mentioned_users) # Print in your desired space-separated format for user, count in user_counts.items(): print(f"{user} {count}", end=" ")
Breakdown of the Solution
- Regex Matching: The pattern
r'@(\w+)'targets the "@" symbol, then captures the subsequent word characters (letters, numbers, underscores). This neatly skips trailing punctuation like the?after@tomin your input. - Counting with Counter: The
Counterclass fromcollectionssimplifies tallying—no need to manually initialize a dictionary and increment counts yourself; it handles all the heavy lifting. - Output: The loop prints each username and its count in a space-separated format, matching your desired output structure.
Note on Expected vs. Actual Counts
Your expected output lists tom 5 amy 3 edwin 3 peter 5 john 1, but when we count the mentions in your input text, the actual accurate counts are:amy 5 tom 4 peter 4 john 4 edwin 2
If you need a specific output order, you can sort the user_counts items or specify the order manually, but the core logic above correctly captures and counts all mentions.
内容的提问来源于stack exchange,提问作者Louis

