如何用Python从TXT文件提取TikTok数据分析工具中的话题标签?
First, let's fix a couple of small issues in your existing code that might be causing unnecessary headaches:
- You’re using
datetimebut haven’t imported it at the top of your script. - The placeholder
.lines (like underos.mkdir("./data")) are invalid syntax and need to be removed.
Now, onto your main problem: extracting hashtags from description strings. TikTok hashtags follow a consistent pattern—they start with # and are followed by letters, numbers, or underscores. A regular expression is the perfect tool to match this pattern reliably.
Step 1: Create a Reusable Hashtag Extraction Function
Add this simple function to pull hashtags from any text string using re.findall():
import re def extract_hashtags(text): # Matches any sequence starting with # followed by word characters (letters, numbers, underscores) return re.findall(r'#\w+', text)
Step 2: Integrate the Function into Your Workflow
You have two practical options here: extract hashtags while processing the CSV (before writing to TXT), or extract them directly from an existing TXT file. Let’s cover both:
Option 1: Extract Hashtags During CSV Processing (and Save to TXT/CSV)
Modify your code to extract hashtags as you process each post, then save them alongside your other data:
from tiktok_bot import TikTokBot import csv import os import sys import re import datetime # Don't forget this import! def getData(): return datetime.datetime.now().strftime("%Y-%m-%d") def buildFileName(type): return f"./data/{getData()}{type}.csv" def buildText(type): return f"./data/{getData()}{type}.txt" def extract_hashtags(text): return re.findall(r'#\w+', text) # Create data folder if it doesn't exist try: os.mkdir("./data") except OSError as e: print("Directory exists") # Write CSV with post data + extracted hashtags with open(buildFileName("_shares"), mode='w', newline='') as csv_file: fieldnames = ['User ID', 'URL', 'Description', 'Comments', 'Likes', 'Hashtags'] writer = csv.DictWriter(csv_file, fieldnames=fieldnames) writer.writeheader() for post in most_shared_posts: hashtags = extract_hashtags(post.desc) writer.writerow({ 'User ID': str(post.author_user_id), 'URL': str(post.share_url), 'Description': str(post.desc), 'Comments': post.statistics.comment_count, 'Likes': post.statistics.digg_count, 'Hashtags': ', '.join(hashtags) # Save hashtags as a comma-separated string }) # Write descriptions and their hashtags to TXT with open(buildText("_shares"), "w") as txt_file: with open(buildFileName("_shares"), mode='r') as csv_file: csv_reader = csv.DictReader(csv_file, delimiter=',') next(csv_reader) # Skip the header row for line in csv_reader: desc = line['Description'] hashtags = extract_hashtags(desc) txt_file.write(f"Description: {desc}\n") txt_file.write(f"Extracted Hashtags: {', '.join(hashtags)}\n\n") # Print to console for quick verification print(f"Description: {desc}") print(f"Extracted Hashtags: {', '.join(hashtags)}\n")
Option 2: Extract Hashtags from an Existing TXT File
If you already have a TXT file filled with descriptions, you can read it directly and pull hashtags:
import re def extract_hashtags(text): return re.findall(r'#\w+', text) # Path to your saved TXT file txt_path = "./data/2024-05-20_shares.txt" with open(txt_path, "r") as txt_file: for line in txt_file: line = line.strip() # Adjust this check based on how your TXT is formatted if line.startswith("Description:"): desc = line.replace("Description:", "").strip() hashtags = extract_hashtags(desc) print(f"Hashtags found: {', '.join(hashtags)}")
Quick Breakdown of the Regex
#: Matches the literal hashtag symbol.\w+: Matches one or more word characters (letters, numbers, underscores)—this covers all valid TikTok hashtag characters.
If you ever need to handle edge cases (like hashtags with accented characters), you can expand the regex to r'#[\w\dÀ-ÿ]+', but #\w+ works for the vast majority of real-world TikTok posts.
内容的提问来源于stack exchange,提问作者user12372145

