如何使用Awk提取文本文件中Tweet含指定单词的对应日期(排除用户名中的匹配)
Fixing Your Awk Command to Extract Target Dates
Let's get this sorted out for you! Your current awk command has two key issues that are preventing it from working as expected:
- It uses exact matching (
==) for the Tweet column, which only catches rows where the tweet is exactly "Bob" instead of containing the word. - It doesn't exclude rows where the username is
@Bob.
Here's a corrected awk command that addresses both problems, plus some robustness improvements:
awk -F',' ' BEGIN { IGNORECASE = 1 # Ignore case for matching OFS = " " # Set output separator to space } # Exclude rows where username is @Bob (handles optional surrounding spaces) $5 !~ /^[[:space:]]*@Bob[[:space:]]*$/ && # Match rows where tweet contains the whole word "Bob" (not substrings like Bobby) $6 ~ /\bBob\b/ { dates = dates $3 OFS # Collect matching dates } END { sub(/ $/, "", dates) # Remove trailing space print dates # Output the result } ' Data.txt
Breakdown of the command:
-F',': Sets the field separator to a comma, which matches your input file's structure.IGNORECASE = 1: Makes the matching case-insensitive (so "bob", "BOB", etc. will also be caught if needed).$5 !~ /^[[:space:]]*@Bob[[:space:]]*$/: Ensures we skip any row where the username field ($5) is exactly@Bob, even if there are extra spaces around it (common in CSV-style files with inconsistent spacing).$6 ~ /\bBob\b/: Uses a regular expression with word boundaries (\b) to match the whole word "Bob" in the tweet field ($6). This avoids false matches for words like "Bobby" or "Bobcat".- Collecting and formatting output: We gather all matching dates into a variable, strip the trailing space at the end, and print them as a single space-separated line.
Testing with your sample data:
Running this command on your provided input will output:
02/03/09 01/09/10
Which matches your expected result perfectly.
内容的提问来源于stack exchange,提问作者homies
相关产品推荐
相关产品推荐

