You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从超450万行异常CSV文件中提取Tweet ID的技术方案咨询

Extract Tweet IDs from Large, Malformed CSV

Since your CSV is massive (4.5M+ lines) with potential formatting issues (newlines in columns, special characters), treating it as raw text is the most reliable approach. Below are efficient solutions in bash, Perl, and Python that focus on extracting numeric Tweet IDs (assuming they're 15-20 digits long, which covers modern Twitter IDs—adjust the regex if your IDs are shorter).

Bash Solution

Bash's grep is lightning-fast for large files and requires no scripting. This command extracts all numeric sequences matching the Tweet ID length range and writes them to a new file:

grep -oE '[0-9]{15,20}' your_input.csv > tweet_ids.txt
  • Pros: No dependencies, runs in seconds even on huge files.
  • Note: If your CSV has other long numeric values (like timestamps), adjust the digit range to match your known Tweet ID length to avoid false positives.

Perl Solution

Perl is optimized for text processing and handles large files efficiently by processing line-by-line (avoids loading the entire file into memory):

#!/usr/bin/perl
use strict;
use warnings;

# Open input and output files
open my $in_fh, '<', 'your_input.csv' or die "Couldn't open input file: $!";
open my $out_fh, '>', 'tweet_ids.txt' or die "Couldn't open output file: $!";

# Iterate over each line, extract all matching Tweet IDs
while (my $line = <$in_fh>) {
    while ($line =~ /(\d{15,20})/g) {
        print $out_fh "$1\n";
    }
}

# Clean up file handles
close $in_fh;
close $out_fh;
  • Pros: Extremely fast, memory-efficient, and flexible (easily tweak the regex if needed).
  • Usage: Save as extract_ids.pl, make executable with chmod +x extract_ids.pl, then run ./extract_ids.pl.

Python Solution

Python is readable and cross-platform, and this script processes the file line-by-line to avoid memory overload:

import re

# Define the regex pattern for Tweet IDs (adjust digit range if needed)
TWEET_ID_PATTERN = re.compile(r'\d{15,20}')

# Open files in context managers (auto-closes them)
with open('your_input.csv', 'r', encoding='utf-8', errors='ignore') as infile, \
     open('tweet_ids.txt', 'w') as outfile:
    
    for line in infile:
        # Find all matching IDs in the current line
        tweet_ids = TWEET_ID_PATTERN.findall(line)
        for tid in tweet_ids:
            outfile.write(f"{tid}\n")
  • Pros: Easy to modify, handles encoding issues gracefully with errors='ignore', works on Windows/macOS/Linux.
  • Note: If you know the exact column where Tweet IDs are stored, you could use a CSV parser that handles quoted fields (like csv.reader with quoting=csv.QUOTE_ALL), but this raw text approach is more reliable if the CSV is severely malformed.

内容的提问来源于stack exchange,提问作者wxs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:07:36