如何用Python/R等大数据友好语言统计多文件第10列元素频率?
Python Solution
This implementation processes files line-by-line to keep memory usage low, making it ideal for large datasets. It tracks total occurrences of each variant and the unique files they appear in.
import os from collections import defaultdict # Track variant counts and their associated files counts = defaultdict(int) variant_files = defaultdict(set) # Get all text files in the current directory txt_files = [f for f in os.listdir('.') if f.endswith('.txt')] for filename in txt_files: with open(filename, 'r') as f: next(f) # Skip header line for line in f: line = line.strip() if not line: continue parts = line.split('\t') if len(parts) < 10: continue # Skip lines with missing columns variant = parts[9] counts[variant] += 1 variant_files[variant].add(filename) # Generate formatted output output = [] output.append(f"{'Variant':<24} {'Nº of repeats':<16} Location") for variant in counts: file_list = ','.join(sorted(variant_files[variant])) output.append(f"{variant:<24} {counts[variant]:<16} {file_list}") # Write results to file with open('var_freq.txt', 'w') as out_file: out_file.write('\n'.join(output) + '\n')
Key Details:
- Uses line-by-line reading to avoid loading entire files into memory
- Automatically deduplicates filenames per variant using sets
- Skips empty/malformed lines to prevent errors
- Sorts file lists for consistent output order
R Solution
This approach uses hash environments for fast lookups and processes files line-by-line to handle large datasets efficiently.
# Initialize hash maps to track counts and files count_env <- new.env(hash = TRUE) file_env <- new.env(hash = TRUE) # Get all text files in the current directory txt_files <- list.files(pattern = "\\.txt$", full.names = FALSE) for (filename in txt_files) { con <- file(filename, "r") readLines(con, n = 1) # Skip header while (length(line <- readLines(con, n = 1)) > 0) { line <- trimws(line) if (nchar(line) == 0) next parts <- strsplit(line, "\t")[[1]] if (length(parts) < 10) next variant <- parts[10] # Update count if (exists(variant, envir = count_env)) { count_env[[variant]] <- count_env[[variant]] + 1 } else { count_env[[variant]] <- 1 file_env[[variant]] <- character(0) } # Add filename if not already present if (!filename %in% file_env[[variant]]) { file_env[[variant]] <- c(file_env[[variant]], filename) } } close(con) } # Convert results to formatted output variants <- ls(count_env) counts <- sapply(variants, function(x) count_env[[x]]) files <- sapply(variants, function(x) paste(sort(file_env[[x]]), collapse = ",")) output_lines <- c(sprintf("%-24s %-16s %s", "Variant", "Nº of repeats", "Location")) for (i in seq_along(variants)) { line <- sprintf("%-24s %-16d %s", variants[i], counts[i], files[i]) output_lines <- c(output_lines, line) } writeLines(output_lines, "var_freq.txt")
Key Details:
- Uses hash environments for O(1) lookup speed with large numbers of variants
- Reads files incrementally to minimize memory usage
- Deduplicates filenames per variant
- Sorts file lists for consistent output
内容的提问来源于stack exchange,提问作者WindSur
相关产品推荐
相关产品推荐

