You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python/R等大数据友好语言统计多文件第10列元素频率?

Python Solution

This implementation processes files line-by-line to keep memory usage low, making it ideal for large datasets. It tracks total occurrences of each variant and the unique files they appear in.

import os
from collections import defaultdict

# Track variant counts and their associated files
counts = defaultdict(int)
variant_files = defaultdict(set)

# Get all text files in the current directory
txt_files = [f for f in os.listdir('.') if f.endswith('.txt')]

for filename in txt_files:
    with open(filename, 'r') as f:
        next(f)  # Skip header line
        for line in f:
            line = line.strip()
            if not line:
                continue
            parts = line.split('\t')
            if len(parts) < 10:
                continue  # Skip lines with missing columns
            variant = parts[9]
            counts[variant] += 1
            variant_files[variant].add(filename)

# Generate formatted output
output = []
output.append(f"{'Variant':<24} {'Nº of repeats':<16} Location")
for variant in counts:
    file_list = ','.join(sorted(variant_files[variant]))
    output.append(f"{variant:<24} {counts[variant]:<16} {file_list}")

# Write results to file
with open('var_freq.txt', 'w') as out_file:
    out_file.write('\n'.join(output) + '\n')

Key Details:

  • Uses line-by-line reading to avoid loading entire files into memory
  • Automatically deduplicates filenames per variant using sets
  • Skips empty/malformed lines to prevent errors
  • Sorts file lists for consistent output order

R Solution

This approach uses hash environments for fast lookups and processes files line-by-line to handle large datasets efficiently.

# Initialize hash maps to track counts and files
count_env <- new.env(hash = TRUE)
file_env <- new.env(hash = TRUE)

# Get all text files in the current directory
txt_files <- list.files(pattern = "\\.txt$", full.names = FALSE)

for (filename in txt_files) {
    con <- file(filename, "r")
    readLines(con, n = 1)  # Skip header
    
    while (length(line <- readLines(con, n = 1)) > 0) {
        line <- trimws(line)
        if (nchar(line) == 0) next
        
        parts <- strsplit(line, "\t")[[1]]
        if (length(parts) < 10) next
        
        variant <- parts[10]
        
        # Update count
        if (exists(variant, envir = count_env)) {
            count_env[[variant]] <- count_env[[variant]] + 1
        } else {
            count_env[[variant]] <- 1
            file_env[[variant]] <- character(0)
        }
        
        # Add filename if not already present
        if (!filename %in% file_env[[variant]]) {
            file_env[[variant]] <- c(file_env[[variant]], filename)
        }
    }
    close(con)
}

# Convert results to formatted output
variants <- ls(count_env)
counts <- sapply(variants, function(x) count_env[[x]])
files <- sapply(variants, function(x) paste(sort(file_env[[x]]), collapse = ","))

output_lines <- c(sprintf("%-24s %-16s %s", "Variant", "Nº of repeats", "Location"))
for (i in seq_along(variants)) {
    line <- sprintf("%-24s %-16d %s", variants[i], counts[i], files[i])
    output_lines <- c(output_lines, line)
}

writeLines(output_lines, "var_freq.txt")

Key Details:

  • Uses hash environments for O(1) lookup speed with large numbers of variants
  • Reads files incrementally to minimize memory usage
  • Deduplicates filenames per variant
  • Sorts file lists for consistent output

内容的提问来源于stack exchange,提问作者WindSur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 04:15:39