基于R/Python的不同数据量txt与csv文件绘图技术问询
R Solutions
1. Importing Files & Selecting Specific Columns
Whether you’re working with CSV or TXT files, the core workflow is similar—though TXT files often need extra parameters to handle delimiters properly.
CSV Files: Use
read.csv()(base R) orread_csv()(from thereadrpackage, faster for large datasets) to import and pick columns directly:# Base R: Select columns by name csv_data <- read.csv("your_data.csv", usecols = c("timestamp", "value")) # readr (tidyverse): More efficient for big files, uses tidy syntax library(readr) csv_data <- read_csv("your_data.csv", col_select = c(timestamp, value))TXT Files: Use
read.table()(base R) orread_delim()(readr) and specify the delimiter (e.g., tab, space):# Base R: Tab-separated TXT, select columns 1 and 3 txt_data <- read.table("your_data.txt", sep = "\t", header = TRUE, usecols = c(1,3)) # readr: Flexible delimiter handling for space-separated TXT txt_data <- read_delim("your_data.txt", delim = " ", col_select = c(id, measurement))
Note: Handling single vs multiple files works exactly the same way—just repeat the import/column selection step for each file. The only difference is adjusting file-specific parameters (like delimiter for TXT).
2. Downsampling the Large CSV (Every n Rows)
To match the smaller 80k-row dataset, calculate n as the ratio of total rows to target rows: 3600000 / 80000 = 45. This means you’ll pick every 45th row.
Base R:
# Import full large dataset large_csv <- read.csv("large_data.csv") # Calculate n (round to ensure clean row counts) n <- round(nrow(large_csv) / 80000) # Select every nth row downsampled_csv <- large_csv[seq(1, nrow(large_csv), by = n), ]dplyr (tidyverse):
library(dplyr) downsampled_csv <- large_csv %>% slice(seq(1, n(), by = n))
Python Solutions
1. Importing Files & Selecting Specific Columns
Using pandas is the standard approach—it handles both CSV and TXT files seamlessly.
CSV Files:
import pandas as pd # Select columns by name during import csv_data = pd.read_csv("your_data.csv", usecols=["timestamp", "value"]) # Or select columns after importing the full dataset csv_data = pd.read_csv("your_data.csv") csv_data = csv_data[["timestamp", "value"]]TXT Files: Use
pd.read_csv()with the correct delimiter, orpd.read_table():# Space-separated TXT, select columns 0 and 2 (0-indexed) txt_data = pd.read_csv("your_data.txt", sep="\s+", usecols=[0, 2]) # Tab-separated TXT, select by column name txt_data = pd.read_table("your_data.txt", usecols=["id", "measurement"])
Note: Importing and selecting columns for multiple files follows the same logic as single files—you can loop through file paths and apply the same steps to each.
2. Downsampling the Large CSV (Every n Rows)
Calculate n the same way: 3600000 / 80000 = 45. Use pandas’ indexing to pick every nth row efficiently.
import pandas as pd # Import the large dataset large_csv = pd.read_csv("large_data.csv") # Calculate n n = round(len(large_csv) / 80000) # Select every nth row downsampled_csv = large_csv.iloc[::n] # Verify row count print(f"Downsampled row count: {len(downsampled_csv)}")
To save memory (critical for 3.6M rows), you can even downsample during import:
# Read only every nth row directly, skipping others downsampled_csv = pd.read_csv("large_data.csv", skiprows=lambda x: x % n != 0 and x != 0)
内容的提问来源于stack exchange,提问作者user9613333

