如何处理列数可变的原始文本文件数据表及解析指定格式文本?
Great question! Let's break this down into two clear, actionable parts: handling your variable-column data table and parsing that structured release notes file.
First, you’ll need to identify how your data is delimited (common options for text files: commas, tabs, spaces, or fixed-width columns). Once you know that, here are reliable approaches:
1. Identify the Delimiter
- For a quick check, use the
filecommand (Linux/macOS) to guess the format:file your_data.txt - If space-separated, inspect sample lines to see if columns use single variable spaces or consistent multiple spaces (fixed width).
2. Command-Line Tools
- CSV/TSV with variable columns: Use
csvkit(install via pip) to inspect and manipulate. For example, count columns per line:
Or usecsvstat --count your_data.txt | grep "Number of fields"awkto flag lines with unexpected column counts:awk -F',' '{if (NF != 5) print NR ": " $0}' your_data.txt # Replace 5 with expected columns - Space-separated/fixed-width: Use
awkfor dynamic field handling. For fixed-width columns, extract values withsubstr:awk '{print substr($0,1,10), substr($0,11,15)}' your_data.txt
3. Python (Pandas) Approach
Pandas handles variable columns gracefully by inferring structure or filling gaps with NaN:
import pandas as pd # Read whitespace-separated data (handles variable spaces between columns) df = pd.read_csv( "your_data.txt", sep=r'\s+', # Regex for one or more whitespace characters engine='python', header=None, # Omit if your file has a header row on_bad_lines='warn' # Warn instead of crashing on mismatched columns ) # Preserve exact fields from every line (even with varying counts): rows = [] with open("your_data.txt", 'r') as f: for line in f: rows.append(line.strip().split()) # Split on whitespace df = pd.DataFrame(rows) # Pandas auto-fills missing columns with NaN
Your release file has a consistent sectioned format, so regex is perfect for extracting specific fields. Here’s how to do it:
1. Quick Command-Line Extractions
- Grab the release version and date:
grep "RELEASE VERSION:" release_notes.txt | awk '{print "Version: " $3, "Date: " substr($0, index($0,"(")+1, length($0)-index($0,"(")-1)}' - Extract the variable format note:
grep -A 2 "NOTES:" release_notes.txt | tail -n 2
2. Python Programmatic Parsing
This code parses all key sections into a reusable dictionary:
import re def parse_release_notes(file_path): release_data = {} with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # Extract release version number and date release_match = re.search(r'RELEASE VERSION: (\d+) \((.*?)\)', content) if release_match: release_data['version_number'] = release_match.group(1) release_data['version_date'] = release_match.group(2) # Extract study variable name var_match = re.search(r'VERSION STUDY VARIABLE: (.*?)\n', content) if var_match: release_data['study_variable'] = var_match.group(1).strip() # Extract question text question_match = re.search(r'QUESTION:\n--------- (.*?)\nNOTES:', content, re.DOTALL) if question_match: release_data['question'] = question_match.group(1).strip() # Extract notes (including variable format) notes_match = re.search(r'NOTES:\n------ (.*?)\nTYPE:', content, re.DOTALL) if notes_match: release_data['notes'] = notes_match.group(1).strip() # Extract data type type_match = re.search(r'TYPE:\n----- (.*?)\n', content) if type_match: release_data['data_type'] = type_match.group(1).strip() return release_data # Usage release_info = parse_release_notes("release_notes.txt") print(release_info)
Running this outputs a structured dictionary like:
{ 'version_number': '20150514', 'version_date': 'May 14, 2015', 'study_variable': 'Version Number Of Release', 'question': 'Version of Cumulative Data File', 'notes': '该变量在数据中的呈现形式为:ANES_cdf_VERSION:YYYY-mmm-DD,其中mmm为标准三字符月份缩写(Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec)。', 'data_type': 'Character-1' }
Key Tips
- For variable-column tables: If columns have specific meanings, define a schema and map fields explicitly instead of relying on auto-inference.
- For release notes: If the format changes across versions, adjust regex patterns to match new section structures.
内容的提问来源于stack exchange,提问作者oleks5412

