请求将数据集按指定列数生成规范表格并解析指定TXT文件内容
Got it, let's break down how to solve your two technical requests properly—converting that space-separated TXT dataset into a structured,规范表格 while preserving every technical term, number, special format, and even the truncated content at the end.
Step 1: Define Column Mapping
First, let's map each space-separated field in the TXT to meaningful column names (adjust these to match your pre-specified column requirements):
- ID
- Firm
- LastName
- FirstInitial
- CreateDate
- CreateTime
- UpdateDate
- UpdateTime
- InitialRatingCode
- InitialRating
- CurrentRatingCode
- CurrentRating
- Value1
- Value2
- EffectiveDate
- EffectiveTime
Step 2: Python Script to Parse and Convert
Here's a straightforward script that reads the TXT content, splits each line correctly, and outputs a formatted Markdown table. It preserves all original content without modifying any details:
# Replace this with reading from your actual file if needed txt_content = """0003 MPARTNER SALZ S 150112 22:30:45 160304 08:38:13 2 BUY 2 BUY 12380 165426 150109 08:00:00 0003 SPROTTSE HUGHES S 140407 02:30:50 141120 13:55:06 2 BUY 2 BUY 3764 57379 140401 10:05:00 0003 SPROTTSE HUGHES S 141223 09:06:13 160715 08:42:56 3 MARKETPERFORM 3 HOLD 3764 57379 141223 08:02:00 001V MPARTNER PEARLSTEIN D 140821 02:44:05 150312 09:17:13 2 BUY 2 BUY 12380 163717 140820 08:16:00 001V MPARTNER PEARLSTEIN D 151016 15:07:40 160411 08:40:35 2 BUY 2 BUY 12380 16...""" # Split content into lines, skip empty lines lines = [line.strip() for line in txt_content.split("\n") if line.strip()] # Define your pre-specified columns (adjust as needed) columns = [ "ID", "Firm", "LastName", "FirstInitial", "CreateDate", "CreateTime", "UpdateDate", "UpdateTime", "InitialRatingCode", "InitialRating", "CurrentRatingCode", "CurrentRating", "Value1", "Value2", "EffectiveDate", "EffectiveTime" ] # Split each line into fields (whitespace-separated) data_rows = [line.split() for line in lines] # Generate Markdown table def build_markdown_table(cols, rows): # Header row table = f"| {' | '.join(cols)} |\n" # Separator row table += f"| {' | '.join(['---']*len(cols))} |\n" # Add data rows, handle truncated fields gracefully for row in rows: # Fill empty spots if a line is truncated (like the last one) padded_row = row + [""] * (len(cols) - len(row)) table += f"| {' | '.join(padded_row)} |\n" return table # Get the formatted table final_table = build_markdown_table(columns, data_rows) print(final_table)
Step 3: Resulting Formatted Table
Running the script will output this clean,规范表格 that retains every original detail:
| ID | Firm | LastName | FirstInitial | CreateDate | CreateTime | UpdateDate | UpdateTime | InitialRatingCode | InitialRating | CurrentRatingCode | CurrentRating | Value1 | Value2 | EffectiveDate | EffectiveTime |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0003 | MPARTNER | SALZ | S | 150112 | 22:30:45 | 160304 | 08:38:13 | 2 | BUY | 2 | BUY | 12380 | 165426 | 150109 | 08:00:00 |
| 0003 | SPROTTSE | HUGHES | S | 140407 | 02:30:50 | 141120 | 13:55:06 | 2 | BUY | 2 | BUY | 3764 | 57379 | 140401 | 10:05:00 |
| 0003 | SPROTTSE | HUGHES | S | 141223 | 09:06:13 | 160715 | 08:42:56 | 3 | MARKETPERFORM | 3 | HOLD | 3764 | 57379 | 141223 | 08:02:00 |
| 001V | MPARTNER | PEARLSTEIN | D | 140821 | 02:44:05 | 150312 | 09:17:13 | 2 | BUY | 2 | BUY | 12380 | 163717 | 140820 | 08:16:00 |
| 001V | MPARTNER | PEARLSTEIN | D | 151016 | 15:07:40 | 160411 | 08:40:35 | 2 | BUY | 2 | BUY | 12380 | 16... |
Key Details to Note
- All technical terms (e.g.,
MARKETPERFORM,HOLD), numbers with leading zeros (e.g.,0003,001V), time stamps (e.g.,22:30:45), and even the truncated16...are preserved exactly as in the original TXT. - If your pre-specified columns have different names, order, or count, simply modify the
columnslist to match your requirements. - For large files, replace the hardcoded
txt_contentwith file reading logic:with open("your_input_file.txt", "r") as f: txt_content = f.read()
内容的提问来源于stack exchange,提问作者JB_G

