You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pandas加载特殊格式的TXT数据?已有参考但格式不符

Loading Custom Key-Value TXT Data with Pandas

Got it, let's tackle this—your TXT data uses that line-by-line key-value format (super common in datasets like Amazon reviews), and standard pandas read_csv/read_table won't handle it directly. No worries, we can build a straightforward parser to load this into a clean DataFrame.

Step 1: Understand Your Data Structure

First, let's confirm the pattern from your sample:

product/productId: B003AI2VGA
review/userId: A141HP4LYPWMSR
review/profileName: Brian E. Erland "Rainbow Sphinx"
review/helpfulness: 7/7
review/score: 3.0
review/time: 1182729600
review/summary: "There Is So Much Darkness Now ~ Come For The Miracle"
review/text: Synopsis: On the daily trek from Juarez, Mexico to ...

Each review is made up of multiple lines, where each line follows [key]: [value]. Reviews are likely separated by blank lines (if not, we'll adjust the code below).

Step 2: Custom Parser Function

Here's a reusable function to load and process this data:

import pandas as pd

def load_review_dataset(file_path):
    reviews = []
    current_review = {}
    
    with open(file_path, 'r', encoding='utf-8') as file:
        for line in file:
            stripped_line = line.strip()
            
            # Blank line means we've finished one review
            if not stripped_line:
                if current_review:
                    reviews.append(current_review)
                    current_review = {}
                continue
            
            # Split only at the FIRST colon (avoids breaking values with colons like review/text)
            key, value = stripped_line.split(':', 1)
            # Clean up extra spaces and quotes around values
            clean_key = key.strip()
            clean_value = value.strip().strip('"')
            
            current_review[clean_key] = clean_value
        
        # Catch the last review if the file doesn't end with a blank line
        if current_review:
            reviews.append(current_review)
    
    # Convert list of review dictionaries to DataFrame
    return pd.DataFrame(reviews)

# Load your data
df = load_review_dataset("your_reviews.txt")
print(df.head())

Step 3: Adjust for No Blank Lines Between Reviews

If your data doesn't have blank lines separating reviews (just continuous key-value lines), and you know each review has exactly 8 fields (like your sample), modify the function to group lines by field count:

def load_review_dataset_no_blank_lines(file_path):
    reviews = []
    current_review = {}
    fields_per_review = 8
    field_counter = 0
    
    with open(file_path, 'r', encoding='utf-8') as file:
        for line in file:
            stripped_line = line.strip()
            if not stripped_line:
                continue
            
            key, value = stripped_line.split(':', 1)
            clean_key = key.strip()
            clean_value = value.strip().strip('"')
            
            current_review[clean_key] = clean_value
            field_counter += 1
            
            # When we hit the field count, save the review and reset
            if field_counter == fields_per_review:
                reviews.append(current_review)
                current_review = {}
                field_counter = 0
        
        # Add any incomplete review at the end (if total lines aren't a multiple of 8)
        if current_review:
            reviews.append(current_review)
    
    return pd.DataFrame(reviews)

Step 4: Clean Up Data Types

Once loaded, you'll want to convert fields to their proper types (e.g., scores to floats, timestamps to dates):

# Convert review score to float
df['review/score'] = df['review/score'].astype(float)

# Convert Unix timestamp to readable datetime
df['review/time'] = pd.to_datetime(df['review/time'], unit='s')

# Check the updated data types
print(df.dtypes)

This should give you a fully functional DataFrame ready for analysis!

内容的提问来源于stack exchange,提问作者Coding_Rabbit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:28:06