如何使用pandas加载特殊格式的TXT数据?已有参考但格式不符
Got it, let's tackle this—your TXT data uses that line-by-line key-value format (super common in datasets like Amazon reviews), and standard pandas read_csv/read_table won't handle it directly. No worries, we can build a straightforward parser to load this into a clean DataFrame.
Step 1: Understand Your Data Structure
First, let's confirm the pattern from your sample:
product/productId: B003AI2VGA
review/userId: A141HP4LYPWMSR
review/profileName: Brian E. Erland "Rainbow Sphinx"
review/helpfulness: 7/7
review/score: 3.0
review/time: 1182729600
review/summary: "There Is So Much Darkness Now ~ Come For The Miracle"
review/text: Synopsis: On the daily trek from Juarez, Mexico to ...
Each review is made up of multiple lines, where each line follows [key]: [value]. Reviews are likely separated by blank lines (if not, we'll adjust the code below).
Step 2: Custom Parser Function
Here's a reusable function to load and process this data:
import pandas as pd def load_review_dataset(file_path): reviews = [] current_review = {} with open(file_path, 'r', encoding='utf-8') as file: for line in file: stripped_line = line.strip() # Blank line means we've finished one review if not stripped_line: if current_review: reviews.append(current_review) current_review = {} continue # Split only at the FIRST colon (avoids breaking values with colons like review/text) key, value = stripped_line.split(':', 1) # Clean up extra spaces and quotes around values clean_key = key.strip() clean_value = value.strip().strip('"') current_review[clean_key] = clean_value # Catch the last review if the file doesn't end with a blank line if current_review: reviews.append(current_review) # Convert list of review dictionaries to DataFrame return pd.DataFrame(reviews) # Load your data df = load_review_dataset("your_reviews.txt") print(df.head())
Step 3: Adjust for No Blank Lines Between Reviews
If your data doesn't have blank lines separating reviews (just continuous key-value lines), and you know each review has exactly 8 fields (like your sample), modify the function to group lines by field count:
def load_review_dataset_no_blank_lines(file_path): reviews = [] current_review = {} fields_per_review = 8 field_counter = 0 with open(file_path, 'r', encoding='utf-8') as file: for line in file: stripped_line = line.strip() if not stripped_line: continue key, value = stripped_line.split(':', 1) clean_key = key.strip() clean_value = value.strip().strip('"') current_review[clean_key] = clean_value field_counter += 1 # When we hit the field count, save the review and reset if field_counter == fields_per_review: reviews.append(current_review) current_review = {} field_counter = 0 # Add any incomplete review at the end (if total lines aren't a multiple of 8) if current_review: reviews.append(current_review) return pd.DataFrame(reviews)
Step 4: Clean Up Data Types
Once loaded, you'll want to convert fields to their proper types (e.g., scores to floats, timestamps to dates):
# Convert review score to float df['review/score'] = df['review/score'].astype(float) # Convert Unix timestamp to readable datetime df['review/time'] = pd.to_datetime(df['review/time'], unit='s') # Check the updated data types print(df.dtypes)
This should give you a fully functional DataFrame ready for analysis!
内容的提问来源于stack exchange,提问作者Coding_Rabbit

