技术问询:如何将网页非标准格式心脏病数据集导入Pandas?
Got it, let's tackle this tricky dataset import step by step! The Hungarian heart disease dataset from UCI has that annoying quirk where the web view splits each actual data row into 10 separate lines, so we need to fix that first before loading it into Pandas. Here's how to do it:
Step 1: Fetch the raw data
First, we'll pull the raw text content directly from the source using the requests library—this avoids getting tripped up by the web page's artificial line breaks. If you don't have requests installed, run pip install requests first.
import requests import pandas as pd from io import StringIO # Grab the raw data text url = "http://archive.ics.uci.edu/ml/machine-learning-databases/heart-disease/hungarian.data" response = requests.get(url) raw_lines = response.text.splitlines()
Step 2: Merge lines into actual data rows
Each real data entry is spread across 10 consecutive lines in the web display. We'll group the raw lines into chunks of 10, then concatenate each chunk into a single line (replacing line breaks with spaces to keep the column separators consistent):
# Group raw lines into chunks of 10 (each chunk = one full data row) merged_rows = [] for i in range(0, len(raw_lines), 10): chunk = raw_lines[i:i+10] # Join the chunk into one line, stripping extra whitespace from each line merged_row = ' '.join(line.strip() for line in chunk) merged_rows.append(merged_row) # Combine all merged rows into a single text block ready for Pandas clean_data = '\n'.join(merged_rows)
Step 3: Load into Pandas
Now we can use pd.read_csv with our cleaned text. We'll handle a few key details:
- Multiple spaces as column separators (
sep='\s+') ?characters as missing values (na_values='?')- Assign the official column names (the dataset doesn't include headers, so we'll use the names from UCI's documentation)
# Official column names from the UCI dataset description column_names = [ "age", "sex", "cp", "trestbps", "chol", "fbs", "restecg", "thalach", "exang", "oldpeak", "slope", "ca", "thal", "num" ] # Load the cleaned data into a DataFrame df = pd.read_csv( StringIO(clean_data), sep='\s+', na_values='?', names=column_names ) # Verify the result print(df.head())
Quick Notes:
- If you run into stray empty lines messing up the chunking, add a quick filter to skip empty lines before grouping:
raw_lines = [line for line in response.text.splitlines() if line.strip()] - The
numcolumn is the target variable—0 means no heart disease, 1-4 indicate different levels of disease presence.
内容的提问来源于stack exchange,提问作者TJE

