You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:如何将网页非标准格式心脏病数据集导入Pandas?

Importing the Hungarian Heart Disease Dataset into Pandas

Got it, let's tackle this tricky dataset import step by step! The Hungarian heart disease dataset from UCI has that annoying quirk where the web view splits each actual data row into 10 separate lines, so we need to fix that first before loading it into Pandas. Here's how to do it:

Step 1: Fetch the raw data

First, we'll pull the raw text content directly from the source using the requests library—this avoids getting tripped up by the web page's artificial line breaks. If you don't have requests installed, run pip install requests first.

import requests
import pandas as pd
from io import StringIO

# Grab the raw data text
url = "http://archive.ics.uci.edu/ml/machine-learning-databases/heart-disease/hungarian.data"
response = requests.get(url)
raw_lines = response.text.splitlines()

Step 2: Merge lines into actual data rows

Each real data entry is spread across 10 consecutive lines in the web display. We'll group the raw lines into chunks of 10, then concatenate each chunk into a single line (replacing line breaks with spaces to keep the column separators consistent):

# Group raw lines into chunks of 10 (each chunk = one full data row)
merged_rows = []
for i in range(0, len(raw_lines), 10):
    chunk = raw_lines[i:i+10]
    # Join the chunk into one line, stripping extra whitespace from each line
    merged_row = ' '.join(line.strip() for line in chunk)
    merged_rows.append(merged_row)

# Combine all merged rows into a single text block ready for Pandas
clean_data = '\n'.join(merged_rows)

Step 3: Load into Pandas

Now we can use pd.read_csv with our cleaned text. We'll handle a few key details:

  • Multiple spaces as column separators (sep='\s+')
  • ? characters as missing values (na_values='?')
  • Assign the official column names (the dataset doesn't include headers, so we'll use the names from UCI's documentation)
# Official column names from the UCI dataset description
column_names = [
    "age", "sex", "cp", "trestbps", "chol", "fbs", "restecg",
    "thalach", "exang", "oldpeak", "slope", "ca", "thal", "num"
]

# Load the cleaned data into a DataFrame
df = pd.read_csv(
    StringIO(clean_data),
    sep='\s+',
    na_values='?',
    names=column_names
)

# Verify the result
print(df.head())

Quick Notes:

  • If you run into stray empty lines messing up the chunking, add a quick filter to skip empty lines before grouping: raw_lines = [line for line in response.text.splitlines() if line.strip()]
  • The num column is the target variable—0 means no heart disease, 1-4 indicate different levels of disease presence.

内容的提问来源于stack exchange,提问作者TJE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:43:21