You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解析含空格与分号的CSV文件——时间字段解析错误求助

Fixing Time Field Parsing for Your Custom CSV Format

Got it, let's work through why your read_data() function is messing up the time field parsing here. First, let's break down the unique quirk of your CSV:

  • The file uses semicolons (;) as the main field separator, but the first field is a combined date-time string with a space (like 20000530 172700).
  • The most likely issues are either your function is treating spaces as field separators (splitting date and time into two separate fields), or it's using the wrong format to parse the combined date-time value.

Solution 1: Manual Parsing (For Custom read_data() Implementations)

If you're writing the parsing logic yourself, follow this step-by-step approach to keep the date-time field intact:

  1. Split the entire line by semicolons first—this ensures the full date-time string stays as the first field.
  2. Split that first field by the space to separate date and time components.
  3. Parse the date (YYYYMMDD) and time (HHMMSS) into a proper datetime object/struct.

Example in C++:

#include <sstream>
#include <vector>
#include <string>
#include <ctime>
#include <iostream>

int read_data(const std::string& line) {
    // Split line by semicolons to get all top-level fields
    std::vector<std::string> fields;
    std::stringstream line_stream(line);
    std::string field;
    while (std::getline(line_stream, field, ';')) {
        fields.push_back(field);
    }

    // Split date-time string into date and time parts
    std::string datetime_str = fields[0];
    size_t space_idx = datetime_str.find(' ');
    std::string date_part = datetime_str.substr(0, space_idx);
    std::string time_part = datetime_str.substr(space_idx + 1);

    // Parse date (YYYYMMDD) into tm struct
    tm datetime = {0};
    datetime.tm_year = std::stoi(date_part.substr(0,4)) - 1900; // Years since 1900
    datetime.tm_mon = std::stoi(date_part.substr(4,2)) - 1;    // Months are 0-indexed
    datetime.tm_mday = std::stoi(date_part.substr(6,2));

    // Parse time (HHMMSS)
    datetime.tm_hour = std::stoi(time_part.substr(0,2));
    datetime.tm_min = std::stoi(time_part.substr(2,2));
    datetime.tm_sec = std::stoi(time_part.substr(4,2));

    // Parse remaining numeric fields
    double open = std::stod(fields[1]);
    double high = std::stod(fields[2]);
    double low = std::stod(fields[3]);
    double close = std::stod(fields[4]);
    int volume = std::stoi(fields[5]);

    // Verify parsed data (replace with your logic)
    std::cout << "Parsed datetime: " << asctime(&datetime)
              << "Open: " << open << ", High: " << high << "\n";

    return 0;
}

Example in Python:

from datetime import datetime

def read_data(line):
    # Split line by semicolons to preserve the full date-time string
    fields = line.strip().split(';')
    # Parse the date-time string with the correct format
    parsed_dt = datetime.strptime(fields[0], '%Y%m%d %H%M%S')
    
    # Parse numeric fields
    open_price = float(fields[1])
    high_price = float(fields[2])
    low_price = float(fields[3])
    close_price = float(fields[4])
    volume = int(fields[5])
    
    # Verify parsing (replace with your logic)
    print(f"Processed: {parsed_dt}, Open: {open_price}, Volume: {volume}")
    return 0

# Test with a sample line
read_data("20000530 172700;0.930200;0.930200;0.930200;0.930200;0")

Solution 2: Use a CSV Library (Avoid Manual Splitting Errors)

If your language has a built-in CSV parser, explicitly set the delimiter to semicolons. This ensures the date-time string is read as a single, intact field.

Python csv Module Example:

import csv
from datetime import datetime

def read_data(file_path):
    with open(file_path, 'r') as csv_file:
        # Explicitly set delimiter to semicolon
        reader = csv.reader(csv_file, delimiter=';')
        for row in reader:
            # Row[0] is the full date-time string
            parsed_dt = datetime.strptime(row[0], '%Y%m%d %H%M%S')
            open_price = float(row[1])
            # Add your processing logic here
            print(f"Processed entry: {parsed_dt}")
    return 0

Key Mistakes to Avoid

  • Don't split on spaces first: This breaks the date-time field into two separate entries, throwing off all subsequent field positions.
  • Use the correct datetime format: For YYYYMMDD HHMMSS, use %Y%m%d %H%M%S (Python) or manual substring extraction (C++) instead of generic date formats that won't match.

内容的提问来源于stack exchange,提问作者M.E.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:35:33