如何将CSV数据集导入TensorFlow指定的placeholder中
Hey there, I get how frustrating it is when you’re crunched for time and tackling a new framework like TensorFlow out of the blue. Let’s cut straight to the solution since your data is already normalized and ready to go—no extra processing hoops needed.
First, let’s confirm the data split logic to make sure we’re aligned: your 184-column CSV row breaks down exactly into:
- Column 0: Today’s price (
close— single float) - Columns 1-3: 3 buy signals (
signals— 3 boolean values) - Columns 4-183: Past 180 days’ prices (
previous— 180 floats)
Option 1: Use Native CSV Reader (Simple, Small Data)
This builds on the code skeleton you already started with. We’ll convert each row to a numpy array, split it into your three placeholder shapes, and feed it directly:
import csv import numpy as np import tensorflow as tf # Define your placeholders (match your original code) close = tf.placeholder(tf.float32, name='close') signals = tf.placeholder(tf.bool, shape=[3], name='signals') previous = tf.placeholder(tf.float32, shape=[180], name='previous') # Add your model/training operations here (e.g., loss, optimizer) # Example dummy operation to test feeding: test_op = tf.concat([tf.expand_dims(close, 0), previous], axis=0) with tf.Session() as sess: sess.run(tf.global_variables_initializer()) with open('/BTC1.csv') as csv_file: csv_reader = csv.reader(csv_file, delimiter=',') # Skip header row if your CSV has one (remove this line if no header) next(csv_reader) line_count = 0 for row in csv_reader: # Convert the entire row to a numpy float array row_data = np.array(row, dtype=np.float32) # Split into your three variables: close_val = row_data[0] # Convert signal floats (0/1) to booleans signals_val = row_data[1:4].astype(bool) # Grab the remaining 180 values for past prices previous_val = row_data[4:] # Now feed these values into your placeholders and run your model # Replace test_op with your actual training/inference operation result = sess.run(test_op, feed_dict={ close: close_val, signals: signals_val, previous: previous_val }) line_count += 1 print(f"Processed {line_count} rows of data")
Option 2: Use TensorFlow Dataset API (Efficient, Large Data)
Since you checked the official Dataset docs but found them overwhelming, here’s a stripped-down version tailored to your use case—no unnecessary processing, just loading and splitting your prepped data:
import tensorflow as tf # Define your placeholders (if you still need them; Dataset can feed directly to models too) close = tf.placeholder(tf.float32, name='close') signals = tf.placeholder(tf.bool, shape=[3], name='signals') previous = tf.placeholder(tf.float32, shape=[180], name='previous') # Load CSV and define column types (all floats first, we'll convert signals later) dataset = tf.data.experimental.CsvDataset( '/BTC1.csv', record_defaults=[tf.float32] * 184, # Match 184 columns header=True # Set to False if your CSV has no header ) # Split each row into your three variables and convert signals to bool def parse_row(*row): close_tensor = row[0] # Stack signal columns and cast from float (0/1) to boolean signals_tensor = tf.cast(tf.stack(row[1:4]), tf.bool) previous_tensor = tf.stack(row[4:]) return close_tensor, signals_tensor, previous_tensor dataset = dataset.map(parse_row) # Optional: Batch your data for faster training (adjust batch size as needed) dataset = dataset.batch(32) # Create iterator to pull batches of data iterator = dataset.make_initializable_iterator() next_batch = iterator.get_next() with tf.Session() as sess: sess.run(iterator.initializer) try: while True: # Get a batch of data close_batch, signals_batch, previous_batch = sess.run(next_batch) # Feed the batch into your placeholders # Replace with your actual training operation sess.run(your_training_op, feed_dict={ close: close_batch, signals: signals_batch, previous: previous_batch }) except tf.errors.OutOfRangeError: print("All data has been processed!")
Key Notes to Avoid Headaches
- Header Handling: Don’t forget to skip the header row if your CSV has one—both examples include a way to do this.
- Signal Conversion: Since your CSV has float values for signals (I assume 0/1), we convert them to booleans with
astype(bool)(numpy) ortf.cast(TensorFlow). - Shape Matching: Double-check that
row_data[4:]gives exactly 180 values—184 total columns minus 4 (close + 3 signals) equals 180, which matches yourpreviousplaceholder shape.
This should get you feeding data into your placeholders so you can focus on training your model instead of wrestling with data loading.
内容的提问来源于stack exchange,提问作者J. Little

