条件填充Pandas DataFrame出现空DataFrame问题求助
Hey there! Let's break down why your multi-code filtering isn't working and fix it up, plus share some better approaches for handling that large text file of yours.
What's Wrong with Your Current Multi-Code Checks?
Your two attempts at multi-code filtering use incorrect comparison logic:
tag == ('AB-C711', 'AB-D702'): This checks iftagis exactly equal to the tuple itself, not if it's one of the values inside the tuple. No line will ever match this.tag == code_list: Similarly, this checks iftagis exactly equal to the entire list, which is impossible. That's why you end up with an empty DataFrame.
The right way to do this is use the in operator, which checks if tag exists within your target collection of codes.
Fixed Basic Code
First, let's adjust your existing code to use in instead. Also, since you have over 100 codes, converting your list to a set will make the in check way faster (sets use hash tables for O(1) lookups, vs. O(n) for lists):
import pandas as pd filename = "C:/Users/abcd/Downloads/abcd-xyz-433.txt" code_df = pd.read_excel('C:/Users/abcd/Downloads/xyz_codes.xlsx') code_list = code_df['codes'].tolist() # Convert list to set for faster lookups code_set = set(code_list) sample = [] with open(filename, 'r') as f: for line in f: # Extract tag, add strip() to remove any accidental whitespace/newlines tag = line[:45].split('|')[5].strip() # Check if tag is in our target set if tag in code_set: sample.append(line.split('|')) # Convert the filtered lines to a DataFrame result_df = pd.DataFrame(sample) print("Filtered data successfully loaded into DataFrame!")
Better Approach for Large Text Files
If your text file is really large, manual line-by-line reading works but isn't the most efficient. Pandas has built-in tools to handle this with chunked reading, which avoids loading the entire file into memory at once:
import pandas as pd filename = "C:/Users/abcd/Downloads/abcd-xyz-433.txt" code_df = pd.read_excel('C:/Users/abcd/Downloads/xyz_codes.xlsx') code_set = set(code_df['codes'].tolist()) # Adjust chunksize based on your available memory (e.g., 10k rows per chunk) chunksize = 10_000 filtered_chunks = [] # Read the file in chunks and filter each chunk for chunk in pd.read_csv(filename, sep='|', chunksize=chunksize): # Assuming the tag is in the 6th column (index 5) filtered_chunk = chunk[chunk.iloc[:, 5].str.strip().isin(code_set)] filtered_chunks.append(filtered_chunk) # Combine all filtered chunks into one DataFrame result_df = pd.concat(filtered_chunks, ignore_index=True) print(f"Final DataFrame contains {len(result_df)} matching rows!")
This approach has a few key benefits:
- Uses pandas' vectorized operations, which are faster than manual loops
- Chunked reading prevents memory overload with huge files
- The
isinmethod combined with a set keeps filtering efficient
Quick Tips to Avoid Headaches
- Always add
.strip()when extracting the tag—accidental spaces or newline characters can break your matches - If your text file has inconsistent line lengths,
pd.read_csvwill handle column alignment better than manualsplit(); you can addon_bad_lines='skip'to skip malformed rows if needed
内容的提问来源于stack exchange,提问作者Vishwas

