基于另一DataFrame条件读取数据并生成指定文本的技术问题
Fixing Your Pandas Syntax Error & Achieving the Desired Output
First, let's break down the issues in your original code and walk through a corrected, working solution.
Issues in Your Original Code
- Typo: You wrote
sentancesinstead ofsentencesin the loop. - Missing Closing Parentheses: The
printstatement andbetween()call are missing closing parentheses, causing theSyntaxError: unexpected EOF. - Incorrect Iteration: Iterating directly over the DataFrame (
for x in sentences) loops over column names, not rows. You need to access individual rows to get eachstartandstopvalue. - Passing Series to
between(): Thebetween()method expects scalar values (single numbers), but you passed entire Series (sentences['start'],sentences['stop']), which won't work for row-wise filtering.
Corrected Solution
Here's the full working code to achieve your desired output:
import pandas as pd # Construct the words DataFrame words = pd.DataFrame() words['no'] = [1,2,3,4,5,6,7,8,9] words['word'] = ['cat', 'in', 'hat', 'the', 'dog', 'in', 'love', '!', '<3'] # Construct the sentences DataFrame sentences = pd.DataFrame() sentences['no'] = [1,2,3] sentences['start'] = [1, 4, 6] sentences['stop'] = [3, 5, 9] # Initialize a list to hold each sentence segment sentence_segments = [] # Iterate over each row in the sentences DataFrame for _, row in sentences.iterrows(): # Get start and stop values from the current row start_val = row['start'] stop_val = row['stop'] # Filter words where 'no' is between start and stop (inclusive) filtered_words = words[words['no'].between(start_val, stop_val, inclusive='both')]['word'] # Join the filtered words into a single string segment = ' '.join(filtered_words) sentence_segments.append(segment) # Join all segments with ' *** ' separator final_output = ' *** '.join(sentence_segments) # Print the result (matches your expected output) print(final_output) # Write the result to a file with open('output.txt', 'w') as f: f.write(final_output)
Key Details:
iterrows(): This method lets us loop through each row of thesentencesDataFrame, giving us access to the specificstartandstopvalues for each segment.between(start_val, stop_val, inclusive='both'): Filters thewordsDataFrame to only include rows wherenofalls within the current row's range. Theinclusive='both'parameter ensures both endpoints are included (this is default in newer Pandas versions, but we set it explicitly for clarity).- String Joining:
' '.join(filtered_words)combines individual words into a space-separated segment, and' *** '.join(sentence_segments)merges all segments into the final output string. - File Writing: Uses a context manager (
with open(...)) to safely write the result tooutput.txt.
Expected Output
Running this code will produce exactly what you wanted:
cat in hat *** the dog *** in love ! <3
内容的提问来源于stack exchange,提问作者Pythonuser
相关产品推荐
相关产品推荐

