从MongoDB读取数据生成CSV时出现多余引号问题求助
Fixing Extra Quotes in Your MongoDB-to-CSV Export
Hey there, let's get that CSV output looking exactly how you want it! The extra quotes you're seeing are a mix of inefficient DataFrame construction and misconfigured CSV export settings. Here's how to fix it step by step:
What's Causing the Issue?
- Your loop using
docs.append(series_obj)is not just slow (especially with large datasets), but it's also introducing unintended string formatting issues as you manually handle each document. - You're using a comma
','as the CSV separator into_csv, but your desired output uses semicolons';'. - The default quoting behavior in pandas isn't aligned with your needs, leading to duplicate quotes around string fields.
Optimized Solution Code
Replace your entire export logic with this cleaner, more reliable version:
import pandas as pd import csv # Pull data from MongoDB and convert directly to DataFrame (no manual loops!) cursor = collection2.find() docs = pd.DataFrame(list(cursor)) # Convert MongoDB's ObjectId _id to a string (so it exports cleanly) if '_id' in docs.columns: docs['_id'] = docs['_id'].astype(str) # Export to CSV with the exact formatting you want docs.to_csv( "test4.csv", sep=';', # Match your desired semicolon separator quoting=csv.QUOTE_NONNUMERIC, # Only add quotes to non-numeric values quotechar='"', # Use double quotes for quoted fields index=False, # Don't include pandas' default index column encoding='utf-8' # Ensure proper character encoding )
Why This Works
- Direct DataFrame Conversion: Skipping the manual
enumerateandappendsteps eliminates formatting errors from manual series creation, and it's way faster for larger datasets. - Proper Separator: Setting
sep=';'ensures columns are split with semicolons, matching your target format. - Controlled Quoting:
csv.QUOTE_NONNUMERICtells pandas to only wrap non-numeric fields in quotes—exactly what your desired output shows. No more random duplicate quotes! - Clean _id Handling: Converting the MongoDB ObjectId to a string prevents serialization issues and ensures it exports as a plain text value.
Quick Check
If you don't need the _id column in your CSV at all, just add docs = docs.drop(columns=['_id']) right after converting to DataFrame.
内容的提问来源于stack exchange,提问作者Evan
相关产品推荐
相关产品推荐

