如何移除DataFrame列名中的句号?多种尝试未生效
Let's figure out why that period isn't budging! Your issue is tricky because spaces work but periods don't—let's break this down step by step.
First, Verify the "Period" Is Actually a Period
Sometimes what looks like an English period (.) might be a full-stop character from another language (like Chinese 。) or a Unicode lookalike. Let's confirm the exact character in your problematic column name:
for col_name in df.columns: if "Src" in col_name: # Target your specific column "Src. of Business Contact Full Name" print(f"Column Name: {col_name}") print("Character Details:") for char in col_name: print(f" '{char}' | ASCII/Unicode Code: {ord(char)}")
An English period will have an ordinal value of 46. If it's a different code, you'll need to target that specific character in your replace call.
Try a More Explicit Renaming Approach
Your select method should work, but let's test using withColumnRenamed instead—sometimes this behaves more predictably with special characters, especially when dealing with bulk renames:
from pyspark.sql import functions as F # Define a helper function to clean column names def clean_column_name(name): # Replace spaces with underscores (optional, if you want to keep readability) cleaned = name.replace(' ', '_') # Remove periods explicitly cleaned = cleaned.replace('.', '') # Optional: Remove all other non-alphanumeric characters # import re # cleaned = re.sub(r'[^a-zA-Z0-9_]', '', cleaned) return cleaned # Loop through each column and rename for old_name in df.columns: new_name = clean_column_name(old_name) df = df.withColumnRenamed(old_name, new_name)
Fix Your Regex Logic
Your regex [^0-9a-zA-Z$]+ should match periods, but let's simplify it to explicitly target spaces and periods first, to rule out any regex quirks:
import re from pyspark.sql import functions as F # Remove just spaces and periods df = df.select([ F.col(col).alias(re.sub(r'[\s\.]', '', col)) for col in df.columns ]) # Or remove ALL non-alphanumeric characters (if that's your goal) # df = df.select([ # F.col(col).alias(re.sub(r'[^a-zA-Z0-9]', '', col)) # for col in df.columns # ])
Handle Spark's Backtick Quirk
If your column names were imported from a source that uses quoted headers (like CSV), Spark sometimes wraps columns with special characters in backticks internally. Try referencing the column with backticks in your F.col call:
from pyspark.sql import functions as F df = df.select([ F.col(f"`{c}`").alias(c.replace('.', '')) for c in df.columns ])
内容的提问来源于stack exchange,提问作者Marc

