合并两个DataFrame遇问题:数据对齐与玩家ID格式修正求助
Hey there! Let's tackle those two small, solvable issues you're dealing with in your script—they're straightforward fixes once you know the tricks. I'll break each problem down with actionable code snippets tailored to your setup (since you're using requests, BeautifulSoup, and pandas).
1. Fixing DataFrame Alignment in Output
If your merged CSV displays perfectly but the pandas DataFrame looks misaligned in your terminal/console, this is almost always a pandas display setting issue. Pandas sometimes struggles with wide characters or column spacing by default. Try adding these lines before you print or display your DataFrame:
import pandas as pd # Fix alignment for characters that need extra width (like East Asian characters) pd.set_option('display.unicode.ambiguous_as_wide', True) pd.set_option('display.unicode.east_asian_width', True) # Adjust column width to prevent wrapping or truncation pd.set_option('display.max_colwidth', None) pd.set_option('display.width', 1000) # Tweak this number to match your console's width
These settings force pandas to calculate character widths correctly and expand the display area, which should make your DataFrame columns line up neatly just like your CSV does.
2. Converting Player ID from List Format to Plain Value
Seeing ['5452'] instead of just 5452 means your player ID column is storing values as lists (or string representations of lists). Let's fix this depending on which scenario you're in:
Scenario A: The ID is actually a list object
If you're storing the ID as a list in your DataFrame (e.g., you wrapped it in brackets during scraping), use apply() to extract the first element:
# Pull the first item from each list in the player_id column df['player_id'] = df['player_id'].apply(lambda x: x[0] if isinstance(x, list) and len(x) > 0 else x)
Scenario B: The ID is a string that looks like a list (e.g., "['5452']")
If the value is a string (maybe from parsing HTML that returned a list-like string), use ast.literal_eval() to convert it to a real list first:
import ast # Convert string to actual list, then extract the ID df['player_id'] = df['player_id'].apply(lambda x: ast.literal_eval(x)[0] if isinstance(x, str) else x)
Pro Tip for Scraping
You can avoid the list issue entirely at the source when collecting data with BeautifulSoup. Instead of wrapping the ID in a list like this:
player_id = [soup.find('div', class_='player-id').text] # Unnecessary list wrap
Just grab the text directly:
player_id = soup.find('div', class_='player-id').text.strip()
That way, you never store the ID as a list in the first place!
内容的提问来源于stack exchange,提问作者Nick

