You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Pandas DataFrame.info()中提取内存占用值并赋值给变量

Extract Memory Usage from DataFrame.info() and Assign to Variable

Hey there! Let's break down how to pull that memory usage value from DataFrame.info() and assign it to a variable—you've got two solid options here, depending on your needs:

Method 1: Parse the output of DataFrame.info()

If you specifically need to extract the value directly from the output of info(), you can redirect its printed output to a buffer, then parse the text to grab the memory line.

Here's how to do it:

import pandas as pd
from io import StringIO

# Create a sample DataFrame to test with
df = pd.DataFrame({'numbers': range(10000), 'strings': ['test']*10000})

# Capture the output of df.info() into a string buffer
output_buffer = StringIO()
df.info(buf=output_buffer)
info_text = output_buffer.getvalue()

# Iterate through each line to find the memory usage entry
memory_usage = None
for line in info_text.split('\n'):
    if 'memory usage' in line.lower():
        # Split the line to get the value (e.g., "156.3 KB")
        memory_usage = line.split(':')[-1].strip()
        # If you want to split into numeric value and unit separately:
        memory_num = float(memory_usage.split()[0])
        memory_unit = memory_usage.split()[1]
        break

# Assign to your variable
total_memory = memory_usage
print(f"Extracted memory usage: {total_memory}")

Note: Keep in mind that Pandas might tweak the output format of info() across versions, so this parsing logic might need small adjustments if you're working with an older/newer release.

The info() method actually uses memory_usage() under the hood to calculate memory usage. Instead of parsing text, you can get the value directly via this API—this is the more reliable and cleaner approach.

import pandas as pd

df = pd.DataFrame({'numbers': range(10000), 'strings': ['test']*10000})

# Calculate total memory usage (use deep=True to account for object dtype overhead)
total_memory_bytes = df.memory_usage(deep=True).sum()

# Optional: Convert to human-readable format (matches what info() displays)
def format_memory(bytes_val):
    for unit in ['B', 'KB', 'MB', 'GB']:
        if bytes_val < 1024:
            return f"{bytes_val:.1f} {unit}"
        bytes_val /= 1024
    return f"{bytes_val:.1f} TB"

human_readable_memory = format_memory(total_memory_bytes)

# Assign to your variable
df_memory_bytes = total_memory_bytes
df_memory_human = human_readable_memory

print(f"Total memory (bytes): {df_memory_bytes}")
print(f"Total memory (human-readable): {df_memory_human}")

The deep=True flag is important here—it ensures we count the actual memory used by object-type columns (like strings), which is what info() shows by default in modern Pandas versions.


内容的提问来源于stack exchange,提问作者Rodrigo Teixeira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:51:51