'Sample each group after pandas groupby
I know this must have been answered some where but I just could not find it.
Problem: Sample each group after groupby operation.
import pandas as pd
df = pd.DataFrame({'a': [1,2,3,4,5,6,7],
'b': [1,1,1,0,0,0,0]})
grouped = df.groupby('b')
# now sample from each group, e.g., I want 30% of each group
Solution 1:[1]
pandas >= 1.1: GroupBy.sample
This works like magic:
# np.random.seed(0)
df.groupby('b').sample(frac=.3)
a b
5 6 0
0 1 1
pandas <= 1.0.X
You can use GroupBy.apply with sample. You do not need to use a lambda; apply accepts keyword arguments:
df.groupby('b', group_keys=False).apply(pd.DataFrame.sample, frac=.3)
a b
5 6 0
0 1 1
Sources
This article follows the attribution requirements of Stack Overflow and is licensed under CC BY-SA 3.0.
Source: Stack Overflow
| Solution | Source |
|---|---|
| Solution 1 |
